NVIDIA tunes Vera Rubin for the agent era of much heavier token demand
NVIDIA tunes Vera Rubin for the agent era of much heavier token demand
NVIDIA published a set of technical announcements around Vera Rubin and its rack-scale inference systems, with a clear infrastructure message: AI agents do not behave like short chatbot sessions. They generate more tokens, maintain long contexts, call tools and may coordinate subagents. If that pattern becomes normal in companies, the relevant cost will not be just having powerful accelerators, but producing more useful work per watt, per rack and per dollar.
What happened
On August 24, NVIDIA published performance data and architecture details for Vera Rubin NVL72 aimed at agent workloads. The company said that, in internal measurements using real-world agentic coding trajectories preserved in SemiAnalysis' AgentX benchmark, Vera Rubin NVL72 delivered up to 30 times more throughput per megawatt than GB300 NVL72 on those workloads. It also said token costs could be up to 35 times lower.
The figure should be read carefully: it is a result measured by NVIDIA, not an independent audit published as a peer-reviewed paper. Even so, it points to the technical direction the company wants to sell as the market shifts from massive training toward intensive inference.
In a second announcement, NVIDIA said Groq 3 LPX for Vera Rubin is in full production and positioned it as a component for fast token generation in long-context workloads. According to the company, in an Artificial Analysis benchmark using Gemma 4 31B, the system reached 3,400 output tokens per second for 100,000-token context use cases. The same post mentions related adoption or deployments by SpaceXAI, CoreWeave and Nebius.
Why it matters
NVIDIA's thesis is that agents change the economics of inference. An assistant summarizing one paragraph may finish in a few turns. An agent that researches, queries databases, reads news, reviews documents, invokes tools, calls subagents and synthesizes a recommendation accumulates context and tokens at each step. NVIDIA cites OpenRouter data to argue that agentic workloads can consume 15 times more tokens than a simple chat request.
That makes energy efficiency and network architecture central editorial issues. A power-constrained AI factory cannot grow indefinitely by simply adding racks. It needs more work per watt, better utilization, lower-latency networking and software that keeps the infrastructure busy.
What changes in the architecture
NVIDIA's third announcement focuses on XPUs and NVLink Fusion. The company argues that hyperscalers and AI-native companies designing custom silicon do not only need a specialized chip; they also need scale-up and scale-out networking, rack design, production software and a supplier ecosystem able to operate at AI-factory scale. The proposal is to combine custom XPUs with mature NVIDIA infrastructure to reduce the time and risk of bringing those designs into production.
For large customers, the promise is flexibility: use custom silicon where it creates differentiation, while relying on proven components for the rest of the platform. For NVIDIA, the message also protects its role even as major buyers explore in-house chips.
What remains unclear
It remains unclear how these results will compare once more labs, clouds and analysts test equivalent agent workloads under public conditions. It should also not be assumed that every company will need AI-factory-scale infrastructure. Many applications will still be handled with smaller models, context compression, caching or less autonomous workflows.
The news does confirm one trend: AI competition is moving from the isolated model toward the full system that supports long, expensive, tool-using agents. In that layer, energy, memory, networking, software and token cost are becoming as important as the model name.
Written by Nova Rivera — Product and automation perspective.
Sources consulted
NVIDIA Blog: Vera Rubin NVL72 efficiency for AI agents; NVIDIA Blog: Vera Rubin LPX/Spectrum-X/NVLink Fusion; NVIDIA Blog: XPUs and AI factories. Exact canonical links appear in the Sources section below.
Sources: NVIDIA Blog, NVIDIA Blog, NVIDIA Blog