AI-Portable
Article image for Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI Articles
Not Applicable

Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI

What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale. These factories are now tasked with…

Condensed by AI-Portable from Editorial queue.

The NVIDIA Rubin GPU, central to the Vera Rubin platform, delivers up to 10x agentic throughput per unit energy compared to Blackwell, enabled by 336 billion transistors, 224 SMs, 896 Tensor Cores with expanded precision, third-generation Transformer Engine, and 288 GB HBM4 memory providing 22 TB/s bandwidth.

Innovations such as enhanced Tensor Memory Accelerator, inline descriptor updates, activation sparsity, adaptive compression, and fine-grained dependent kernel triggering optimize MoE scaling, long-context attention, and minimize kernel transition latency, collectively maximizing tokens/sec and tokens/watt for agentic inference workloads.

At rack scale, NVIDIA Vera Rubin NVL72 integrates liquid cooling, power smoothing with DSX MaxLPS, cable-free MGX architecture, and hot-swappable NVLink switch trays, enabling multitrillion-parameter models, up to 40% more GPUs within the same power envelope, and resilient, high-throughput agentic AI supercomputing.

AI-generated content may summarize information incompletely. Verify important information. Learn more

What began as discrete AI model training and human-facing chat interfaces has evolved into always-on AI factories dedicated to producing intelligence at scale. These factories are now tasked with powering agentic workflows that reason, plan, use tools, verify intermediate results, and execute complex multistep tasks across vast contexts.

Agentic workloads are not defined by a single prompt and response, but by sustained inference across many reasoning steps. They demand low per-step latency, high decode throughput, efficient long-context attention, large KV cache capacity, and the ability to scale models across tightly coupled GPU domains. The data center must be reimagined as a single unit of compute, a vision realized with the NVIDIA Vera Rubin platform .

At the core of the platform is the NVIDIA Rubin GPU, designed to deliver up to 10x more agentic throughput per unit of energy than NVIDIA Blackwell (Figure 1). Enhanced Tensor Cores with expanded precision flexibility, a new HBM4 memory subsystem, and the third-generation Transformer Engine—delivering up to 50 petaflops of NVFP4 performance—work together to accelerate agentic workloads efficiently.

This post examines how the NVIDIA Rubin GPU and its co-designed scale-up system address the end-to-end bottlenecks of agentic inference, from data movement and compute efficiency to long-context execution and rack-scale deployment.

Original source ↗