Adaptation Curve

Topic · L2 · Chips & Hardware

Seven generations, one trajectory: more memory, bigger racks.

H100 to Rubin Ultra in five years. The number to watch isn't teraflops — it's the coherent rack domain, the GPUs that behave as one. It's what makes a trillion-parameter MoE trainable at all.

The rack domain staircase: 8 → 72 → 576

2020202220242026202883272576GPUs / coherent domainA100/H100 · 8GB200 NVL72 · 72Rubin Ultra NVL576 · 5769× jump8× jump
The rack domain column is the story. 8 held for three generations, then jumped to 72 (GB200) and is heading to 576 (Rubin Ultra). That staircase is what unlocked ultra-sparse trillion-param models.

Bandwidth ladder — link bandwidth by tier

  • Tier 1 (per-link)20000 GB/s
  • Tier 1 (per-link)25000 GB/s
  • Tier 2 (per-link)7700 GB/s
  • Tier 2 (per-link)13000 GB/s
  • Tier 3 (per-link)10000 GB/s
  • Tier 3 (per-link)10000 GB/s
  • Tier 4 (per-link)1800 GB/s
  • Tier 4 (per-link)3600 GB/s

on-die SRAM (10s TB/s) → in-rack NVLink (single-digit TB/s) → cross-rack (~100 GB/s). Each step an order of magnitude slower — which is exactly why MoE won.

Open full Bandwidth Ladder →
The bandwidth ladder, folded in. on-die SRAM (10s TB/s) » in-rack NVLink (single-digit TB/s) » cross-rack (~100 GB/s). Each step an order of magnitude slower — which is exactly why MoE (mostly-idle experts you can park in cheap memory) won.