Topic · L2 · Chips & Hardware
Seven generations, one trajectory: more memory, bigger racks.
H100 to Rubin Ultra in five years. The number to watch isn't teraflops — it's the coherent rack domain, the GPUs that behave as one. It's what makes a trillion-parameter MoE trainable at all.
The rack domain staircase: 8 → 72 → 576
The rack domain column is the story. 8 held for three generations, then jumped to 72 (GB200) and is heading to 576 (Rubin Ultra). That staircase is what unlocked ultra-sparse trillion-param models.
Bandwidth ladder — link bandwidth by tier
- Tier 1 (per-link)20000 GB/s
- Tier 1 (per-link)25000 GB/s
- Tier 2 (per-link)7700 GB/s
- Tier 2 (per-link)13000 GB/s
- Tier 3 (per-link)10000 GB/s
- Tier 3 (per-link)10000 GB/s
- Tier 4 (per-link)1800 GB/s
- Tier 4 (per-link)3600 GB/s
on-die SRAM (10s TB/s) → in-rack NVLink (single-digit TB/s) → cross-rack (~100 GB/s). Each step an order of magnitude slower — which is exactly why MoE won.
Open full Bandwidth Ladder →The bandwidth ladder, folded in. on-die SRAM (10s TB/s) » in-rack NVLink (single-digit TB/s) » cross-rack (~100 GB/s). Each step an order of magnitude slower — which is exactly why MoE (mostly-idle experts you can park in cheap memory) won.