Decode throughput modeled at batch 1, 80B parameters, FP8 mode (BF16-native decode is 7,219 tokens/s). Rubin / MI455X throughput estimated from 2026 vendor bandwidth and TDP specs; GPU cost from Morgan Stanley VR200 NVL72 rack estimates (÷ 72); Sophon figures modeled, pre-silicon.
Roofline model, 80B FP8 decode, single accelerator. Etched Sohu and Cerebras WSE-3 figures estimated from public claims (no published BOM or power data); Cerebras shown as a 2-wafer SRAM-resident system (~46 kW, ~120× Sophon’s power). Dotted lines: Sophon node roadmap, 22 nm→N3 (projected / extrapolated). Whitepaper §5.A.5c, Figure 8b.
To train a 100T MoE in a 1 GW build — how long it takes and the fleet you must buy, Sophon (28 nm) vs HBM4 GPUs:
All three fill the same 1 GW build and train the same ≈ 2.5×10²⁸-FLOP, 48× MoE. At 28 nm, Sophon’s ≈3.6–3.8× more, lower-power dies (472 W vs 1,700–1,800 W) deliver roughly the same aggregate FLOP/s — near-parity in training time (≈4–5 months for all three) — but the fleet costs ~2.9–3.7× less to build; the 7 nm node then finishes ≈2.7–3.3× faster (≈1.5 months). Batched training is compute-bound and near per-FLOP parity — the 174× advantage above is on serving, which is memory-bound.
2D-TMD has already flown — a wafer-scale MoS₂ system ran nine months on orbit (Nature, 2026). The accelerator that survives without shielding mass becomes the default for space compute. Read the full analysis · why space needs this chip