For the first time, the AMD roadmap is a rack, not a card. That changes how 2027 capacity gets quoted — and it puts a floor under how cheap the current Instinct generation can get for inference buyers who move now.
At Hot Chips 2026, AMD walked through the Instinct MI400 family and the Helios rack in detail. The MI455X carries 432 GB of HBM4 at 23.3 TB/s (up from 288 GB of HBM3E on the MI355X), 256 work-group processors, 192 MB of L2, and roughly 40 petaflops of MXFP4 compute per GPU. Helios packs 72 of them into a single liquid-cooled rack: 2.9 exaflops, 31 TB of HBM4, 1.7 PB/s of memory bandwidth, 260 TB/s of scale-up and 43 TB/s of scale-out bandwidth. AMD has not committed to a shipping date beyond systems it “expects to ship”; the roadmap points to 2027.
The presentation sat next to NVIDIA’s own Vera Rubin NVL72 session, which is the point. Helios is AMD’s first credible answer at the rack-scale level where the hyperscale and neocloud market now buys, and it lands with 6th-gen EPYC hosts and an open networking stack. Meanwhile the current generation is already in the field: MI355X posts competitive MLPerf inference numbers against H200, and MI300X rentals are clearing at $1.44–$5.20 per GPU-hour against H100 at roughly $2.89 and up, with B200/B300 still carrying a supply premium.
AMD is not going to displace NVIDIA in training allocation in 2027. But a second rack-scale source with more HBM per GPU changes the negotiation for anyone whose workload is memory-bound inference — which, increasingly, is everyone running agents at scale.
What it means if you’re buying: Two things at once. For 2027 commitments, insist on a two-stack quote: Helios/MI455X against Vera Rubin NVL72, priced per token delivered rather than per GPU-hour, because HBM capacity per rack is what sets inference cost and AMD has more of it. For capacity you need this quarter, the Instinct discount is the opportunity — MI300X/MI355X clusters are renting at a fraction of Blackwell rates, and for inference-heavy, memory-bound workloads the gap in delivered throughput is far smaller than the gap in price.
What to do about it: Our GPU book today is NVIDIA-weighted — LYNX (2K× H200, US) and AX3-PTR-5001 (H200 141 GB hardware) for Hopper inference, ORION 1024 and MAGNETAR for B300 terms — and buyers who want an Instinct quote alongside should ask; we source both stacks. Operators sitting on idle MI300X or MI355X: there is a price gap on the market for exactly that capacity, and the Neocloud Watch shows which of the 20 largest neoclouds are already running EPYC and Instinct. List it behind a codename and we’ll put it in front of inference buyers this week.
Original reporting · servethehome.com ↗ · Summary and analysis by Ax3