MLPerf inference results published this week show AMD's latest Instinct accelerator pulling ahead on LLM inference — though training remains an NVIDIA stronghold.
New MLPerf inference submissions show AMD's Instinct MI355X delivering up to 1.3× the throughput of NVIDIA's H200 on several large-language-model serving benchmarks, the strongest showing yet for AMD's accelerator line.
The gains concentrate in memory-bandwidth-bound inference at long context lengths, where the MI355X's HBM capacity advantage tells. Training submissions, by contrast, continue to favor NVIDIA's ecosystem, where software maturity and interconnect scale remain decisive.
Analysts noted that inference now represents the fastest-growing share of AI compute demand, which makes a credible second source commercially meaningful even without training parity.
What it means: For inference-heavy buyers, MI3XX capacity just became a legitimate lever on price — a credible alternative pressures H200 rates even if you never deploy it.
What to do about it: Quote both stacks. Ax3 brokers MI3XX and H200 capacity side by side; buyers running serving workloads should ask us to price the same requirement across both before committing.
Original reporting · theregister.com ↗ · Summary and analysis by Ax3