BlueOnyx
HardwareInfrastructureDatacenterSemiconductorsVendors

72 GPUs, One Rack: AMD Challenges NVIDIA's AI Infrastructure Dominance

Théodore BaillyPublished on 1 octobre 20265 min read
Carte graphique GPU sur fond jaune vif

Introduction

For the past five years, deploying AI infrastructure at scale has meant, in practice, negotiating with a single vendor. The NVIDIA ecosystem — DGX, HGX, NVLink, CUDA — has dictated terms to buyers ranging from hyperscalers to enterprise IT leaders across manufacturing and beyond. On September 30, 2026, a $1.2 billion order placed by Vultr with HPE changed that equation: this is the first commercial deployment of the AMD Helios AI Rack, a fully integrated rack-scale platform that stakes a credible claim against the green empire on its own turf.

A Rack Built as a System, Not a Stack of Components

What sets the AMD Helios apart is not raw accelerator performance alone — though that performance is real. Each rack integrates 72 AMD Instinct MI455X GPUs, built on the CDNA 5 architecture fabbed at 2 nanometers by TSMC, with 432 GB of HBM4 memory and 23.3 TB/s of memory bandwidth per chip. Rack-level compute density reaches 1.4 ExaFLOPS at FP8, with 31 TB of HBM4 pooled across the system.

But the Helios philosophy goes beyond silicon. AMD EPYC Venice server processors — the first x86 chip in volume production on TSMC's 2nm node — handle local orchestration. AMD Pensando Vulcano NICs manage the fabric. The HPE Juniper Networking QFX5252 switch handles inter-rack connectivity. The entire system is liquid-cooled end-to-end. This is no longer a bag of components for your team to integrate: it is a validated, fully supported, deliverable unit.

HPE plays the role of reference integrator here, leveraging its acquisition of Juniper Networks to own the full network-to-compute stack. The competitive logic is straightforward: NVIDIA offers its DGX SuperPOD under the same packaged model. AMD now counters with an equivalent proposition — backed by a distribution partner with serious reach.

What the First Commercial Order Signals

The contract size — $1.2 billion, for deployments across Vultr's U.S. data centers — is a clear market signal. It means a mid-market cloud provider evaluated the AMD Helios ecosystem as mature enough for production, and chose not to wait. This is precisely the kind of third-party validation that large operators and enterprise IT leaders watch for before diversifying their vendor base.

One detail deserves close attention from infrastructure teams: the AMD Helios relies on open standards for GPU interconnect, where NVIDIA depends on its proprietary NVLink. If that open architecture holds up under sustained production workloads, it structurally shifts buyer negotiating leverage and reduces the risk of long-term single-vendor lock-in.

What IT Leaders Should Take Away

For teams planning AI infrastructure investments on a 2027 horizon, this first commercial delivery has three concrete implications. First, a rack-scale alternative to NVIDIA now exists beyond the press release. Second, HPE has confirmed its position as an integrator capable of owning the full cycle — from design through production deployment, support included. Third, the presence of a second viable ecosystem in procurement processes restores pricing leverage and delivery timeline flexibility.

The real unknown remains the software ecosystem: compilers, libraries, and framework support for ML workloads running PyTorch or JAX. NVIDIA spent a decade building CUDA's dominance. AMD is advancing with ROCm, but software hegemony is won differently than silicon. For IT leaders, that is where the final decision will be made — and Vultr's real-world results over the coming quarters will be the deciding evidence.

Share

72 GPUs, One Rack: AMD Challenges NVIDIA's AI Infrastructure Dominance