Aug. 23 at 12:52 PM
$AMD Helios vs.
$NVDA Vera Rubin,the rack-scale AI battle is much closer than many investors realize.
đ¨Lengthy but informative postđ¨
The next phase of AI infrastructure isnât simply about who makes the fastest GPU. The competition is shifting toward complete rack-scale systems: accelerators + CPUs + memory + networking + cooling + software, engineered to behave like one massive AI computer.
Thatâs where AMD Helios and NVIDIA Vera Rubin collide.
đ´ AMD HELIOS
Helios is AMDâs answer to NVIDIAâs rack-scale architecture, combining Instinct MI400-series accelerators, EPYC CPUs, high-speed networking, open standards, and liquid cooling into an integrated system.
AMD is no longer trying to sell customers an isolated GPU. It wants to sell an entire AI rack architecture.
Helios is designed around 72 GPUs per rack, putting it directly into the same rack-scale conversation as NVIDIA.
AMD is also leaning heavily into open infrastructure.
Rather than forcing customers into a proprietary stack, Helios is designed around technologies such as UALink, Ethernet-based networking, ROCm and broader open ecosystem standards.
That could become one of AMDâs biggest competitive weapons.
Hyperscalers donât necessarily want one vendor controlling the accelerator, CPU, networking, interconnect AND software layers forever. Helios gives them another path.
đ˘ NVIDIA VERA RUBIN
Vera Rubin represents NVIDIA pushing the opposite strategy to its logical extreme.
Rubin combines NVIDIAâs next-generation Rubin GPUs, Vera CPUs, NVLink rack-scale interconnect, networking and the enormous CUDA software ecosystem.
The rack effectively becomes the computer.
And NVIDIAâs biggest advantage remains brutally simple:
CUDA.
NVIDIA has spent years building a software moat encompassing CUDA, libraries, optimized kernels, networking, inference software and developer tooling.
That means Rubin isnât merely competing against MI400.
AMD is competing against NVIDIAâs entire installed ecosystem.
THE ARCHITECTURAL DIFFERENCE
Helios:
72 AMD Instinct GPUs
AMD EPYC CPUs
ROCm
UALink/open interconnect strategy
High-speed Ethernet networking
Open rack architecture
Liquid cooling
Designed for massive training + inference deployments
Vera Rubin:
72 Rubin GPUs in the NVL72 configuration
Vera CPUs
NVLink
CUDA
Spectrum-X / NVIDIA networking ecosystem
Integrated rack architecture
Liquid cooling
Designed for enormous AI factories
Both companies are essentially saying:
Stop thinking about GPUs. Start thinking about AI supercomputers measured in racks and megawatts.
MEMORY IS BECOMING A HUGE BATTLEGROUND
AI models are exploding in size.
That makes HBM capacity and bandwidth increasingly important.
More memory per accelerator means larger models can remain resident in high-bandwidth memory, potentially reducing communication overhead and improving inference efficiency.
AMD has been particularly aggressive about pushing memory capacity across its Instinct roadmap.
That could make Helios especially interesting for large-model inference, where memory economics can matter almost as much as raw compute.
NVIDIA counters with its enormous advantage in NVLink and highly optimized scale-up communication.
AMD can attack with memory + openness + economics.
NVIDIA attacks with interconnect + software + ecosystem integration.
NETWORKING COULD DECIDE MORE THAN PEOPLE THINK
Once youâre connecting 72 GPUs inside a rackâand potentially thousands of racks inside an AI clusterânetworking becomes critical.
NVIDIA owns a tremendous amount of its stack.
GPU â CPU â NVLink â NIC â switches â software.
That vertical integration allows NVIDIA to optimize the system almost end-to-end.
AMDâs approach is different.
Helios represents a more open AI infrastructure model, allowing hyperscalers and OEMs greater flexibility in how systems are assembled and networked.
That could be attractive to companies that donât want their entire AI infrastructure controlled by one supplier.
SOFTWARE: NVIDIA STILL HAS THE ADVANTAGE
CUDA remains the industryâs dominant GPU-computing ecosystem.
Millions of developers already know it.
Thousands of applications are optimized around it.
AMDâs ROCm has improved substantially, but overcoming an ecosystem advantage built over more than a decade doesnât happen overnight.
If two systems deliver comparable hardware performance, NVIDIA can still win because customers value deployment speed, compatibility and software maturity.
AMD doesnât necessarily have to destroy CUDA.
It simply needs ROCm to become good enough that economics begin influencing the purchasing decision.
AND THIS IS WHERE THE ECONOMICS GET INTERESTING
Hyperscalers arenât buying 8 GPUs anymore.
Theyâre contemplating AI factories consuming hundreds of megawattsâor eventually gigawattsâof power.
At that scale, tiny differences become enormous.
GPU price matters.
Performance per watt matters.
Memory capacity matters.
Networking costs matter.
Cooling matters.
Utilization matters.
And ultimately the metric customers care about becomes something like:
tokens per dollar per watt.
If AMD can deliver competitive performance while providing lower acquisition cost or better economics, Helios doesnât need to beat Rubin at everything.
Even taking 10â20% of massive future rack-scale deployments could represent an enormous business.
NVIDIAâS BIGGEST ADVANTAGE
NVIDIA arguably has the strongest AI infrastructure ecosystem ever assembled.
Its advantage isnât merely Rubin.
Itâs:
Rubin + Vera + NVLink + networking + CUDA + libraries + installed base + developer ecosystem.
That combination is extremely difficult to attack.
AMDâS BIGGEST ADVANTAGE
AMD doesnât need to become NVIDIA.
Its opportunity is becoming the credible second ecosystem for hyperscale AI infrastructure.
AMD can offer:
Instinct + EPYC + ROCm + open networking + open rack standards + competitive memory + potentially aggressive economics.
Hyperscalers LOVE second sources.
Competition gives them negotiating leverage and reduces supply-chain concentration.
When companies are spending tens of billions of dollars annually on AI infrastructure, avoiding complete dependence on one vendor becomes strategically valuable.
THEN THEREâS THE SERVER INFRASTRUCTURE LAYER
This is why companies like
$SMCI,
$DELL and
$HPE matter.
Someone still has to turn these chips into deployable AI infrastructure.
That means engineering:
liquid cooling, power delivery, rack integration, networking, storage, serviceability, manufacturing and rapid deployment.
For Supermicro, the ideal outcome isnât necessarily AMD beating NVIDIA or NVIDIA beating AMD.
Itâs BOTH ecosystems exploding.
More Rubin racks.
More Helios racks.
More liquid cooling.
More networking.
More power infrastructure.
More AI factories.
MY SCORECARD
Training leadership: đ˘ NVIDIA Vera Rubin
NVIDIAâs software ecosystem, networking and platform integration remain formidable advantages.
Inference opportunity: đ´ AMD Helios
Memory capacity, accelerator economics and improving ROCm support could make AMD dangerous here.
Software ecosystem: đ˘ NVIDIA
CUDA remains the standard AMD has to chase.
Open ecosystem: đ´ AMD
Helios is positioned around greater infrastructure flexibility and open standards.
Vertical integration: đ˘ NVIDIA
NVIDIA controls an extraordinary percentage of its AI stack.
Potential price/performance disruption: đ´ AMD
AMD doesnât need outright performance leadership if it can deliver compelling TCO and performance/$.
Overall incumbent advantage: đ˘ NVIDIA
Potential market-share disruptor: đ´ AMD
THE BIGGER INVESTMENT THESIS
Investors focusing on âAMD GPU vs NVIDIA GPUâ may increasingly be looking at the wrong battlefield.
The battlefield is becoming:
Helios rack vs Rubin rack.
Then:
Helios cluster vs Rubin cluster.
Eventually:
AMD-powered AI factory vs NVIDIA-powered AI factory.
And hereâs what Iâm most bullish about for the broader AI infrastructure sector:
There doesnât have to be one winner.
AI compute demand could become so enormous that NVIDIA can remain dominant while AMD simultaneously gains billions of dollars of accelerator and rack-scale infrastructure business.
Every generation gets hotter, denser and more complicated.
That means more GPUs.
More HBM.
More networking.
More liquid cooling.
More power.
More racks.
More data centers.
Thatâs bullish for an entire ecosystem:
AVGO MU SNDK TSM WDC
Vera Rubin currently looks like the platform to beat.
But Helios matters because AMD is no longer showing up with just another GPU.
Theyâre showing up with an entire rack-scale AI platform.
If Helios proves competitive on performance per dollar and performance per watt, hyperscalers suddenly have something theyâve wanted for years:
A legitimate alternative to NVIDIA at rack scale.
That competition could define the next several years of the AI infrastructure buildout. #AI #Datacenter