Nvidia Weighs Radical Idea: Less Rubin Ultra Chip Memory
Resumo
Nvidia testa três versões do GPU Rubin Ultra com menos memória de alta largura de banda do que o planejado, devido à escassez de chips de memória avançada que pode afetar o desempenho ou exigir mais chips para executar grandes modelos de IA.

Nvidia is weighing a radical step to deal with a shortage of advanced high-bandwidth memory chips: using less of it than planned in its next-generation graphics processing unit, the Rubin Ultra.
Over the past few weeks, Nvidia has been testing at least three versions of the Rubin Ultra GPU, some of which have less memory than what the company initially announced, according to three people with direct knowledge of the trial. Nvidia is considering the lower-memory versions partly because it may not be able to secure enough advanced memory to supply its original design, according to the three people.
Having less memory could affect the chip’s performance, although Nvidia could compensate in other ways. But an AI firm using the reduced-memory chip to run a large AI model may have to use more chips than it would have otherwise.
Nvidia’s possible move reflects how the AI boom has taxed the chip industry’s ability to produce enough supply to meet soaring demand from data centers. That has created an intense shortage of memory chips, a component of GPUs, which has driven up prices. Higher costs have created ripple effects throughout the tech industry, forcing companies to spend more than planned and also causing companies like Apple to raise retail prices for hardware.
Nvidia didn’t have a comment.
What is extraordinary about this situation is that the shortage of memory chips is in some ways due to the popularity in recent years of Nvidia’s GPUs, which use memory. And now Nvidia is itself having to look at alternative memory configurations. The current three Rubin Ultra testing samples would represent something of a downgrade in memory from what is in Nvidia’s Rubin chip, which is in mass production now and rolling out to customers.
The impact of a memory reduction may not significantly hurt customer demand for the Rubin Ultra, according to two Nvidia customers. Because a chip with less memory would likely cost less than the original design, one customer said the change could even benefit companies seeking to cut spending by giving them a cheaper option.
Another said that while large amounts of memory are important for running the biggest frontier AI models, their company is not worried about the memory specifications of any single chip. Instead, it’s focusing on maintaining a longer-term partnership with Nvidia over several generations of hardware.
Nvidia may also be able to offset some of the impact of less memory through its broader server design. Faster networking and more efficient ways to store data could allow customers to spread models and workloads across more Rubin Ultra chips.
The testing of downgraded memory appears to contrast with Nvidia’s public confidence that it had secured sufficient memory supply. In mid-July, Nvidia’s senior vice president of hardware engineering, Andrew Bell, told a group of reporters that Nvidia has a team that attempts to predict and fix supply chain issues years out.
“We were in front of the memory problem, so it’s not gonna hold us back anytime soon,” Bell said. “The pricing, of course, is a problem for the whole world, and probably the pricing will be the bigger challenge. But for supply, we’re in shape.”
There’s Still Time
To be sure, Nvidia hasn’t finalized all the specifications for the Rubin Ultra, according to the three people and one Nvidia customer. Nvidia doesn’t plan to ship the chip until late next year, which means it still has time to adjust the design based on memory availability, costs and customer demand. SemiAnalysis first reported the possible memory chip downgrade in a note to its clients last week.
Nvidia hasn’t decided on the Rubin Ultra’s selling price. High-bandwidth memory can account for more than half of an advanced AI chip’s component cost, according to Epoch AI, and versions with less memory could be significantly cheaper while remaining suitable for many AI applications.
Nvidia CEO Jensen Huang first unveiled the Rubin Ultra at Nvidia’s annual developer conference in 2025. He said each GPU would contain 1 terabyte of HBM4E–the most advanced version of the high-bandwidth memory used in AI chips, designed to move more data, hold more memory and use power more efficiently than the last version, HBM4. That memory would be spread across 16 memory stacks, according to a slide displayed at the conference.
The versions Nvidia has been testing have fewer stacked layers of memory chips and less memory per die, and in some cases they use HBM4, the older version. Some versions tested have as little as 192GB of total memory, while others have 256GB, compared to the initial 1TB, according to the three people and one Nvidia customer. Nvidia says its Vera Rubin chip comes with up to 288GB of HBM4.
By reducing the amount of memory in each GPU, Nvidia could spread what memory it does have across a greater number of GPUs, maintaining its production volume more easily, according to the people.
Manufacturing Challenges
One of the reasons Nvidia is weighing a memory downgrade is that memory suppliers may struggle to produce enough HBM4E to meet Rubin Ultra’s planned schedule, the people said. The greater capabilities of HBM4E require denser memory chips and faster electrical connections. Packaging the HBM4E when the components of a chip are assembled is also harder.
Nvidia has taken steps to address possible memory supply constraints, including striking a $500 billion partnership with the parent company of memory-making giant SK Hynix late last month. That partnership will aim to co-develop new memory technologies including HBM4 and HBM4E in multiple configurations.
That partnership was designed to “help us secure a stable supply of high-bandwidth memory,” said Raj Mirpuri, Nvidia’s vice president of global AI clouds and infrastructure, in a July briefing with reporters. He added that SK Hynix would be making an investment to increase its production capacity for the HBM that would be made available to Nvidia. That’s after SK Hynix already said in June it plans to double its memory chip capacity over the next five years.