Nvidia is expanding its influence beyond GPUs and interconnects with a new memory architecture designed for custom artificial-intelligence processors. Called NVHBM, the technology is intended to increase memory bandwidth, reduce power consumption, and free up valuable space on accelerator chips.
Nvidia says NVHBM can deliver up to 30% more bandwidth than conventional HBM4E while consuming up to 15% less power. The design may also give chip developers significantly more room for compute hardware and workload-specific features.
A memory design aimed at custom accelerators
NVHBM is being introduced as part of Nvidia’s NVLink Fusion platform. The platform is aimed at hyperscalers and other AI companies that want to develop their own accelerators—often referred to as XPUs—while still connecting those chips to Nvidia’s broader networking and rack-scale infrastructure.
Amazon’s Annapurna Labs is the first publicly announced company working with Nvidia on the technology.
Nvidia will not manufacture the memory itself. Instead, the company is defining the architecture and developing the associated controller and interface technology. Production is expected to come from one or more major memory manufacturers, including Micron, SK hynix, or Samsung.
Moving the memory controller
The central difference between NVHBM and standard HBM designs is the location of the memory controller.
In conventional implementations, the controller is placed on the accelerator’s main compute die. This approach requires additional circuitry and communication paths between the processor and the HBM stacks, taking up space that could otherwise be used for computing resources.
NVHBM moves the controller into the base die of the HBM stack. Nvidia is also supplying a custom physical interface, or PHY, designed specifically for this arrangement.
By relocating these components, Nvidia says accelerator designers can reduce interface complexity on the XPU and dedicate more of the main die to processing, cache, or specialized AI functions.
Nvidia’s performance claims
Nvidia says the architecture could provide the following improvements compared with standard HBM4E:
-
Up to 30% more memory bandwidth.
-
Up to 15% lower HBM power consumption.
-
Up to 25% more usable space on the XPU compute die.
-
Up to 67% less area devoted to the PHY and related interface circuitry.
-
Up to 80% more usable silicon across the overall design.
The company estimates that these benefits could result in approximately 30% higher end-to-end XPU performance when bandwidth, compute-die area, and power savings are considered together.
These figures are Nvidia’s own projections. Actual results will depend on the accelerator architecture, memory configuration, manufacturing process, and the workload being run.
Why the extra die space matters
Higher memory bandwidth is important for AI systems because accelerators often spend significant time moving data between memory and compute units. However, the additional silicon area may prove just as valuable.
Designers could use the freed-up space to add more processing engines, expand on-chip cache, integrate specialized AI functions, or optimize the chip for a particular customer’s applications. All of this could potentially be achieved without increasing the size of the processor package.
HBM already places memory close to the accelerator, reducing the distance data must travel compared with conventional memory modules installed on a motherboard. NVHBM builds on that approach by redesigning how the memory and processor communicate.
Part of Nvidia’s broader strategy
NVHBM is not being presented as a memory product that customers can buy independently. Nvidia is positioning it as one element of NVLink Fusion, a broader effort to support companies developing semi-custom and fully custom AI silicon.
The company says it is establishing a standardized implementation that can be produced by multiple memory suppliers. In theory, this could reduce the engineering and qualification work required when customers source HBM from different manufacturers.
That strategy could help hyperscalers bring custom AI processors to market more quickly while continuing to use Nvidia’s networking and infrastructure technologies.
The outlook for NVHBM
Nvidia’s announcement reflects the growing importance of memory architecture in AI hardware. As models become larger and accelerator designs more specialized, chip performance increasingly depends not only on raw compute capacity but also on bandwidth, power efficiency, and available die area.
NVHBM addresses all three constraints on paper. Whether it becomes a widely adopted technology will depend on manufacturing readiness, partner designs, compatibility with Nvidia’s infrastructure, and independent performance results.
For now, the technology gives Nvidia another way to shape the architecture of the AI data center—even when the final accelerator is designed by a different company.