Optical HBM Sharing for GPU Memory Bottlenecks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current GPU architectures face inefficiencies in memory and compute resource utilization due to limited high bandwidth memory (HBM) availability per GPU, leading to increased latency and reduced system efficiency in AI and ML applications.
Innovation Solution
The system and method enable sharing of HBM between multiple GPUs using both electrical and optical switching, facilitated by an optical physical layer that extends signal reach and density, and incorporates all-to-all connections and broadcasting for enhanced data transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If HBM is shared through networking channels between GPUs, then memory capacity is increased, but data transfer speed decreases and latency increases
Solution Approach 1:
The system segments memory resources by providing each GPU with dedicated HBM2 memory while simultaneously enabling shared access to additional HBM2 memory through optical interconnects. This segmentation allows near-memory access for frequently used data while providing shared capacity for less frequently accessed data, resolving the contradiction between memory capacity and data transfer speed.
Solution Approach 2:
Optical interconnects serve as an intermediary between GPUs and shared HBM2 memory resources, enabling high-speed data transfer that bridges the gap between dedicated and shared memory access. The optical medium provides bandwidth comparable to electrical interconnects while enabling extended reach for shared memory architecture.
2Speed
If dedicated HBM is allocated to each GPU, then data access speed is high, but memory resource utilization efficiency decreases
Solution Approach 1:
The HBM2 memory resources are designed to serve multiple functions: acting as dedicated memory for individual GPUs when needed, and as shared memory pool for multiple GPUs simultaneously. The optical interconnect architecture enables the same memory resources to be dynamically allocated to different compute units based on workload requirements, improving overall utilization efficiency while maintaining high access speeds.
3Length of stationary object
If optical interconnects are used for HBM sharing, then reach and density are extended, but system complexity increases
Solution Approach 1:
The system replaces traditional electrical interconnects with optical interconnects for HBM sharing, substituting the mechanical/electrical signal transmission medium with optical technology. This substitution extends signal reach and increases density while managing complexity through standardized optical interfaces and protocols.
4Speed
If electrical signals are used for HBM access, then bandwidth is high, but reach and density are limited
Solution Approach 1:
The system changes the fundamental parameter of signal transmission from electrical to optical domain. Optical signals exhibit different propagation characteristics with lower attenuation and higher bandwidth potential, enabling extended reach and increased density while maintaining high data transfer rates required for HBM access.
Data Source
AI summary
Disclosed is a system and method for sharing memory and enlarging capacity between computer resources using optical links. In certain embodiments, high bandwidth memory is shared between multiple processing units through electrical and optical switching.


