Chiplet CPU-Accelerator Memory Tunneling Beyond PCIe Limits
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Modern server systems face limitations with Peripheral Component Interconnect Express (PCIe) connections, such as limited shared address space between CPUs and accelerators, high bandwidth and low latency requirements, and restricted port and lane counts that hinder efficient CPU-accelerator balancing.
Innovation Solution
Implementing a system-on-chip (SoC) with a uniform memory access tunneling system that allows accelerators to directly access shared memory via a die-to-die interface, using protocols like AXI or CXL, bypassing the need for accelerator-specific memory.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If PCIe is used to connect accelerators to CPU, then accelerators can be attached to CPU, but shared address space between CPU and accelerators is not allowed
Solution Approach 1:
The patent merges the CPU and accelerator into a unified system-on-chip architecture where both share a common address space through the die-to-die interface. This eliminates the need for separate address spaces and enables direct memory access between CPU and accelerator without requiring complex PCIe mediation layers.
Solution Approach 2:
The die-to-die interface provides universal access for both CPU and accelerator to the same memory resources. The uniform memory access tunneling system allows any device to access any memory location through a standardized interface, making the system more versatile and reducing the need for device-specific connection protocols.
2Speed
If PCIe ports and lanes are used for accelerator connections, then accelerators can communicate with CPU, but aggregate bandwidth is limited and latency increases
Solution Approach 1:
The patent segments the memory access path into separate uniform memory access tunnels for different devices. By dividing the communication pathways and using dedicated tunneling mechanisms, the system achieves higher aggregate bandwidth while maintaining low latency through parallel access paths that bypass traditional PCIe serialization bottlenecks.
Solution Approach 2:
The uniform memory access tunneling system acts as an intermediary layer between CPU and accelerator, providing optimized data transfer pathways. This tunneling mechanism mediates memory access requests to achieve direct high-speed communication without the overhead of traditional PCIe protocol layers.
3Adaptability or versatility
If limited number of PCIe ports and lanes are used, then system complexity is reduced, but number of accelerators that can be attached is limited
Solution Approach 1:
The patent implements dynamic resource allocation where CPU and accelerator share memory resources through the die-to-die interface. The system can dynamically allocate memory regions and adjust access priorities based on workload requirements, enabling more accelerators to be attached without fixed physical port constraints.
Solution Approach 2:
The patent transitions from a physical port-based connection model to a virtual address space-based connection model. By moving the connection limitation from the physical dimension (number of PCIe ports) to the virtual dimension (address space management), the system can support more accelerators through software-defined resource allocation.
4Adaptability or versatility
If PCIe is used for accelerator connections, then accelerators can be integrated, but CPU and accelerator ratio balancing becomes difficult
Solution Approach 1:
The patent enables flexible CPU-to-accelerator ratio configuration by changing the resource allocation parameters in the uniform memory access tunneling system. Administrators can dynamically adjust memory region assignments and access permissions to accommodate different numbers and types of accelerators without physical reconfiguration.
Data Source
AI summary
In some embodiments, a computer-implemented method includes receiving, at a security agent of a host central processing unit (CPU), accelerator firmware from flash memory; determining, at the security agent, whether the accelerator firmware includes a critical accelerator firmware component or a non-critical accelerator firmware component; authenticating, at the security agent, the critical accelerator firmware component instantaneously upon a determination that the accelerator firmware is the critical accelerator firmware component, wherein authenticating the critical accelerator firmware component yields an authenticated critical accelerator firmware component; and providing the authenticated critical accelerator firmware component to an accelerator via a sideband bus for execution at the accelerator.


