Chiplet CPU-Accelerator Memory Tunneling Beyond PCIe Limits

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern server systems face limitations with Peripheral Component Interconnect Express (PCIe) connections, such as limited shared address space between CPUs and accelerators, high bandwidth and low latency requirements, and restricted port and lane counts that hinder efficient CPU-accelerator balancing.

Innovation Solution

Implementing a system-on-chip (SoC) with a uniform memory access tunneling system that allows accelerators to directly access shared memory via a die-to-die interface, using protocols like AXI or CXL, bypassing the need for accelerator-specific memory.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If PCIe is used to connect accelerators to CPU, then accelerators can be attached to CPU, but shared address space between CPU and accelerators is not allowed

Engineering Contradiction:
Improveshared address space accessVSAvoidconnection interface complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges the CPU and accelerator into a unified system-on-chip architecture where both share a common address space through the die-to-die interface. This eliminates the need for separate address spaces and enables direct memory access between CPU and accelerator without requiring complex PCIe mediation layers.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The die-to-die interface provides universal access for both CPU and accelerator to the same memory resources. The uniform memory access tunneling system allows any device to access any memory location through a standardized interface, making the system more versatile and reducing the need for device-specific connection protocols.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Speed

If PCIe ports and lanes are used for accelerator connections, then accelerators can communicate with CPU, but aggregate bandwidth is limited and latency increases

Engineering Contradiction:
Improvedata transfer bandwidthVSAvoidaccess latency
Core Design Contradiction:
SpeedVSLoss of time

Solution Approach 1:

The patent segments the memory access path into separate uniform memory access tunnels for different devices. By dividing the communication pathways and using dedicated tunneling mechanisms, the system achieves higher aggregate bandwidth while maintaining low latency through parallel access paths that bypass traditional PCIe serialization bottlenecks.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The uniform memory access tunneling system acts as an intermediary layer between CPU and accelerator, providing optimized data transfer pathways. This tunneling mechanism mediates memory access requests to achieve direct high-speed communication without the overhead of traditional PCIe protocol layers.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If limited number of PCIe ports and lanes are used, then system complexity is reduced, but number of accelerators that can be attached is limited

Engineering Contradiction:
Improvenumber of accelerators per CPUVSAvoidinterface configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic resource allocation where CPU and accelerator share memory resources through the die-to-die interface. The system can dynamically allocate memory regions and adjust access priorities based on workload requirements, enabling more accelerators to be attached without fixed physical port constraints.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent transitions from a physical port-based connection model to a virtual address space-based connection model. By moving the connection limitation from the physical dimension (number of PCIe ports) to the virtual dimension (address space management), the system can support more accelerators through software-defined resource allocation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Adaptability or versatility

If PCIe is used for accelerator connections, then accelerators can be integrated, but CPU and accelerator ratio balancing becomes difficult

Engineering Contradiction:
ImproveCPU-to-accelerator ratio flexibilityVSAvoidsystem configuration ease
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent enables flexible CPU-to-accelerator ratio configuration by changing the resource allocation parameters in the uniform memory access tunneling system. Administrators can dynamically adjust memory region assignments and access permissions to accommodate different numbers and types of accelerators without physical reconfiguration.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12481604B2Integrated chiplet-based central processing units with accelerators for system security
Publication Date: 2025.11.25 META PLATFORMS INC
  • US12481604B2 patent drawing
  • US12481604B2 patent drawing
  • US12481604B2 patent drawing

AI summary

In some embodiments, a computer-implemented method includes receiving, at a security agent of a host central processing unit (CPU), accelerator firmware from flash memory; determining, at the security agent, whether the accelerator firmware includes a critical accelerator firmware component or a non-critical accelerator firmware component; authenticating, at the security agent, the critical accelerator firmware component instantaneously upon a determination that the accelerator firmware is the critical accelerator firmware component, wherein authenticating the critical accelerator firmware component yields an authenticated critical accelerator firmware component; and providing the authenticated critical accelerator firmware component to an accelerator via a sideband bus for execution at the accelerator.