NTB Protocol Translation for CXL Memory Borrowing Across Hosts
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing memory systems in compute architectures face challenges in scalability, integration, and inter-node data sharing, particularly for applications like AI and HPC, with legacy IO-fabric solutions like PCIe falling short, and advanced interconnects like CXL and UALink lacking efficient protocol management and translation.
Innovation Solution
Implementing a system-level architectural solution that leverages memory fabric interconnects for scalable memory provisioning and sharing, utilizing RPUs and memory fabric switches to facilitate dynamic memory pooling and intent-based protocol translations between CXL and non-CXL protocols, enabling seamless communication and resource management across compute elements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If legacy IO-fabric solutions like PCIe are used, then backward compatibility is maintained, but scalability and performance for AI/HPC workloads is insufficient
Solution Approach 1:
The patent introduces CXL translators as intermediary components that bridge legacy PCIe protocols with modern CXL protocols. These translators enable heterogeneous systems to coexist by converting between different communication protocols, allowing scalable memory expansion while maintaining compatibility with existing PCIe infrastructure.
Solution Approach 2:
The system implements universal protocol translation capabilities that allow a single memory fabric to support multiple protocols (PCIe, CXL.io, CXL.cache, CXL.mem) simultaneously. This multi-functional approach enables the same infrastructure to serve both legacy and next-generation workloads, achieving scalability without sacrificing backward compatibility.
2Productivity
If advanced interconnects like CXL are deployed, then scalability and data throughput are improved, but protocol management complexity increases
Solution Approach 1:
CXL translators serve as protocol intermediaries that automatically handle translation between different CXL protocol versions and types. This abstraction layer simplifies protocol management by centralizing translation logic in dedicated components, allowing the memory fabric to maintain high throughput while reducing the complexity burden on individual devices.
Solution Approach 2:
The system implements intent-based protocol selection where the memory fabric receives high-level access intentions from compute elements and automatically determines the appropriate protocol translation path. This feedback-driven approach simplifies protocol management by using semantic intent rather than requiring explicit protocol configuration at each layer.
3Adaptability or versatility
If memory disaggregation is implemented, then adaptability and resource sharing are improved, but system complexity increases
Solution Approach 1:
The CXL translator infrastructure provides universal access to memory resources across different compute elements (CPUs, GPUs, FPGAs) through a unified protocol translation layer. This universal interface simplifies resource sharing by abstracting the underlying physical connections, allowing memory disaggregation without proportionally increasing system complexity.
Solution Approach 2:
CXL translators act as intermediary components that mediate between compute elements and memory resources, handling protocol translations and access coordination. This intermediary layer simplifies the overall system architecture by centralizing complexity in translation components while presenting simple, standardized interfaces to both compute elements and memory resources.
4Adaptability or versatility
If protocol translation is added to existing memory systems, then interoperability is improved, but processing overhead increases
Solution Approach 1:
CXL translators are positioned as dedicated intermediary components that handle protocol translation in a centralized manner. By isolating translation logic in specialized hardware components rather than embedding it throughout the entire memory subsystem, the system achieves improved interoperability while minimizing translation overhead through optimized hardware acceleration.
Solution Approach 2:
The patent replaces software-based protocol translation with hardware-accelerated translation mechanisms within CXL translators. This substitution of mechanical (software) processing with hardware-based processing significantly reduces translation latency while maintaining comprehensive protocol interoperability across different CXL protocol types.
Data Source
AI summary
Protocol translation embodiments enabling Memory as a Service (MaaS) and symmetric memory access across heterogeneous computing infrastructures comprising CPUs, GPUs, accelerators, fabric attached memory, and interconnects such as CXL, UALink, NVLink, and Ethernet. The embodiments facilitate NTB or Host-to-Host memory borrowing and symmetric memory flows by translating between CXL and non-CXL protocols, enabling seamless memory sharing between CXL-native and legacy systems, such as PCIe. The embodiments receive non-CXL.cache requests from a first entity, translate them to CXL.cache Device-to-Host Requests for transmission to a second entity, and conversely translate CXL.mem requests to non-CXL.mem requests. This protocol translation enables memory pooling across diverse data center architectures where CPUs, GPUs, and accelerators may share memory resources through Host-to-Host memory borrowing. The embodiments support composable infrastructure for AI/ML workloads, cloud-native applications. The embodiments provide transparent memory disaggregation for multi-tenant environments, enabling dynamic memory provisioning with symmetric memory flows between heterogeneous compute nodes.


