CXL.mem and CXL.cache Translation for Dynamic Memory Pooling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing memory systems in compute architectures face challenges in scalability, efficiency, and interoperability, particularly with protocols like CXL.io, CXL.cache, and CXL.mem, which are not efficiently managed, limiting the performance of applications such as AI workloads and HPC.

Innovation Solution

Introduce Resource Provisioning Units (RPUs) and Memory Fabric Switches to facilitate dynamic memory pooling and protocol translations between CXL.io, CXL.cache, and CXL.mem, enabling seamless communication across different fabric types, including UALink and NVLink.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple CXL protocols (CXL.io, CXL.cache, CXL.mem) are used to support different memory access patterns, then protocol versatility and adaptability are improved, but device complexity and difficulty of protocol management increase

Engineering Contradiction:
Improveprotocol versatilityVSAvoidprotocol management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a protocol translator as an intermediary component that mediates between different CXL protocols (CXL.io, CXL.cache, CXL.mem). This translator automatically converts transactions between protocols, eliminating the need for complex manual protocol management while preserving the versatility of multiple protocols for different memory access patterns.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The protocol translator is designed with multi-functionality to handle multiple CXL protocol types simultaneously. It provides universal support for translating between CXL.io, CXL.cache, and CXL.mem protocols, allowing a single component to manage diverse protocol requirements without increasing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Speed

If HBM technology is used to provide high-speed, low-power memory access, then memory bandwidth and speed are improved, but scalability and cost are worsened due to integration limits

Engineering Contradiction:
Improvememory access speedVSAvoidscalability
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The patent segments the memory system into multiple accessible memory spaces (HBM memory space, CXL memory space, CXL.cache memory space) that can be independently managed. This segmentation allows the system to leverage HBM's high-speed performance for critical workloads while maintaining scalability through CXL-based memory expansion for larger datasets, thus resolving the contradiction between speed and scalability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The protocol translator acts as an intermediary that enables seamless data access between HBM memory and CXL memory spaces. It translates memory access requests between different memory hierarchies, allowing the system to scale beyond HBM's integration limits while preserving HBM's high-speed performance for latency-sensitive operations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If memory disaggregation is implemented to decouple memory from compute nodes, then adaptability and resource sharing are improved, but system complexity and protocol translation requirements increase

Engineering Contradiction:
Improveresource sharing capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The protocol translator serves as a mediating component in memory disaggregated architectures, automatically handling protocol translation between different memory spaces (HBM, CXL.mem, CXL.cache). This intermediary approach enables resource sharing across compute nodes while abstracting away the complexity of protocol translations, thus improving adaptability without proportionally increasing system complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The protocol translation mechanism operates autonomously to manage memory access requests in disaggregated systems. It self-manages the complexity of translating between different protocols and memory hierarchies, allowing the system to benefit from memory disaggregation's resource sharing capabilities without requiring manual intervention or complex configuration.

Inventive Principle:
Principle #25Self-service

4Adaptability or versatility

If CXL.io Configuration Request TLPs are processed directly, then compatibility with existing PCIe configurations is maintained, but translation overhead and processing time increase

Engineering Contradiction:
Improveprotocol compatibilityVSAvoidtransaction processing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The protocol translator performs preliminary action by pre-establishing translation rules and mappings for CXL.io Configuration Request TLPs. This allows for efficient, predictable translation without runtime delays, maintaining compatibility with existing PCIe configurations while minimizing processing time through optimized translation logic.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces manual mechanical protocol translation with an automated software-based translation mechanism. This substitution reduces processing time by eliminating the need for complex hardware-level protocol conversion while maintaining full compatibility with CXL.io and PCIe configurations through intelligent software translation.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12360925B2Translating between CXL.mem and CXL.cache read transactions
Publication Date: 2025.07.15 HYATT GAYA OPAL MS
  • US12360925B2 patent drawing
  • US12360925B2 patent drawing
  • US12360925B2 patent drawing

AI summary

Memory has been playing a major role in the performance, scalability and applicability of General Compute systems, and more recently, in realizing Generative Artificial Intelligence (GenAI) and High-Performance Computing (HPC) systems that scale to thousands of GPUs, CPUs and special-purpose Accelerators. Embodiments herein disclose efficient software-defined protocol terminations and protocol translations utilizing Compute Express Link (CXL), including translations between CXL.mem and CXL.cache protocols. Also disclosed are CXL-based systems, Resource Provisioning Units (RPUs), and Memory Fabric Switches enabling dynamic memory pooling and sharing, host-to-host communication utilizing CXL.mem, CXL.cache and CXL.io, intent-based protocol translations, and optionally seamless interactions between CXL, UALink, NVLink, and/or Ethernet protocols, utilizing a broad range of semantics including IO, Cache, and Memory, optimizing memory access and reducing latency. Some embodiments also enable scalability, flexibility and security in high-performance architectures suited for data centers and next-generation computing environments.