Memory provisioning for ai-scale fabrics based on ualink, nvlink, ESUN / sue, CXL, pcie, chi, UPI, and / or infinity fabric
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- UNIFABRIX LTD
- Filing Date
- 2026-02-03
- Publication Date
- 2026-08-06
Smart Images

Figure IB2026051019_06082026_PF_FP_ABST
Abstract
Description
Memory Provisioning for Al-Scale Fabrics Based on UALink, NVLink, ESUN / SUE, CXL, PCIe, CHI, UPI,and / or Infinity FabricCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This Application claims priority to: U.S. Provisional Patent Application No. 63 / 752,940, filed 3 Feb 2025, U.S. Provisional Patent Application No. 63 / 784,089, filed 5 Apr 2025, U.S. Provisional Patent Application No.63 / 811,859, filed 25 May 2025, U.S. Provisional Patent Application No. 63 / 826,342, filed 18 Jun 2025, U.S. Provisional Patent Application No. 63 / 856,653, filed 3 Aug 2025, U.S. Provisional Patent Application No. 63 / 874,393, filed 2 Sep 2025, U.S. Provisional Patent Application No. 63 / 895,053, filed 7 Oct 2025, U.S. Provisional Patent Application No. 63 / 906,709, filed 28 Oct 2025, and U.S. Provisional Patent Application No. 63 / 931,124, filed 4 Dec 2025.BACKGROUND
[0002] Modem datacenters face unprecedented challenges in memory resource utilization and sharing as workloads become increasingly memory -intensive and distributed. Applications spanning artificial intelligence (Al), machine learning (ML), Large Language Model (LLM) inference, database analytics, and virtualized environments require flexible access to large memory pools that may exceed the capacity limitations of individual servers, whether the servers are CPU-based, GPU-based, or accelerator-based. These evolving demands have driven the development of memory disaggregation technologies that decouple memory resources from compute nodes, enabling more efficient utilization of datacenter infrastructure.
[0003] Compute Express Link (CXL) has emerged as a promising interconnect technology for memory expansion and pooling, providing protocols such as CXL.io, CXL.mem, and CXL.cache that enable high-bandwidth, low-latency communication between processors and memory devices. CXL allows hosts to access memory resources beyond their local physical limitations through standardized interfaces and protocols. However, current CXL implementations face challenges when hosts need to share memory resources, such as in scenarios requiring physical address space isolation and translation between different physical address spaces.
[0004] Traditional memory architectures bind memory resources tightly to specific processors, creating inefficiencies when workloads have varying memory requirements. While CXL enables memory expansion utilizing device attachment, existing solutions typically require each host to manage its own view of memory resources without efficient mechanisms for sharing memory pools among hosts. This limitation becomes apparent in multi-tenant environments, containerized applications, and distributed computing scenarios wherein different hosts may benefit from accessing shared memory resources.
[0005] Moreover, address translations in current systems primarily focus on virtual -to -physical translations within a single host domain through Memory Management Units (MMUs). When hosts attempt to access shared memory resources, the lack of host-to-host physical address translation capabilities creates barriers to memory sharing. Hosts operate within their own Host Physical Address (HP A) spaces, and coordinating access to shared resources across these disparate physical address spaces remains challenging.SUMMARY
[0006] Some of the disclosed embodiments introduce novel system-level architectural solutions leveraging RPUs to enable dynamic memory sharing and pooling across multiple hosts in datacenter environments. These embodiments provide host-to-host physical address translation capabilities that allow different hosts to access shared memory resources through protocols based on CXL, overcoming traditional boundaries between isolated Host Physical Address (HP A) spaces. By implementing RPUs with CXL devices coupled to processor coherent interconnects, the embodiments enable memory sharing between hosts while maintaining address space isolation and security. Some embodiments optionally support Multi-Headed Device (MHD) configurations, enabling multiple hosts to simultaneously access the same memory resources through separate CXL Endpoints. The embodiments address challenges in memory disaggregation and resource utilization for contemporary and future workloads including AI / ML training and inference, LLM deployment, distributed databases, containerized applications, edge computing, and emerging computational paradigms. The host-to-host physical address translations enable efficient memory sharing in multi-tenant environments, cloud-native architectures, and heterogeneous computing systems where different hosts (whether CPU-based, GPU-based, or accelerator-based) require flexible access to shared memory pools. By decoupling memory resources from individual host boundaries while maintaining compatibility with existing operating systems and MMU -based virtual memory systems, the embodiments provide scalable solutions for memory -intensive applications. The integration of RPUs within the processor's coherent interconnect fabric enables low-latency memory access across host boundaries, optionally supporting real-time analytics, in-memory computing, and distributed shared memory models.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The embodiments are herein described by way of example only, with reference to the accompanying drawings. No attempt is made to show structural details of the embodiments in more detail than is necessary for a fundamental understanding of the embodiments. In the drawings:
[0008] FIG. 1 A illustrates one embodiment of a system comprising a memory switch, a memory pool, or a Global Fabric -Attached Memory Device;
[0009] FIG. IB illustrates one embodiment of a system comprising a memory pool coupled to hosts and to a memory expander;
[0010] FIG. 2A illustrates one embodiment of a system comprising a memory pool comprising two or more MxPUs;
[0011] FIG. 2B illustrates one embodiment of a system comprising a memory pool comprising at least one MxPU and at least one xPU or CPU;
[0012] FIG. 3A illustrates one embodiment of a system comprising a memory pool comprising a processor, DRAM, and an RPU performing host-to-host physical address translations;
[0013] FIG. 3B illustrates one embodiment of a system comprising a memory pool comprising a CXL Multi Headed Device (MHD) comprising a processor coupled to DRAM;
[0014] FIG. 4 illustrates one embodiment of a system comprising an Al memory switch or a memory pool, comprising a CXL Multi Headed Device (MHD);
[0015] FIG. 5 A illustrates one embodiment of a system comprising an RPU that translates between UALink and a Coherent Interconnect Interface;
[0016] FIG. 5B illustrates one embodiment of a TFD showing address translation between UALink UPLI and ARM CHI ReadOnce;
[0017] FIG. 6A illustrates one embodiment of a system functioning as a UALink memory switch appliance or a UALink memory pool;
[0018] FIG. 6B illustrates one embodiment of a TFD depicting a multi-entity memory access scenario whereinGPUs access memory through UPLI-to-ARM CHI translations;
[0019] FIG. 7A illustrates one embodiment of a system that may function as an NVLink memory switch appliance or an NVLink memory pool;
[0020] FIG. 7B illustrates one embodiment of a TFD depicting a multi-entity memory access scenario wherein GPUs access memory mapped to physical address spaces through NVLink to ARM CHI translations;
[0021] FIG. 8 illustrates one embodiment of a system comprising a processor comprising an NVLink interface, processing cores, LLC, a CXL RP, and memory controllers coupled via memory channels to memory;
[0022] FIG. 9 illustrates one embodiment of a TFD demonstrating translations from NVLink traffic to traffic conforming to a protocol utilized by a processor’s coherent interconnect;
[0023] FIG. 10A illustrates one embodiment of a system wherein an entity is coupled via an IEEE 802.3 PHY to an RPU comprising a CXL device coupled to an ARM architecture processor;
[0024] FIG. 10B illustrates one embodiment of a TFD demonstrating translating CXL.mem messages to ARM CHI requests;
[0025] FIG. 11 A illustrates one embodiment of a system comprising a processor having multiple interfaces;
[0026] FIG. 1 IB illustrates one embodiment of a system comprising a processor capable of servicing external requests through CCGs optimized for handling CXL.mem traffic;
[0027] FIG. 12 illustrates one embodiment of a multi-host memory pooling or sharing utilizing a switch-based topology with physical layers based on IEEE 802.3 PMA;
[0028] FIG. 13A illustrates one embodiment of a system that translates between NVLink-based traffic and CHIbased coherent interconnect traffic;
[0029] FIG. 13B illustrates one embodiment of a TFD showing the translation of NVLink read request to CHI ReadOnce request;
[0030] FIG. 14A illustrates one embodiment of a system that translates between NVLink-based traffic and ARM CHI traffic;
[0031] FIG. 14B illustrates one embodiment of an RPU that translates between NVLink traffic and CHI traffic, utilizing an intermediate protocol based on ARM AMBA ACE -Lite;
[0032] FIG. 15 A illustrates one embodiment of a system that translates between NVLink traffic and CHI -based traffic;
[0033] FIG. 15B illustrates one embodiment of an RPU that translates between NVLink traffic and CHI traffic;
[0034] FIG. 16A illustrates one embodiment of a TFD showing translating an NVLink read request to a PCIe UIO read request to an ARM CHI ReadOnce request;
[0035] FIG. 16B illustrates one embodiment of a TFD showing translating an NVLink read request to a CXL.cache RdCurr request to an ARM CHI ReadOnce request;
[0036] FIG. 17A illustrates one embodiment of a system comprising an external entity coupled to an optional NVLink switch coupled to a processor comprising an RPU comprising an NVLink interface, a Request Agent (RA) Proxy, and a Home Agent (HA) Proxy;
[0037] FIG. 17B illustrates one embodiment of a system comprising a processor comprising NVLink chiplets (such as NVLink Fusion) to translate between NVLink and CHI;
[0038] FIG. 18A illustrates one embodiment of a system comprising an xPU comprising an RPU that translates between NVLink traffic and CHI traffic;
[0039] FIG. 18B illustrates one embodiment of a system comprising an entity including NVLink and CXL ports coupled to CHI interfaces that enable memory access via a processor’s coherent interconnect;
[0040] FIG. 19A illustrates one embodiment of a system comprising a processor comprising an NVLink chiplet coupled via NVLink-C2C to the processor’s coherent interconnect;
[0041] FIG. 19B illustrates one embodiment of a system comprising an xPU coupled to a GPU utilizing an RPU that translates between NVLink traffic and CHI -based traffic;
[0042] FIG. 20A illustrates one embodiment of GPU / CPU coupled to an xPU comprising dies coupled by chip-to-chip interfaces;
[0043] FIG. 20B illustrates one embodiment of a custom accelerator comprising an NVLink Fusion chiplet;
[0044] FIG. 21 A illustrates one embodiment of a system functioning as a multi-protocol memory switch appliance or a multi-protocol memory pool utilizing NVLink-based interfaces;
[0045] FIG. 2 IB illustrates one embodiment of a TFD depicting a multi -entity memory access scenario wherein separate NVLink and UALink transactions utilize the same coherent interconnect infrastructure for memory access;
[0046] FIG. 22 illustrates one embodiment of a heterogeneous computing system comprising an NVLink chiplet coupled to an accelerator based on ARM mesh architecture;
[0047] FIG. 23A illustrates one embodiment of a memory switch configured to provide memory to its coupled entities;
[0048] FIG. 23B illustrates one embodiment of a TFD demonstrating NVLink requests from entities to access memory;
[0049] FIG. 24A illustrates one embodiment of a system that translates between CXL.mem and CXL.cache;
[0050] FIG. 24B illustrates one embodiment of a TFD demonstrating translations between CXL.mem M2S MemRd Request and CXL.cache D2H RdCurr Request;
[0051] FIG. 24C illustrates one embodiment of a TFD demonstrating translations between CXL.mem M2S MemRd Request and CXL.cache D2H RdShared Request;
[0052] FIG. 25A illustrates one embodiment of a system that translates between first and second CXL.cache;
[0053] FIG. 25B illustrates one embodiment of a transaction flow diagram (TFD) demonstrating translations between CXL.cache H2D Snplnv Request and CXL.cache D2H CLFlush Request;
[0054] FIG. 25C illustrates one embodiment of a TFD demonstrating translations between CXL.cache H2D SnpCur Request and CXL.cache D2H RdCurr Request;
[0055] FIG. 26A illustrates one embodiment of a system that translates between first and second CXL.mem;
[0056] FIG. 26B illustrates one embodiment of a transaction flow diagram (TFD) demonstrating translations between CXL.mem M2S MemRdData Request and CXL.mem M2S MemRd Request, with optional speculative memory reads;
[0057] FIG. 27A illustrates one embodiment of a system that translates between a UALink-based protocol and a PCIe-based protocol;
[0058] FIG. 27B illustrates one embodiment of a TFD demonstrating translations between UPLI Request ReqCmd(Read) and PCIe MRd; and
[0059] FIG. 27C illustrates one embodiment of a TFD demonstrating translations between UPLI Request ReqCmd(Read) and PCIe UIOMRd.DETAILED DESCRIPTION
[0060] In one embodiment, a system comprising: a processor comprising a coherent interconnect; the processor is coupled to memory having a capacity of at least 64GB; wherein the processor is configured to utilize physical addresses within a Host Physical Address (HP A) space to access the memory, and to execute an operating system (OS) that utilizes a virtual address space; a memory management unit (MMU) configured to enable access to the memory based on mapping addresses within the virtual address space to physical addresses within the HPA space; a resource provisioning unit (RPU) comprising a Compute Express Link (CXL) device configured to communicate with an entity according to a protocol based on CXL; and wherein the RPU is further coupled to the coherent interconnect and configured to perform host-to-host physical address translations, whereby the host-to-host physical address translations enable the entity to access the memory via the CXL device.
[0061] Optionally, the OS may utilize the MMU for virtual to physical address mapping to access the memory, wherein the MMU translates OS-level virtual addresses to physical addresses within the HPA space. Processes, applications and user programs executing under the control of the OS may utilize the MMU to access the memoryutilizing virtual addresses while the MMU enforces memory protection and isolation between different processes or applications. Device drivers operating within the OS kernel space may utilize the MMU for accessing memory -mapped device registers and for managing DMA buffers. When the processor supports virtualization, hypervisors may utilize the MMU to manage memory mappings for virtual machines (VMs), wherein hypervisors and / or guest OSs may further utilize the MMU to manage memory mappings for processes within the VMs, optionally supporting nested virtualization that may include multiple levels of address translations. In some embodiments, an MMU may translate from addresses within a physical address space, such as a Guest Physical Address (GPA) space, to addresses within another physical address space, such as an HPA space. Infrastructure code or firmware running on hidden cores may utilize the MMU for accessing memory regions allocated for infrastructure tasks such as memory telemetry collection or memory pool management operations. And hardware components such as DMA engines within the system may utilize the MMU or IOMMU functionality to perform address translations when moving data between different memory regions.
[0062] Optionally, the processor, MMU, and RPU may be implemented as a semiconductor device that combines processing capabilities with memory pooling functionality. The processor may be a multi -core processor based on x86, ARM, RISC-V, or other instruction set architectures, and may include various levels of cache hierarchy. The HPA space utilized by the processor is the physical address space the processor utilizes to access the memory. The RPU may be implemented as dedicated hardware logic, firmware running on dedicated cores, or a combination thereof, and may maintain translation tables or use programmable mappings to convert between different HPA spaces used by external entities and the local HPA space of the processor.
[0063] Optionally, the messages received by the RPU, such as the messages conforming to the CXL protocol, may include additional messages that do not carry HPA, and such messages may be processed by the RPU without performing host-to-host physical address translations. Additionally or alternatively, the RPU may further process additional messages that carry virtual addresses instead of host physical addresses, and the messages carrying host physical addresses may coexist with other types of messages that may be processed differently by the RPU, such that the description of messages carrying host physical addresses does not limit the presence or processing of other types of messages that may be communicated with the entity and through the processor. Furthermore, the RPU may apply different processing methods to different types of messages according to their content and / or requirements, which may include forwarding messages without modification, modifying message contents without performing address translations, or performing other types of translations or modifications that may differ from the above described host-to-host physical address translations.
[0064] Optionally, the entity utilizes a second HPA space, and the host-to-host physical address translations translate physical addresses within the second HPA space to physical addresses within the HPA space. Optionally, the second HPA space utilized by the entity may have a different size, layout, or addressing scheme compared to the HPA space utilized by the processor. The host-to-host physical address translations may include offset calculations, range remapping, or lookup table operations to convert addresses between the two HPA spaces. The RPU may support configurable translation windows that define which portions of the entity’s HPA space are mapped to the processor's HPA space, and may implement protection logic to prevent unauthorized access to memory regions outside the allocated ranges.
[0065] Optionally, the system further comprises a CXL Root Port configured to communicate with a CXL memory expander that utilizes a Device Physical Address (DPA) space; and wherein at least one of the operating system, system firmware, or the memory expander is configured to map between physical addresses within the HPA space and physical addresses within the DPA space, which enable the entity to utilize the memory and / or the CXL memory expander. Optionally, the CXL memory expander may be a CXL type-3 device that provides additional memory capacity to the system. The DPA space of the memory expander represents the device -local physical addresses used internally by the expander. The OS or system firmware may maintain mapping tables that associate HPA ranges with DPA ranges of the memory expander, enabling transparent access to the expanded memory. Additionally or alternatively, HPA to DPA mapping may further be maintained by the memory expander, such as via internalfirmware, software, or hardware of the expander.
[0066] Optionally, the RPU further comprises a second CXL device configured to communicate with a second entity utilizing a second protocol based on CXL, whereby the second entity utilizes a third HPA space; and wherein the RPU is further configured to translate physical addresses within the third HPA space to physical addresses within the HPA space, which enable the second entity to utilize the CXL memory expander. Optionally, the system may support multiple entities accessing the CXL memory expander utilizing coordinated address translations. Different entities may have their own portions of the memory expander's capacity utilizing separate HDM regions or virtual CXL devices exposed by the RPU. Additionally or alternatively, the memory expander may expose multiple HDM regions, or may expose multiple logical devices (LDs), which may be mapped via RPU translations to multiple entities. The RPU may maintain separate translation contexts for separate entities, ensuring that memory accesses from different entities are properly isolated while still allowing shared access to designated memory regions when configured for multi-entity sharing. The system may implement Quality -of-Service (QoS) mechanisms to fairly allocate memory expander bandwidth among multiple entities.
[0067] Optionally, the RPU further comprises a second CXL device configured to communicate with a second entity utilizing a second protocol based on CXL, whereby the second entity utilizes a third HPA space, and the RPU is further configured to translate physical addresses within the third HPA space to physical addresses within the HPA space, which enable the second entity to utilize the memory. Optionally, when supporting multiple entities accessing the memory (e.g., DRAM), the system may implement memory partitioning schemes to allocate specific memory regions to different entities. The RPU may enforce access controls to enable entities to access only their respective allocated memory regions. The system may support dynamic reallocation of memory between entities based on workload demands or administrative policies, and may implement memory tiering and migration capabilities to move data between different entities' allocated regions such as when workload access patterns change or reconfiguration occurs.
[0068] Optionally, the entity comprises a host coupled to the processor via at least one of a CXL root port or a CXL switch, and the second protocol based on CXL is different from the protocol based on CXL. Optionally, supporting different CXL protocols for different entities may enable heterogeneous system configurations wherein entities with varying capabilities can utilize or share the memory pool. For example, one entity may use CXL.mem for simple memory expansion while another entity uses CXL. cache for cache -coherent shared memory. The RPU may maintain protocol-specific state machines and translation logic for different supported protocol combinations, enabling interoperability between entities using different CXL protocol subsets.
[0069] Optionally, the processor comprises a modified processing unit (MxPU), the memory comprises dynamic random-access memory (DRAM), and the RPU enables the entity to utilize DRAM having a capacity of at least 256GB of the DRAM. Optionally, the MxPU may be derived from an established CPU or GPU design with modifications to support CXL device functionality and host-to-host address translations. The large DRAM capacity (>256GB) may be achieved through multiple memory channels supporting high-capacity DRAM modules. The MxPU may implement memory compression, deduplication, or other techniques to effectively increase the usable memory capacity exposed to entities beyond the physical DRAM capacity.
[0070] Optionally, the memory comprises dynamic random-access memory (DRAM) that is coupled via memory channels to the processor, and the CXL device comprises a Global Fabric -Attached Memory (G-FAM) Device (GFD). Optionally, the memory channels may include multiple channels transmitting in parallel to increase memory bandwidth and reduce latency. The memory channels may support one or more DRAM modules, such as DIMMs or RDIMMs, and may implement various memory technologies including DDR4, DDR5, LPDDR4, LPDDR5, or future memory standards. The memory channels may include memory controllers integrated within the processor or implemented as separate components within the system, and may support features such as ECC, memory interleaving, and channel bonding for improved performance and reliability.
[0071] Optionally, the protocol based on CXL utilizes CXL.mem, and the CXL device exposes at least one Hostmanaged Device Memory (HDM) address region to the entity. Optionally, when operating according to CXL.mem,the CXL device (such as CXL EP) may expose one or more HDM regions that appear as memory -mapped regions to the coupled entity. The HDM regions may be configured with specific address ranges, access permissions, and memory attributes through HDM decoders. The entity may access these HDM regions using standard memory load / store operations, which are translated by the entity's CXL root port into CXL.mem transactions. The system may support multiple HDM regions with different characteristics, such as volatile memory regions backed by the memory and persistent memory regions backed by storage -class memory.
[0072] Optionally, the protocol based on CXL utilized CXL.io, and the host-to-host physical address translation translates from physical addresses carried in CXL.io UIOMRd Transaction Layer Packets (TLPs) received from the entity to physical addresses within the HPA space. Optionally, when operating according to CXL.io, the system may process various types of TLPs including memory read / write TLPs, configuration TLPs, and message TLPs. The UIOMRd TLPs may carry physical addresses within the entity’s physical address space that require translation to the local HPA space. The RPU may intercept these TLPs, extract the physical addresses, perform the necessary translations, and generate corresponding transactions in the local HPA space. The system may also support other CXL.io transaction types such as UIOMWr for memory writes and may implement flow control and credit management according to CXL specifications.
[0073] Optionally, the processor comprises multiple cores, from which at least one is a hidden core; and wherein the RPU is further configured to utilize the hidden core for internal tasks, wherein the internal tasks comprise at least one of internal firmware processing, CXL Fabric Manager (FM) API processing, processing in memory (PIM), nearmemory processing, or housekeeping tasks. Optionally, the RPU is configured to utilize at least one hidden core for internal tasks, which may include processing internal firmware, handling CXL Fabric Manager (FM) API processing, processing in memory (PIM), near-memory processing, and / or performing housekeeping tasks. By utilizing hidden cores to these specific functions, the processor may improve its performance and enable efficient operation without overburdening non-hidden cores that may be allocated to running user workloads. Additionally, utilizing the hidden core(s) for the RPU tasks can allow a CPU vendor to differentiate the processor from other CPUs while maintaining compatibility with existing / established designs, applications, and software code base that was developed for established CPUs.
[0074] Optionally, the hidden core is isolated from user access and visibility, providing user-infrastructure isolation. Optionally, the processor’s hidden core(s) are isolated from user access and visibility, providing userinfrastructure isolation. This isolation ensures that the user cannot affect the execution of code on the hidden cores, enhancing the security and reliability of the system. By separating the visible user-controlled cores from the hidden vendor-controlled cores, the processor can effectively protect critical infrastructure functions from undesired interference or tampering by potentially malicious user code.
[0075] Optionally, the processor comprises multiple cores, from which at least one is hidden and is utilized for collection of memory telemetry. Optionally, at least one of the processor’s hidden core(s) is utilized for collection of memory telemetry. By running memory telemetry on the hidden core(s), the system can effectively monitor and manage memory resources, such as memory resources in a memory pool, without burdening the user-accessible cores, which allows for efficient resource utilization and prevents memory management tasks from interfering with user code execution.
[0076] Optionally, the processor comprises multiple cores, from which at least one is a hidden core utilized for secure key storage and management for encrypting and decrypting data transmitted according to the protocol based on CXL, leveraging user-infrastructure isolation provided by the hidden core. Optionally, at least one of the processor’s hidden core(s) is utilized for secure key storage and management, specifically for encrypting and decrypting data transmitted according to the protocol based on CXL. By leveraging the user-infrastructure isolation provided by the hidden core(s), the system prevents sensitive cryptographic keys used for securing data transmitted according to the protocol based on CXL from being accessible to user code. This isolation enhances the security of the data transmitted between the processor and the entity, protecting it from potential compromise by malicious user code. The hidden core(s) may perform the cryptographic operations on the data themselves, improving confidentiality,integrity, and / or replay protection. Alternatively, the hidden core(s) may utilize hardware -accelerated cryptographic engine(s) for performing at least part of the cryptographic operations on the data, while the hidden core(s) remain responsible for the management of the secure keys and for controlling the processing flows of the data. In this approach, the cryptographic accelerator may handle the data processing while the hidden core(s) handle the control, following a Control / Data Plane separation. Furthermore, the infrastructure code running on the hidden core(s) may participate in enabling support for confidential computing over memory exposed / provisioned by the RPU via the CXL device of the system.
[0077] Optionally, the system further comprises a hardware -accelerated cryptographic engine, wherein the hidden core is configured to utilize the hardware -accelerated cryptographic engine for performing at least part of the cryptographic operations on the data transmitted according to the protocol based on CXL. Optionally, the system includes one or more hardware -accelerated cryptographic engines that can be utilized by the hidden core(s) for performing at least part of the cryptographic operations on the data transmitted according to the protocol based on CXL. The hidden core(s) are responsible for managing the secure keys and controlling the processing flows of the data, while the cryptographic engine(s) handle the actual data processing. This approach features control / data plane separation, wherein the hidden core(s) act as the control plane, and the cryptographic engines serve as the data plane. By offloading the computationally intensive cryptographic operations to hardware accelerators, the system may achieve higher performance and efficiency in securing the data transmitted according to the protocol based on CXL.
[0078] Optionally, the hidden core enables support for confidential computing over memory exposed by the RPU via the CXL device; whereby confidential computing performs computation within a secure isolated environment to protect data in use. Optionally, the hidden core(s) of the processor enable support for confidential computing over memory exposed / provisioned by the RPU via the CXL device. Confidential computing is a security paradigm that aims to protect data in use by performing computation within a secure, isolated environment, such as a Trusted Execution Environment (TEE). In Confidential computing, data remains encrypted and confidential even during processing, protecting sensitive information from unauthorized access, modification, or disclosure. This may be achieved utilizing a combination of hardware -based security features, such as encrypted memory regions and secure enclaves, and optional software -based logic that enforce access controls and data isolation. By enabling computation on encrypted data without exposing the plaintext contents, confidential computing provides a higher level of security and privacy compared to traditional computing models that only protect data at rest and in transit. The infrastructure code running on the hidden core(s) participates in setting up and managing the secure environment required for confidential computing, including provisioning encrypted memory regions, managing encryption keys, and keeping sensitive data protected from unauthorized access. By leveraging the user-infrastructure isolation provided by the hidden core(s), the system can create a trusted execution environment for confidential computing, enabling secure processing of sensitive data within the memory exposed by the RPU utilizing the protocol based on CXL.
[0079] Optionally, the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core for error handling and / or correction tasks within a memory pool comprising the memory, enhancing data integrity and reliability. Optionally, the error handling and correction tasks performed by hidden cores may include detecting and correcting single -bit and multi-bit errors, managing spare memory regions for replacing faulty memory locations, and maintaining error logs for system analysis. The hidden cores may implement scrubbing routines (e.g., patrol scrub) that periodically read and correct memory contents to prevent error accumulation. The system may support various error correction codes and advanced ECC schemes suitable for large-scale memory pools.
[0080] Optionally, the error handling and / or correction tasks further comprise predictive failure analysis (PF A) operations, configured to predict and handle imminent failure of memory components within the memory pool, thereby preempting potential data loss and system downtime. Optionally, the error handling and correction tasks executed by the hidden core(s) of the processor include predictive failure analysis operations designed to anticipate and address imminent failures of memory components within the memory pool. By implementing the PF A, the system may proactively identify potential faults before they manifest into actual failures, enabling timely interventions thatmitigate the risk of data loss and system downtime. The PFA may not only enhance the reliability and data integrity of the memory system but also improve overall system resilience in high-performance computing architectures.
[0081] Optionally, the memory comprises dynamic random-access memory (DRAM), and the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core for controlling or managing memory access scheduling within a memory pool comprising the DRAM, to improve memory utilization and throughput. Optionally, memory access scheduling controlled or managed by hidden cores, such as via utilizing a hardware -based memory controller or a memory access scheduler managed by hidden cores, may optimize memory bandwidth utilization by reordering memory requests based on factors such as request priority, memory bank availability, and access patterns. The hidden cores may implement and apply sophisticated scheduling algorithms that consider Quality -of-Service (QoS) requirements, minimize memory access conflicts, and maximize row buffer hit rates. The scheduling may also account for thermal constraints and power management goals while maintaining fair access for the memory pool clients.
[0082] Optionally, the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core for managing security protocols within a memory pool comprising the memory, including data encryption and / or access controls. Optionally, security protocol management by hidden cores may include encryption algorithms for data at rest and in transit, managing security keys and certificates, and enforcing access control policies. The hidden cores may support various security standards such as CXL Integrity and Data Encryption (IDE) for protecting data transmitted over CXL links. The memory pool may include secure enclaves or trusted execution environments to protect sensitive data and cryptographic operations from unauthorized access.
[0083] Optionally, the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core for configuration management tasks within a memory pool comprising the memory, including dynamic allocation and deallocation of memory resources. In further embodiments, one or more of the hidden cores of the processor may be utilized for advanced infrastructure management tasks within a memory pool based on the processor and the memory. These tasks may include one or more of: (i) error handling and correction, which enhances data integrity and reliability by promptly addressing memory errors, (ii) memory access scheduling, which improve the allocation and utilization of memory resources based on current demand and operational priorities, (iii) security management, which secures the memory pool by implementing robust encryption and access controls to safeguard data, and / or (iv) configuration management, which dynamically adjusts memory settings to adapt to varying workload requirements. One or more of these tasks may be employed to maintain the overall efficiency, security, and / or performance of the system, such as in environments requiring high-speed, high-integrity memory operations, thereby enhancing the system’s capabilities and distinguishing it from architectures based on conventional CPU / GPU (where CPU / GPU refers to CPU and / or GPU).
[0084] Optionally, the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core for memory tiering tasks. Optionally, memory tiering tasks performed by hidden cores may include classifying memory regions into different performance tiers based on their underlying technology characteristics. The hidden cores may monitor access patterns to different memory regions, such as via utilizing hardware -based telemetry collectors and analyzers, and dynamically adjust tier assignments to optimize overall system performance. The system may support various memory technologies and / or speeds in different tiers, suchas high-bandwidth DRAM (e.g., MRDIMMs) in tier 1, ordinary DRAM (e.g., RDIMMs) in tier 2, and persistent memory or storage-class memory (SCM) in lower tiers.
[0085] Optionally, the memory tiering tasks further comprise migration of data between memory tiers based on hotness level of the data, thereby increasing performance of memory accesses from the entity to hot data. Optionally, the hidden core(s) of the processor may enable support for memory tiering, wherein memory regions or subsets of memory regions exposed to entities, may be mapped to memory resources based on parameters such as the hotness of the data in these memory regions, e.g., the frequency at which the data is used. In one embodiment, the hidden core(s) may utilize memory telemetry to map hot data to higher-performance memory tiers, whereas colder data may bemapped to slower memory such as Flash memory coupled to the processor. In other embodiments, the hidden core(s) may utilize memory mapping based on priority or Service-Level Agreement (SLA) associated with the data, e.g., in cases wherein the system is configured to prioritize particular workloads, virtual machines, users, or tenants, that utilize the data. Yet in other embodiments, the hidden core(s) may migrate data between memory tiers, such as migrating hot data from a lower-performance memory tier to a higher-performance memory tier.
[0086] Optionally, the system further comprises a direct Memory Access (DMA) engine, wherein the hidden core is configured to utilize the DMA engine for migrating data between memory tiers. Optionally, the hidden core(s) of the processor may utilize a DMA engine for data migration between memory tiers, offloading the data movement task from the hidden core(s) to a dedicated engine, thereby providing faster migration of data and freeing the hidden core(s) to perform additional tasks.
[0087] In various embodiments, hidden cores are isolated from the user's access and visibility, while visible cores are available for user utilization. This isolation may be achieved utilizing different techniques, such as utilizing Type 1 hypervisors, Type 2 hypervisors, hardware partitioning, software partitioning, asymmetric multiprocessing (AMP), firmware configuration, CPU microcode updates, custom CPUs, security extensions, and / or a combination thereof.
[0088] In a first example, a Type 1 hypervisor may be utilized to create hidden and visible cores. A Type 1 hypervisor, such as VMware ESXi or Microsoft Hyper- V, runs on the hardware and manages virtual machines (VMs). The hypervisor can allocate specific processing cores to VMs using techniques such as CPU affinity or core pinning. For instance, certain cores may be designated as hidden and assigned to a VM that is not accessible or visible to the user. These hidden cores may run system management tasks or specialized applications such as CXL memory management or memory pool operations, while the visible cores are allocated to user-accessible VMs running general-purpose operating systems (GPOS). The hypervisor prevents the user from direct access to the hidden cores, maintaining isolation.
[0089] In a second example, a Type 2 hypervisor may be utilized to achieve similar isolation. A Type 2 hypervisor, such as VMware Workstation or Oracle VirtualBox, mns on a host OS and supports guest OSes, wherein the host OS manages the visible cores accessible to the user. The Type 2 hypervisor can then create additional VMs using hidden cores, which run separate OSes or specialized tasks. The overhead of the Type 2 hypervisor is higher compared to a Type 1 hypervisor, but it may provide additional flexibility in managing user-visible and hidden cores.
[0090] In a third example, hardware partitioning, also known as hardware -assisted virtualization in some systems, may be utilized to divide processing cores to isolated partitions at the hardware level, wherein the isolated partitions run different operating systems. It may be used in various scenarios wherein isolation between partitions is required, including high-reliability and safety -critical systems. For instance, one partition with hidden cores may run an RTOS or embedded OS for critical system functions, while another partition with visible cores runs a GPOS for user applications. Hardware partitioning enables isolation, as the partitions are managed by the hardware, preventing user access to the hidden cores.
[0091] In a fourth example, software partitioning, such as the Jailhouse hypervisor, may be utilized to create isolated partitions while offering lower overhead compared to full virtualization. This approach allocates specific cores to different partitions, wherein hidden cores may run dedicated tasks or specialized applications. For example, Jailhouse can configure certain cores to run an RTOS or bare -metal applications, isolating them from user access; and visible cores can run a GPOS that is available for user applications.
[0092] In a fifth example, Asymmetric Multiprocessing (AMP) may be utilized to run different OSes on different cores without a hypervisor. In this configuration, certain cores may mn an RTOS or embedded OS, while other cores may run a GPOS. Communication between the operating systems may be achieved utilizing shared memory or interprocess communication logic. For instance, Linux may run on the visible cores for user applications, while an RTOS may run on the hidden cores for real-time tasks. AMP provides a straightforward method to isolate hidden cores from user access while leveraging the specific strengths of different operating systems.
[0093] In a sixth example, firmware configuration may be utilized to achieve hidden and visible cores. By accessing the Basic Input / Output System (BIOS) or the Unified Extensible Firmware Interface (UEFI) settings, certainCPU cores can be disabled, making them invisible to the OS. While this method can prevent the OS from utilizing the disabled cores, it is noted that depending on the embodiment, these cores may still be accessible utilizing other means, such as hardware debugging interfaces, and these changes may not be persistent (e.g., rebooting the system could reset the BIOS / UEFI settings, making the hidden cores visible again). Therefore, depending on the specific requirements, additional measures may be necessary to provide complete isolation of the hidden cores.
[0094] In a seventh example, CPU microcode updates provided by the hardware vendor may be employed. These updates can include specific instructions to disable or hide cores at the microcode level, preventing their detection or usage by the operating system. This method provides a secure way to manage core visibility, as the updates are controlled by the CPU manufacturer.
[0095] In an eighth example, custom CPU designed by hardware vendors can be utilized, which include technologies and mechanisms that enable core partitioning and management of core visibility. For example, Intel's Resource Director Technology (RDT) allows for the partitioning of CPU resources, while ARM's Big.LITTLE architecture enables heterogeneous multi-processing, wherein different types of cores can be used for different purposes. These vendor-specific embodiments provide control over core allocation and maintain certain cores hidden from the user.
[0096] In a ninth example, security extensions such as Intel’s Trusted Execution Technology (TXT) or ARM’s TrustZone may be used. These technologies create secure execution environments that isolate specific cores for security-sensitive operations. The hidden cores may only be accessible within the secure environment, protecting them from user interference and enabling secure execution of critical tasks.
[0097] In one embodiment, a method comprising: accessing memory coupled to a processor utilizing physical addresses within a Host Physical Address (HP A) space; wherein the processor comprises a coherent interconnect; mapping addresses within a virtual address space to physical addresses within the HPA space; whereby the addresses within the virtual address space are utilized by an operating system (OS) of an apparatus comprising the processor; communicating, by a Compute Express Link (CXL) device of a resource provisioning unit (RPU), with an entity coupled to the apparatus according to a protocol based on CXL; wherein the RPU is coupled to the coherent interconnect; and performing, by the RPU, host-to-host physical address translations which enable the entity to access the memory via the CXL device.
[0098] Optionally, the entity comprises a second host that utilizes a second HPA space, and the host-to-host physical address translations are translating physical addresses within the second HPA space to physical addresses within the HPA space. Optionally, the method further comprises communicating, via a CXL Root Port, with a CXL memory expander that utilizes a Device Physical Address (DPA) space; and wherein at least one of the operating system or system firmware is mapping between physical addresses within the HPA space and physical addresses within the DPA space, whereby the mapping enables the second host to utilize the memory and / or the CXL memory expander.
[0099] In one embodiment, an apparatus comprising: a processor comprising a coherent interconnect; the processor is coupled to memory having a capacity of at least 64GB; wherein the processor is configured to utilize physical addresses within a first Host Physical Address (HPA) space to access the memory, and to execute an operating system (OS) that utilizes a virtual address space; a memory management unit (MMU) configured to enable access to the memory, based on mapping addresses within the virtual address space to physical addresses within the first HPA space; a resource provisioning unit (RPU), coupled to a Compute Express Link (CXL) device configured to exchange messages conforming to a protocol based on CXL which utilizes a second HPA space; and wherein the RPU is further coupled to the coherent interconnect and configured to translate physical addresses within the second HPA space to physical addresses within the first HPA space.
[0100] In one embodiment, a system designed to function as a Multi-Headed Device (MHD), comprising: a processor comprising a coherent interconnect; the processor is coupled to dynamic random-access memory (DRAM) having a capacity of at least 32GB; wherein the processor is configured to utilize physical addresses within a Host Physical Address (HPA) space to access the DRAM, and to execute an operating system (OS) that utilizes a virtualaddress space; a memory management unit (MMU) configured to enable access to the DRAM, based on mapping addresses within the virtual address space to physical addresses within the HPA space; first and second Compute Express Link (CXL) Endpoints configured to communicate with hosts coupled to the system according to a protocol based on CXL; and a resource provisioning unit (RPU) configured to perform host-to-host physical address translations which enable the hosts to access the DRAM utilizing messages conforming to the protocol based on CXL.
[0101] The CXL Specification revision 3.2 defines a Multi-Headed Device (MHD) in section 2.5 as a CXL type-3 device with multiple CXL ports, referred to as heads. The CXL specification currently defines two types of MHDs that are distinguished by how they present themselves on each head: (i) a MH-SLD, which presents Single Logical Devices (SLDs) on the heads, and has a 1 : 1 mapping between heads and LDs, and (ii) a MH-MLD, which may present Multi-Logical Devices (MLDs) on any of their heads, wherein a head in a Multi-Headed Device has at least one and no more than 16 Logical Devices mapped.
[0102] Optionally, the DRAM is coupled via at least four memory channels to the processor; wherein the DRAM has a memory capacity exceeding 128 GB, 256 GB, 512 GB, or 1 TB; and wherein the DRAM comprises mainstream DRAM modules exhibiting an average unit price per gigabyte that does not exceed three times an average unit price per gigabyte of a lowest-cost DRAM module technology in volume production for servers in data centers.
[0103] FIG. 1A illustrates one embodiment of a system comprising a memory switch, a memory pool, a Global Fabric -Attached Memory (GF AM) Device (GFD), a memory expander (ME), or a memory expansion device, which comprise a processor, memory (such as DRAM), and an RPU coupled to an entity such as a host. The processor may include processing cores and cache hierarchies that utilize a first HPA space for accessing system resources. The memory may be coupled to the processor via memory channels, such as DDR4 or DDR5 channels, providing high-bandwidth memory access. The RPU may include, or be coupled to, a CXL device (such as a CXL EP), and may be integrated within the same semiconductor device as the processor or implemented as a separate component. The RPU may perform host-to-host physical address translations between the entity's HPA space and the processor's HPA space. The entity may be coupled to the memory pool via the CXL device that supports one or more CXL protocols, enabling the entity to access the memory based on the address translations performed by the RPU.
[0104] FIG. IB illustrates one embodiment of a system comprising a memory pool coupled to hosts and to a memory expander, wherein the memory pool is based on a processor (such as an MxPU) comprising an RPU and CXL devices. The memory pool may include memory tiers, such as a first memory tier (denoted as " 1 ") comprising DRAM coupled via memory channels to the MxPU, and a second memory tier (denoted as "2") comprising DRAM associated with the memory expander. The MxPU may include a CXL RP for coupling to the memory expander, enabling the memory pool to extend its capacity beyond the directly attached DRAM. Multiple hosts may be coupled to the memory pool via separate CXL devices (such as CXL EPs) within the MxPU, wherein the hosts utilize their respective HPA spaces. The RPU within the MxPU may perform different host-to-host physical address translations for the different coupled hosts, enabling concurrent access to both memory tiers while maintaining isolation between different hosts' physical address spaces.
[0105] FIG. 2A illustrates one embodiment of a system comprising a memory pool comprising two or more MxPUs. The memory pool may utilize a chipset-based architecture wherein a collection of electronic components such as MxPUs, xPUs, CPUs, and memory buffers, works together on a platform for realizing a memory pool functionality. The memory pool may include memory tiers, such as a first memory tier (denoted as "1") comprising DRAM coupled via memory channels to the first MxPU, a second memory tier (denoted as "2") comprising DRAM associated with the memory expander coupled to the first MxPU, a third memory tier (denoted as "3") comprising DRAM coupled via memory channels to the second MxPU, and a fourth memory tier (denoted as "4") coupled to the memory buffer that is coupled to the second MxPU. The MxPUs may be interconnected via an ISoL, such as UPI, UXI, Infinity Fabric, or CHI C2C, enabling coherent communication between the MxPUs. Each MxPU may include its own RPU for performing host-to-host physical address translations and CXL devices (such as CXL EPs) for coupling to external hosts, allowing at least some of the external hosts to access the distributed memory resources across memory tiers. The memory buffers may provide additional memory capacity and may include buffer controllogic for managing data flow between different memory tiers.
[0106] FIG. 2B illustrates one embodiment of a system comprising a memory pool comprising at least one MxPU and at least one xPU (that may be a CPU). The memory pool may utilize a chipset-based architecture. The memory pool may include memory tiers, such as a first memory tier (denoted as " 1 ") comprising DRAM coupled to the MxPU, a second memory tier (denoted as "2") comprising DRAM associated with the memory expander coupled to the MxPU, a third memory tier (denoted as "3") comprising DRAM coupled to the xPU / CPU, and a fourth memory tier (denoted as "4") coupled to the memory buffer. The MxPU may include CXL devices (such as CXL EPs) and serve as the primary interface for external hosts to access the memory pool via protocols based on CXL, while the xPU / CPU may provide additional processing capabilities and memory resources. The RPU within the MxPU may coordinate address translations to enable external hosts to access memory resources across the tiers, including memory attached to the xPU / CPU. This embodiment may optimize cost and performance by combining specialized MxPUs for memory pooling with established xPUs / CPUs for processing tasks and additional memory capacity.
[0107] FIG. 3A illustrates one embodiment of a system comprising a memory pool comprising a processor, DRAM, and an RPU. The RPU may include or be coupled to a CXL device. The RPU performs host-to-host physical address translations that enable an entity, external to the memory pool, to access the DRAM coupled to the processor. The processor may include multiple cores, wherein some of the cores may be hidden from the user and may serve for executing infrastructure tasks related to operations, administration and management (0 AM) of the memory pool.
[0108] FIG. 3B illustrates one embodiment of a system comprising a memory pool comprising a CXL Multi Headed Device (MHD), such as Multi-Headed Single Logical Device (MH-SLD) or Multi-Headed Multi-Logical Device (MH-MLD), comprising a processor coupled to DRAM. The processor includes one or more processing cores wherein each processing core may include an MMU. The MHD further comprises CXL endpoints, wherein at least some of the endpoints may be associated with logical devices such as SLDs or MLDs, and an RPU configured to perform host-to-host physical address translations that enable entities external to the MHD to access the DRAM. Optionally, some of the illustrated blocks may be omitted, combined, or implemented as discrete chiplets, IP blocks, or firmware-assisted logic. The number and type of cores is implementation-dependent and may include general-purpose CPUs, vector engines, Al accelerators, or heterogeneous combinations thereof. In alternative or additional embodiments, one or more cores execute processing-in-memory (PIM) operations, for example, reductions, searches, or machine-learning kernels, directly against data resident in the DRAM, thereby reducing link bandwidth consumption. By virtue of the address-translation logic in the RPU, the MHD can expose the DRAM as a shared or partitionable pool that is concurrently accessible by entities via the CXL endpoints, which enables memory pooling, memory sharing, multi-tenant isolation, and / or dynamic capacity provisioning within a CXL-based system.
[0109] FIG. 4 illustrates one embodiment of a system comprising an Al memory switch or a memory pool, comprising a CXL Multi Headed Device (MHD) coupled to two external entities. The memory pool may include additional MHDs coupled to additional entities. The memory pool may utilize a chipset-based architecture wherein a collection of electronic components such as MxPUs, xPUs, CPUs, and memory buffers, works together on a platform for realizing a memory pool functionality. The MHD comprises an MxPU coupled to DRAM, wherein the DRAM may be internal to the MHD, such as mounted on a PCB alongside the MxPU, possibly within an MHD enclosure, or the DRAM may be external to the MHD, such as in pluggable memory modules (e.g., EDSFF). The MxPU may be derived from an established processor design, such as a CPU design that utilizes a combination of at least one compute die and at least one I / O die that may communicate with each other utilizing an on-package interconnect such as AMD Infinity Fabric, ARM CHI C2C, or NVIDIA NVLink-C2C. An RPU, optionally implemented in a separate die / chiplet, or embedded into an I / O die and / or into a compute die, performs host-to-host physical address translations that enable entities coupled to the memory pool via the CXL Endpoints to access the DRAM coupled to the MxPU. The MxPU may include one or multiple chip-to-chip interfaces, such as ISoL, that may provide interconnection of multiple MxPU instances in various topologies to create a larger logical MHD, a distributed MHD, or a memory pool that may serve additional external entities and provide larger memory capacities. The chip-to-chip interface may utilize the same communication protocol utilized by the on-package interconnect links, such as AMD Infinity Fabric, ARM CHI C2C,or NVIDIA NVLink-C2C. Processing cores in the MxPU, optionally hidden cores utilized for infrastructure tasks, may provide Processing In Memory (PIM) services to data residing in the DRAM.
[0110] The explosive growth of artificial intelligence workloads, such as Large Language Models (LLMs) and Generative Al (GenAI) applications, has reshaped datacenter architectures, demanding unprecedented levels of computational power and memory bandwidth. These compute -intensive workloads, alongside High-Performance Computing (HPC) applications such as climate modeling, genomics research, and real-time analytics, require massive parallelization across multiple accelerators while maintaining low-latency access to large memory pools. The convergence of Al training, inference at scale, and traditional HPC workloads has created a paradigm shift where memory bandwidth and capacity have become as important as raw computational throughput, driving the need for advanced interconnect technologies that can efficiently bridge the gap between accelerators and memory resources.
[0111] Ultra Accelerator Link (UALink) has emerged as a high-speed interconnect technology designed to address the demanding requirements of accelerator-based computing architectures. UALink provides protocols and mechanisms for high-bandwidth, low-latency communication between accelerators and other system components, enabling efficient data movement across the compute fabric. As datacenters increasingly adopt heterogeneous computing models combining CPUs, GPUs, and domain-specific accelerators, interconnect technologies should support flexible resource allocation and dynamic memory sharing across different processing elements. The physical address spaces utilized by different accelerators and host processors often operate independently, creating challenges for unified memory access and resource pooling.
[0112] Current interconnect solutions face limitations in enabling efficient memory sharing and resource provisioning across heterogeneous compute environments. The lack of hardware -accelerated address translations between accelerator interconnects and host processor coherent fabrics creates bottlenecks when accelerators need direct access to host memory resources. Furthermore, existing architectures struggle to support multiple accelerators simultaneously accessing shared memory pools when different accelerators utilize different physical address spaces. These challenges underscore the need for architectural solutions that can integrate accelerator interconnects with host processor memory subsystems, providing hardware -accelerated address translation and resource provisioning capabilities that enable efficient memory sharing across heterogeneous compute elements.
[0113] Some of the disclosed embodiments introduce novel system-level architectural solutions leveraging RPUs to enable dynamic memory sharing between accelerators and host processors through UALink -based interconnects. These embodiments provide hardware -accelerated physical address translation capabilities that allow entities connected via UALink to efficiently access host memory resources through the processor's coherent interconnect fabric. Implementing RPUs that perform address space translations between UALink-based protocols and the processor's physical address space, enable memory sharing across heterogeneous compute environments while maintaining compatibility with existing operating systems and MMU -based virtual memory systems. The embodiments address challenges in fields such as memory disaggregation and resource utilization for workloads including AI / ML training and inference, distributed computing, and / or high-performance analytics. Some embodiments optionally support multiple entities accessing shared memory resources through separate UALink-based ports with independent physical address spaces. The integration of RPUs within the processor's coherent interconnect fabric enables low-latency memory access for accelerator-based workloads, optionally supporting in-memory computing paradigms and distributed shared memory models.
[0114] In one embodiment, an apparatus comprises a processor comprising a coherent interconnect, where the coherent interconnect couples processing cores to memory controllers that are coupled to memory channels capable of supporting memory having a capacity of at least 64GB, and the processor is configured to utilize physical addresses within a physical address space (PAS) to access the memory and to execute an operating system (OS) that utilizes a virtual address space. The apparatus further comprises a memory management unit (MMU) configured to enable the OS to access the memory based on mapping addresses within the virtual address space to physical addresses within the PAS. Additionally, the apparatus comprises a resource provisioning unit (RPU) comprising an Ultra Accelerator Link-based port (UALink-based port) configured to communicate with an entity coupled to the apparatus accordingto a UALink-based protocol, wherein the RPU is further coupled to the coherent interconnect and configured to translate physical addresses associated with the UALink-based protocol to physical addresses within the PAS, whereby the physical address translations enable the entity to access the memory via the UALink-based port, the coherent interconnect, and the memory controllers.
[0115] In another embodiment, an apparatus comprises a processor comprising a coherent interconnect that couples processing cores to memory controllers coupled to memory channels capable of supporting memory having a capacity of at least 64GB, wherein the processor is configured to utilize physical addresses within a first physical address space (PAS1) to access the memory and to execute an operating system (OS) that utilizes a virtual address space. The apparatus includes a memory management unit (MMU) configured to enable the OS to access the memory based on mapping addresses within the virtual address space to physical addresses within the PAS1. The apparatus further comprises first and second resource provisioning units (RPUs) comprising first and second respective UALink-based ports configured to communicate, according to a UALink-based protocol, with first and second respective entities coupled to the apparatus, whereby the first and second entities utilize second and third respective physical address spaces (PAS2, PAS3). The first and second RPUs are further coupled to the coherent interconnect, wherein the PAS 1, PAS2, and PAS3 are different, and whereby the apparatus is capable of enabling the first and second entities to access portions of the memory via the first and second UALink-based ports, the coherent interconnect, and the memory controllers.
[0116] In yet another embodiment, a method comprises operating a processor comprising a coherent interconnect that couples processing cores to memory controllers that communicate with memory channels coupled to memory having a capacity of at least 64GB . The method further comprises utilizing, by the processor, physical addresses within a physical address space (PAS) to access the memory, and executing, by the processor, an operating system (OS) that utilizes a virtual address space. The method includes mapping addresses within the virtual address space to physical addresses within the PAS, which enables the OS to access the memory. Additionally, the method comprises communicating according to a UALink-based protocol with an entity via a UALink-based port, and performing physical address translations from physical addresses associated with the UALink-based protocol to physical addresses within the PAS, whereby the physical address translations enable the entity to access the memory via the UALink-based port, the coherent interconnect, and the memory controllers.
[0117] Modem datacenters require efficient mechanisms for memory resource sharing between accelerators and host processors to support AI / ML workloads, HPC applications, and distributed computing environments. Embodiments herein disclose systems incorporating RPUs that enable entities to access host memory through UALink-based interconnects. The processor utilizes a coherent interconnect coupling processing cores to memory controllers, with an MMU mapping virtual addresses to physical addresses within the processor's physical address space. The RPU performs hardware -accelerated physical address translations between UALink-associated addresses and the processor's physical address space, enabling entities to access memory via the UALink port, coherent interconnect, and memory controllers. Some embodiments support multiple RPUs with independent UALink ports serving entities with distinct physical address spaces, enabling dynamic resource allocation and memory pooling, which address memory disaggregation challenges for GenAI inference, LLM training, and next-generation datacenter architectures requiring flexible memory sharing across heterogeneous compute elements.
[0118] In one embodiment, an apparatus comprising: a processor comprising a coherent interconnect, the coherent interconnect couples processing cores to memory controllers that are coupled to memory channels capable of supporting memory having a capacity of at least 64GB; wherein the processor is configured to utilize physical addresses within a physical address space (PAS) to access the memory, and to execute an operating system (OS) that utilizes a virtual address space; a memory management unit (MMU) configured to enable the OS to access the memory, based on mapping addresses within the virtual address space to physical addresses within the PAS; a resource provisioning unit (RPU) comprising an Ultra Accelerator Link -based port (UALink-based port) configured to communicate with an entity coupled to the apparatus according to a UALink-based protocol; and wherein the RPU is further coupled to the coherent interconnect and configured to translate physical addresses associated with theUALink-based protocol to physical addresses within the PAS; whereby the physical address translations enable the entity to access the memory via the UALink-based port, the coherent interconnect, and the memory controllers. Optionally, the address translations between the physical addresses may enable isolation between different address domains while allowing controlled access to system resources. The translation logic may support various mapping schemes including offset-based translation, page-table-based translation, or range-based translation. The RPU may include translation lookaside buffers (TLBs) or other caching mechanisms to optimize translation performance for frequently accessed address ranges.
[0119] Optionally, the UALink-based protocol conforms to UALink Protocol Level Interface (UPLI), the physical addresses associated with the UPLI comprise network physical addresses (NPAs), and the physical addresses within the PAS comprise system physical addresses (SPAs) or host physical addresses (HP As). Optionally, embodiments may utilize a global or a flat addressing model, wherein a single address space may include addresses that may be utilized for accessing memory within a system domain, and may also include addresses associated with UPLI that may be utilized for accessing memory in different system domains, wherein physical address translations may be performed between physical addresses within that single address space. Alternatively, embodiments may utilize multiple physical address spaces, such as NPA space (wherein NPAs may be utilized for accessing memory in different system domains) and SPA space (wherein SPAs may be utilized for accessing memory within a system domain), wherein physical address translations may be performed from NPAs to SPAs. In some embodiments, the NPAs may represent addresses within a global or a flat UALink fabric address space that may span multiple nodes or devices, or may represent addresses within a destination UALink Accelerator referenced by a destination identifier in the UPLI. SPAs or HP As (an implementation choice) may represent the local addressing scheme utilized by the processor or by the system node (SN). The translation from NPAs to SPAs / HPAs may include routing information extraction, node identifier processing, and / or address offset calculations to map fabric -side addresses to local memory locations.
[0120] Optionally, in addition to the physical address translations, the RPU is further configured to translate between first fields conforming to UALink-based protocol message formats, and second fields conforming to message formats of a protocol utilized by the coherent interconnect. Optionally, the protocol utilized by the coherent interconnect is based on Coherent Hub Interface (CHI -based protocol), and the RPU is further configured to translate read requests corresponding to the UALink-based protocol to requests corresponding to the CHI-based protocol carrying ReadOnce for non-cacheable data access or ReadShared for cacheable data access. Optionally, the selection between ReadOnce and ReadShared may be determined by cache allocation hints, memory region attributes, or explicit indicators in the UPLI request. The RPU may additionally translate UPLI write requests to other CHI write opcodes based on write granularity and coherency requirements. The translation may preserve transaction ordering by utilizing CHI's ordering rules and potentially implementing additional ordering enforcement logic when UALink ordering requirements exceed those provided by CHI.
[0121] Optionally, the protocol utilized by the coherent interconnect is based on an Intel Coherent Processor Interconnect Protocol (ICPIP -based protocol) for scalable multiprocessors with a shared physical address space, and wherein the RPU is further configured to translate read requests corresponding to the UALink-based protocol to requests corresponding to the ICPIP-based protocol carrying RdCur opcodes, and maintain coherency state information for physical addresses within the PAS that are associated with the coherent interconnect. Optionally, when the coherent interconnect is based on UPI protocol, the coherency state information maintained by the RPU may include caching states or similar cache coherency states. Examples of ICPIP include Intel’s QPI, UPI, KTI, UXI, and future Intel’s Coherent Processor Interconnect Protocols. The translation to ICPIP RdCur opcodes, such as UPI RdCur opcodes, may be accompanied by snoop responses handling when the requested data exists in other processor caches. The RPU may include state tracking mechanisms to optimize subsequent accesses to the same cache lines.
[0122] Optionally, the protocol utilized by the coherent interconnect is based on Infinity Fabric protocol (IF -based protocol); and wherein the RPU is further configured to translate write requests corresponding to the UALink-based protocol to write commands corresponding to the IF -based protocol while preserving write ordering required by the entity. Optionally, the preservation of write ordering may include tracking write dependencies and enforcingcompletion ordering as specified by the UALink memory model. The RPU may translate UPLI write requests that include Byte Enables (which indicate partial writes) to appropriate Infinity Fabric write command types while maintaining producer-consumer ordering relationships. For example, the UPLI 64-bit byte enable field (OrigDataByteEn), which allows for individual bytes within a data beat to written or not in a write transfer, may be translated by the RPU to the appropriate Infinity Fabric write command type.
[0123] Optionally, the RPU is further configured to translate a request corresponding to the UALink -based protocol to at least one message corresponding to the protocol utilized by the coherent interconnect; and wherein the at least one message causes prefetch to a cache of the processor. Optionally, the RPU may translate UPLI requests, such as UPLI prefetch hints that may be carried in vendor-defined commands, to messages of a protocol utilized by the coherent interconnect that effectively prefetch data into a cache of the processor, such as LLC Prefetch RFO (LlcPrefRFO), LLC Prefetch Code (LlcPrefCode), or LLC Prefetch Data (LlcPrefData) opcodes of a protocol based on Intra-Die Interconnect (IDI), which is the protocol used by some Intel processor cores.
[0124] Optionally, the RPU is further configured to: translate Tags associated with transactions corresponding to the UALink-based protocol to Tags utilized by the coherent interconnect, and maintain a mapping between the Tags associated with the transactions and the Tags utilized by the coherent interconnect. Optionally, the Tag translation logic may accommodate different Tag formats and sizes between the UPLI and coherent interconnect domains. Tags may be used to identify a transaction, such as when supporting outstanding requests in-flight through the RPU, or may be used to convey properties associated with messages or transactions, such as trace Tags used for debugging and performance measurements, or authorization Tags used for security. Tags may be referenced by different names in different embodiments, such as by the name Transaction identifier (TxniD) in some ARM CHi implementations. The mapping between UPLI Tags and coherent interconnect Tags may include using on-silicon SRAM, content-addressable memory (CAM) or Ternary Content-Addressable Memory (TCAM) structures, hash tables, or indexed arrays. The RPU may handle Tag exhaustion scenarios by including flow control mechanisms that prevent new transactions when available Tags are depleted.
[0125] Optionally, the RPU is further configured to: maintain a Tag allocation table to track outstanding transactions from the entity, allocate coherent interconnect Tags from a pool of available Tags upon receiving requests conforming to the UALink-based protocol, and release the Tags upon completion of corresponding transactions. Optionally, the Tag allocation table may be sized to support the maximum number of outstanding transactions allowed by the UALink specification or by configured limits. The Tag pool management may implement various allocation techniques including round-robin, least-recently-used, or priority -based allocation. The RPU may monitor Tag utilization to detect potential bottlenecks and may include Tag recycling logic to handle long-latency transactions efficiently.
[0126] Optionally, the entity is configured to access the memory utilizing read and write requests conforming to the UALink-based protocol, wherein the read and write requests are translated by the RPU; and the processing cores are configured to access entity -attached resources by issuing coherent interconnect requests that the RPU is further configured to translate to transactions conforming to the UALink-based protocol, wherein the transactions target the entity. Optionally, the bidirectional access capability may enable various computing paradigms including memory pooling, memory sharing, resource disaggregation, and heterogeneous computing. When processing cores access entity -attached resources, such as High-Bandwidth Memory (HBM) resources, the RPU may handle different memory attributes, caching policies, and ordering requirements between the two domains. The translation of coherent interconnect requests to transactions conforming to the UALink-based protocol may include protocol-specific adaptations to maintain correctness across domain boundaries.
[0127] Optionally, the entity comprises entity -attached memory; and wherein the RPU is further configured to map a portion of the entity -attached memory into the PAS, enabling the processing cores to access the entity -attached memory utilizing load and store operations. Optionally, the mapping of entity -attached memory into PAS may include establishing memory windows with specific attributes such as cacheability, write-combining behavior, and / or access permissions. The RPU may support dynamic remapping of entity -attached memory regions based on workloadrequirements or system configuration changes. The load and store operations from processing cores may be subject to memory ordering mles enforced by both the coherent interconnect and UPLI.
[0128] Optionally, the RPU is further configured to enforce access control by comparing the physical addresses associated with the UALink-based protocol against a set of predetermined allowed address ranges for the entity, and blocking transactions that fall outside the predetermined allowed address ranges. Optionally, the predetermined allowed address ranges may be configured by privileged software, firmware, or hardware configuration registers. The RPU may support multiple security contexts with different predetermined allowed address ranges for different entities or different operational modes. The blocking of unauthorized transactions may generate error responses, security exceptions, or logging events for system monitoring and debugging purposes.
[0129] Optionally, the RPU is further configured to apply security filtering based on examination of transaction attributes associated with the UALink-based protocol, which include requester identification and access permissions, and selectively allowing or denying transactions based on preconfigured security policies. Optionally, the security filtering may examine additional UPLI transaction attributes, such as vendor-defined commands or fields, virtual channel identifiers, traffic classes, or custom security tokens. The preconfigured security policies may be stored in secure storage within the RPU or loaded from trusted system firmware during initialization. The RPU may support dynamic policy updates under appropriate authentication and authorization mechanisms.
[0130] Optionally, the RPU is further configured to: detect sequential access patterns in requests corresponding to the UALink-based protocol which are received from the entity, and issue prefetch requests that are routed via the coherent interconnect and the memory controllers to retrieve data in advance of anticipated entity requests. Optionally, the prefetch mechanism may utilize various pattern detection algorithms including stride detection, stream detection, or machine learning-based prediction. The RPU may maintain prefetch buffers to store prefetched data and may implement prefetch throttling to prevent memory bandwidth saturation. The prefetch requests may be marked with lower priority than demand requests to minimize interference with the explicit memory accesses.
[0131] Optionally, the memory comprises dynamic random-access memory (DRAM), and the entity comprises a graphics processing unit (GPU) or a central processing unit (CPU) configured to utilize the UALink-based port for accessing the memory; and wherein the RPU enables the entity to access the DRAM with cache-line granularity. An entity, such as a GPU or a CPU, may utilize the UALink port for high-bandwidth memory access to memory resources attached to the processor. Optionally, the cache -line granularity access may align with standard cache line sizes such as 64 bytes, 128 bytes, or 256 bytes, enabling efficient data transfers between the entity and the DRAM. The RPU may support memory consistency maintenance, which includes coordination between the entity's memory model and the processor's memory model, with the RPU translating between different consistency requirements. The high-bandwidth memory access may be optimized utilizing features such as memory interleaving, bank -aware scheduling, or quality-of-service mechanisms that prioritize latency -sensitive or bandwidth-intensive access patterns from the GPU or CPU entity.
[0132] Optionally, the RPU is further configured to coalesce coherent interconnect transactions targeting contiguous or nearby addresses into fewer requests corresponding to the UALink-based protocol; whereby the coalescing reduces transaction overhead and improves memory bandwidth utilization. Optionally, the request coalescing may consider factors including address proximity, request types, and timing windows when determining which requests to combine. The RPU may include write combining buffers for write requests and may support read coalescing for sequential read patterns. Optionally, the RPU is further configured to utilize an intermediate protocol selected from Peripheral Component Interconnect Express (PCIe) or Compute Express Link (CXL) when translating between the UALink-based protocol and a protocol utilized by the coherent interconnect.
[0133] In one embodiment, an apparatus comprising: a processor comprising a coherent interconnect, the coherent interconnect couples processing cores to memory controllers that are coupled to memory channels capable of supporting memory having a capacity of at least 64GB; wherein the processor is configured to utilize physical addresses within a first physical address space (PAS1) to access the memory, and to execute an operating system (OS) that utilizes a virtual address space; a memory management unit (MMU) configured to enable the OS to access thememory, based on mapping addresses within the virtual address space to physical addresses within the PAS1; first and second resource provisioning units (RPUs) comprising first and second respective Ultra Accelerator Link -based ports (UALink-based ports) configured to communicate, according to a UALink -based protocol, with first and second respective entities coupled to the apparatus, whereby the first and second entities utilize second and third respective physical address spaces (PAS2, PAS3); and wherein the first and second RPUs are further coupled to the coherent interconnect; wherein the PAS1, the PAS2, and the PAS3 are different; and whereby the apparatus is capable of enabling the first and second entities to access portions of the memory via the first and second UALink-based ports, the coherent interconnect, and the memory controllers.
[0134] Optionally, the first RPU is configured to translate physical addresses within the PAS2 to physical addresses within the PAS1; and wherein the second RPU is configured to translate physical addresses within the PAS3 to physical addresses within the PAS1; whereby the first and second RPUs enable the first and second entities to access the memory. Optionally, the UALink-based protocol conforms to UALink Protocol Level Interface (UPLI); and in addition to the physical address translations, the first and second RPUs are further configured to translate between first fields conforming to the UPLI message formats, and second fields conforming to second message formats of a protocol utilized by the coherent interconnect.
[0135] Optionally, the protocol utilized by the coherent interconnect is based on Coherent Hub Interface (CHIbased protocol), and at least one of the first and second RPUs is further configured to translate UPLI read requests to CHI-based requests carrying ReadOnce for non-cacheable data access or ReadShared for cacheable data access.
[0136] Optionally, the protocol utilized by the coherent interconnect is based on Intel’s Coherent Processor Interconnect Protocol (ICPIP -based protocol) for scalable multiprocessors with a shared physical address space, and at least one of the first and second RPUs is further configured to translate read requests corresponding to the UPLI to requests corresponding to the ICPIP -based carrying opcodes based on RdCur.
[0137] Optionally, the protocol utilized by the coherent interconnect is based on Infinity Fabric (IF -based); and wherein at least one of the first and second RPUs is further configured to translate write requests corresponding to the UPLI to write commands corresponding to the IF-based while preserving write ordering required by the respective entity.
[0138] Optionally, at least one of the first and second RPUs is further configured to translate Tags associated with transactions corresponding to the UALink-based protocol to Tags utilized by the coherent interconnect, maintain a mapping between the Tags associated with transactions and the Tags utilized by the coherent interconnect, and translate response Tags associated with the coherent interconnect back to response Tags associated with the UALink-based protocol.
[0139] Optionally, the first RPU maintains a first translation table for mapping addresses within the PAS2 to addresses within the PAS1, and the second RPU maintains a second translation table for mapping addresses within the PAS3 to addresses within the PAS1; wherein the first and second translation tables are different and provide isolation between memory accesses from the first and second entities. Optionally, the first RPU is configured to translate addresses within the PAS2 to a first subset of addresses within the PAS1, and the second RPU is configured to translate addresses within the PAS3 to a second subset of addresses within the PAS1, wherein the first and second subsets are non-overlapping. Optionally, the first RPU is configured to translate at least some addresses within the PAS2 to a shared subset of addresses within the PAS1, and the second RPU is configured to translate at least some addresses within the PAS3 to the same shared subset of addresses within the PAS1, enabling the first and second entities to access shared memory regions. Optionally, the PAS2 has a different size than the PAS3, and wherein the PAS2 and the PAS3 have different sizes than the PAS 1 ; and wherein the first and second RPUs are further configured to dynamically modify the address translations between the PAS2 and the PAS 1 , and between the PAS3 and the PAS 1 , based on memory allocation requests or reconfiguration commands.
[0140] In one embodiment, a method comprising: operating a processor comprising a coherent interconnect that couples processing cores to memory controllers that communicate with memory channels coupled to memory having a capacity of at least 64GB; utilizing, by the processor, physical addresses within a physical address space (PAS) toaccess the memory; executing, by the processor, an operating system (OS) that utilizes a virtual address space; mapping addresses within the virtual address space to physical addresses within the PAS, which enables the OS to access the memory; communicating according to a protocol based on Ultra Accelerator Link (UALink-based protocol) with an entity via a UALink-based port; and performing physical address translations from physical addresses associated with the UALink-based protocol to physical addresses within the PAS; whereby the physical address translations enable the entity to access the memory via the UALink-based port, the coherent interconnect, and the memory controllers. Optionally, the coherent interconnect utilizes a protocol based on Coherent Hub Interface (CHIbased protocol), and wherein, in addition to performing the physical address translations, further comprising: (a) translating between a first field conforming to the UALink-based protocol message format, and a second field conforming to the CHI-based protocol message format; and (b) translating UALink-based protocol read requests to CHI-based protocol requests carrying ReadOnce for non-cacheable data access or ReadShared for cacheable data access.
[0141] FIG. 5A illustrates one embodiment of a system where an entity, such as a GPU or accelerator, communicates via a UALink port included in an RPU that further includes a Coherent Interconnect Interface that may utilize a protocol based on ARM CHI. The Coherent Interconnect Interface couples the RPU to an interconnect component, such as a crosspoint (XP), within a coherent interconnect. The Coherent Interconnect Interface performs the necessary protocol conversions between a UALink domain and a coherent interconnect domain, such as between UPLI and ARM CHI, enabling the entity to access memory and other resources coupled to the coherent interconnect. The coherent interconnect may be implemented as a mesh topology connecting various components including processing cores, home nodes (HN), memory controllers (MC), and accelerator cores.
[0142] FIG. 5B illustrates one embodiment of a TFD showing address translation between UALink UPLI and CHI. An entity, such as a GPU, initiates a UPLI request comprising a physical address (AS.2.1), which the RPU translates to a CHI request carrying ReadOnce with a translated physical address (AS.1.1). The transaction flows through the coherent interconnect via a home node to a memory controller, which retrieves the data and returns it, through the coherent interconnect, to the RPU that translates the response back to the UPLI domain for delivery to the requesting entity.
[0143] FIG. 6A illustrates one embodiment of a system that may function as a UALink memory switch appliance or a UALink memory pool, and may include an accelerator, GPU, xPU / MxPU, or a memory switch ASIC, that is coupled to two entities denoted as Entity.1 / GPU.1 and Entity.2 / GPU.2. The accelerator / xPU includes processing cores and memory controllers coupled to a coherent interconnect that may be based on CHI. The accelerator / xPU utilizes translations, performed the RPUs, between UALink-based ports and the coherent interconnect. The first RPU (RPU.1) may enable Entity.1 / GPU.1 to access, via the first UALink port and the coherent interconnect, resources mapped to a physical address space utilized by the coherent interconnect, such as memory resources of the accelerator / xPU. Correspondingly, the second RPU (RPU.2) may enable Entity.2 / GPU.2 to access, via the second UALink port and the coherent interconnect, resources mapped to a physical address space utilized by the coherent interconnect, such as memory resources of the accelerator / xPU.
[0144] FIG. 6B illustrates one embodiment of a TFD depicting a multi-entity memory access scenario wherein first and second entities / GPUs access memory mapped to one or more physical address spaces utilized by the coherent interconnect, through UPLI-to-ARM CHI translations. Entity.1 / GPU.1 initiates a first UPLI Request (Req) with ReqCmd(Read), ReqSrcPhysAccID(a.l) to identify the source entity, such as an accelerator or a GPU, ReqDstPhysAccID(b.l) to identify the destination entity / accelerator, and ReqAddr(AS.2.1) representing a UPLI request address, such as a network physical address (NPA) from a second physical address space. RPU.l translates the first UPLI Request to ARM CHI REQ carrying Opcode(ReadOnce) and Addr(AS.1.1) from a first physical address space utilized by the coherent interconnect. Concurrently or sequentially, Entity.2 / GPU.2 may initiate a second UPLI Request (Req) with ReqCmd(Read), ReqSrcPhysAccID(a.2), ReqDstPhysAccID(b.2), and ReqAddr(AS.3.1) representing a UPLI request address, such as a network physical address (NPA) optionally from a third physical address space or from the second physical address space. RPU.2 translates the second UPLI Request to ARM CHIREQ carrying Opcode(ReadOnce) and Addr(AS.1.2) from the first physical address space utilized by the coherent interconnect. Both transactions flow through the coherent interconnect to one or more home nodes, which may send respective ARM CHI REQ messages to one or more memory controllers with Opcode(ReadNoSnp) and the addresses Addr(AS.1.1) and Addr(AS.1.2), respectively.
[0145] The memory controller(s) retrieve the requested data from the memory and send first and second ARM CHI RD AT messages with Opcode(CompData) carrying *Data.1* and *Data.2*, representing the data retrieved from the addresses AS.1.1 and AS.1.2, respectively. RPU.l translates the first ARM CHI RD AT message to UPLI Read Response / Data (RdRsp) with RdRspSrcPhysAccID(b.l), RdRspDstPhysAccID(a.l), and RdRspData(*Data.l*) for Entity.1 / GPU.l. RPU.2 translates the second ARM CHI RD AT message to UPLI RdRsp with RdRspSrcPhysAccID(b.2), RdRspDstPhysAccID(a.2), and RdRspData(*Data.2*) for Entity.2 / GPU.2. The embodiment demonstrates how entities / GPUs may share access to the same memory through different RPUs that perform both translation between UPLI and ARM CHI, and physical address translations. Alternatively, the embodiment may be viewed as two separate UPLI transactions that utilize the same coherent interconnect infrastructure to access the memory, wherein entities such as accelerators or GPUs may access the memory via a shared or a separate address space that may be translated to a shared coherent interconnect physical address space. Still alternatively, the response and read data paths may be implemented according to other designs, such as wherein the memory controller(s) may send the data to the home node(s) that send it to the respective RPUs, or the home node(s) send responses to the RPUs while the memory controller(s) send the data directly to the RPUs.
[0146] Modem datacenters face challenges in memory resource utilization and sharing as workloads become increasingly memory -intensive and distributed. Applications spanning artificial intelligence (Al), machine learning (ML), Large Language Model (LLM) inference, database analytics, and virtualized environments require flexible access to large memory pools that may exceed the capacity limitations of individual servers. The emergence of generative Al (GenAI) and deep learning workloads has intensified demands for high-bandwidth memory access across heterogeneous computing resources, including CPUs, GPUs, and specialized accelerators. These evolving demands have driven the development of memory disaggregation technologies and high-speed interconnects that enable more efficient utilization of datacenter infrastructure.
[0147] NVLink and NVLink-compatible interconnect technologies have emerged as high-bandwidth, low-latency communication solutions for connecting processors and accelerators within computing systems. These interconnect protocols enable direct memory access between devices, supporting data rates that exceed traditional PCIe capabilities. NVLink-based architectures allow GPUs and other accelerators to communicate with each other and with host processors through dedicated high-speed links. However, integrating NVLink-connected devices with host processor memory systems presents challenges when different components operate within separate physical address spaces.
[0148] Current implementations of NVLink-based systems face limitations when accelerators need to access host memory resources. The separation between device physical address spaces and host physical address spaces creates barriers to efficient memory sharing. While NVLink provides high-bandwidth connectivity, the lack of address translations between NVLink-connected devices and host memory subsystems restricts the ability to leverage host memory as an extension of device memory. This limitation becomes apparent in memory -intensive workloads where accelerators could benefit from accessing large host memory pools beyond their local device memory capacity. There is a need for solutions that enable NVLink-connected devices to access host memory resources through address translations that bridge different physical address spaces while maintaining the high-bandwidth advantages of NVLink-based interconnects.
[0149] Some of the disclosed embodiments introduce novel system-level architectural solutions leveraging RPUs to enable NVLink-connected devices to access host memory resources. These embodiments provide address translation capabilities that allow entities such as GPUs, accelerators, CPUs, and other NVLink-capable devices to access host memory through NVLink-based protocols, overcoming traditional boundaries between device and host physical address spaces. By implementing RPUs that translate between physical addresses associated with NVLink-based protocols and host physical address spaces, the embodiments enable flexible memory architectures whereNVLink-connected devices can utilize host memory as extended memory resources. Some embodiments optionally support memory disaggregation and memory pooling configurations, enabling multiple NVLink-connected devices to share host memory resources. The address translations enable efficient memory sharing for contemporary workloads including AI / ML training and inference, LLM deployment, distributed computing, and high-performance computing applications.
[0150] In one embodiment, an apparatus comprises an integrated circuit comprising processing cores comprising memory management units (MMUs) and coherent caches, wherein the processing cores are configured to respond to snoop requests that utilize physical addresses within a physical address space (PAS), and wherein the MMUs are configured to translate virtual addresses to physical addresses within the PAS. The apparatus further comprises a coherent interconnect coupling the processing cores to memory controllers coupled to memory channels capable of supporting memory having a capacity of at least 64GB, and wherein the processing cores are configured to execute an operating system (OS) that accesses the memory utilizing the physical addresses within the PAS. Additionally, the apparatus comprises a resource provisioning unit (RPU) comprising an NVLink-based interface configured to communicate, according to an NVLink-based protocol, with an entity coupled to the apparatus. And the RPU is further coupled to the coherent interconnect and configured to translate physical addresses associated with the NVLink-based protocol to physical addresses within the PAS, whereby the physical address translations enable the entity to access the memory via the NVLink-based interface and the memory controllers.
[0151] In another embodiment, a method for enabling an entity to access memory via an NVLink-based interface and memory controllers comprises operating a processor comprising processing cores, memory management units (MMUs), and coherent caches, wherein the processing cores respond to snoop requests that utilize physical addresses within a physical address space (PAS), and the MMUs translate virtual addresses to physical addresses within the PAS. The method further comprises communicating, via a coherent interconnect, between the processing cores and the memory controllers that communicate with memory channels coupled to memory having a capacity of at least 64GB. The method additionally comprises executing, by the processing cores, an operating system (OS) that accesses the memory utilizing the physical addresses within the PAS. The method also comprises communicating according to an NVLink-based protocol with the entity via an NVLink-based interface, and translating physical addresses associated with the NVLink-based protocol to physical addresses within the PAS.
[0152] In yet another embodiment, a system comprises a host processor, a memory having a capacity of at least 64GB, and a coherent interconnect architecture coupling processing elements to the memory, wherein the processing elements utilize a local physical address space to access the memory. The system further comprises a resource provisioning unit (RPU) configured to translate physical addresses associated with an NVLink-based protocol, utilized by an entity coupled to the RPU via an NVLink-based interface, to physical addresses within the local physical address space, whereby the physical address translations enable the entity to utilize the memory as disaggregated memory accessed via the NVLink-based interface and the memory controllers.
[0153] Some embodiments disclose systems and methods for efficient memory resource sharing across heterogeneous computing environments to support Al workloads, LLM inference, and high-performance computing applications. The embodiments utilize an RPU that performs address translations between NVLink-based protocols and host physical address spaces, enabling GPUs, accelerators, and other NVLink-capable devices to access host memory resources. The system includes processing cores with MMUs, a coherent interconnect coupling the cores to memory controllers supporting more than 64GB of memory, and an RPU with an NVLink-based interface. The RPU translates physical addresses associated with the NVLink-based protocol to physical addresses within the host's physical address space, enabling entities to access host memory via the NVLink-based interface. The embodiments optionally support memory disaggregation and pooling configurations, enabling flexible memory architectures and improved resource utilization suitable for GenAI workloads, distributed computing, and next-generation datacenter deployments.
[0154] In one embodiment, an apparatus comprising: an integrated circuit comprising processing cores comprising memory management units (MMUs) and coherent caches; wherein the processing cores are configured to respond tosnoop requests that utilize physical addresses within a physical address space (PAS), and wherein the MMUs are configured to translate virtual addresses to physical addresses within the PAS; a coherent interconnect coupling the processing cores to memory controllers coupled to memory channels capable of supporting memory having a capacity of at least 64GB, and wherein the processing cores are configured to execute an operating system (OS) that accesses the memory utilizing the physical addresses within the PAS; a resource provisioning unit (RPU) comprising an NVLink-based interface configured to communicate, according to an NVLink -based protocol, with an entity coupled to the apparatus; and wherein the RPU is further coupled to the coherent interconnect and configured to translate physical addresses associated with the NVLink-based protocol to physical addresses within the PAS; whereby the translate of the physical addresses enables the entity to access the memory via the NVLink-based interface and the memory controllers.
[0155] Optionally, the NVLink-based interface comprises at least one differential pair and is configured to support reliable communication by utilizing at least one of: a replay buffer configured to enable retransmissions of packets that were not positively acknowledged by a receiver, or a Forward Error Correction (FEC) code configured to enable correction of symbol errors. Optionally, in addition to the physical address translations, the RPU is further configured to translate between first fields conforming to the NVLink-based protocol message formats, and second fields conforming to message formats of a protocol utilized by the coherent interconnect.
[0156] Optionally, the protocol utilized by the coherent interconnect is based on Coherent Hub Interface (CHIbased protocol), and the RPU is further configured to translate read requests corresponding to the NVLink-based protocol to requests corresponding to the CHI-based protocol carrying ReadOnce or ReadShared. Optionally, the RPU may further translate CHI responses to NVLink responses, such as CHI responses carrying CompData to NVLink responses. Additionally, the RPU may maintain transaction context to properly correlate requests and responses across the protocol domains. The translation to CHI ReadOnce may be utilized for non-cacheable data accesses, while ReadShared may be utilized for cacheable shared data. The RPU may handle protocol-specific differences in flow control, credit management, and response ordering between the NVLink and CHI domains. The CompData responses from CHI may carry the requested data along with completion status, which the RPU translates into appropriate NVLink response formats.
[0157] Optionally, the protocol utilized by the coherent interconnect is based on an Intel Coherent Processor Interconnect Protocol (ICPIP -based protocol) for scalable multiprocessors with a shared physical address space, and wherein the RPU is further configured to translate memory access requests corresponding to the NVLink-based protocol to requests corresponding to the ICPIP-based protocol, while maintaining coherency state tracking for physical addresses within the PAS that are associated with the coherent caches. Examples of ICPIP include Intel’s Ultra Path Interconnect (UPI), KTI, UXI, and future Intel’s Coherent Processor Interconnect Protocols. Optionally, the coherency state tracking between NVLink and ICPIP domains may include monitoring cache line states and ensuring consistency across protocol boundaries. The RPU may include state machines to track outstanding transactions and their coherency implications. The translation may accommodate differences in data transfer granularity and response timing between NVLink and ICPIP protocols.
[0158] Optionally, the protocol utilized by the coherent interconnect is based on Infinity Fabric (IF -based), and wherein the RPU is further configured to translate NVLink-based traffic to IF -based traffic, while preserving memory ordering required by the entity. Optionally, the preservation of memory ordering may include tracking command dependencies and enforcing completion ordering as required by both NVLink and Infinity Fabric specifications. The RPU may include ordering enforcement logic that respect producer-consumer relationships and memory barrier semantics across the protocol boundary. The RPU may translate NVLink commands that include partial write indicators to appropriate Infinity Fabric write command types while maintaining data integrity.
[0159] Optionally, the RPU is further configured to translate commands or encodings associated with the NVLink-based protocol to commands or opcodes associated with a protocol utilized by the coherent interconnect, based on a mapping between request types of the NVLink-based protocol and corresponding request types of the protocol utilized by the coherent interconnect. Optionally, the mapping may be implemented utilizing lookup tables, state machines, orprogrammable translation logic. The RPU may handle various NVLink categories including memory reads, memory writes, and atomic operations, translating them to appropriate coherent interconnect opcodes while preserving transaction semantics.
[0160] Optionally, the RPU is further configured to translate a request corresponding to the NVLink -based protocol to at least one message corresponding to the protocol utilized by the coherent interconnect; wherein the at least one message causes prefetch to a cache of a processor comprising the processing cores. Optionally, the RPU may translate NVLink requests, such as requests carrying explicit or implicit prefetch hints, to messages of a protocol utilized by the coherent interconnect that effectively prefetch data into a cache of the processor, enabling reduced memory access latency for anticipated future accesses. An example of a prefetch hint may include a case wherein the RPU detects a pattern of reading pairs of addresses that are adjacent to each other or separated by a distinguishable stride.
[0161] Optionally, the RPU is further configured to utilize an intermediate protocol selected from Peripheral Component Interconnect Express (PCIe) or Compute Express Link (CXL) when translating between the NVLink -based protocol and a protocol utilized by the coherent interconnect. Optionally, the use of an intermediate protocol may facilitate translation by leveraging existing protocol conversion logic. When utilizing PCIe as an intermediate protocol, the RPU may translate NVLink traffic to PCfe Transaction Layer Packets (TLPs) and subsequently to coherent interconnect transactions. When utilizing CXL as an intermediate protocol, the RPU may leverage CXL.cache or CXL.mem as appropriate for the transaction type. The intermediate protocol stage may enable reuse of existing protocol bridges and translation logic.
[0162] Optionally, the RPU is further configured to maintain mappings between transaction identifiers utilized by the NVLink-based protocol and transaction identifiers utilized by the coherent interconnect, enabling correlation of requests and responses across domains. Optionally, the transaction identifier mappings may accommodate different identifier formats, sizes, and allocation schemes between NVLink and the coherent interconnect. Transaction identifiers may be used to identify a transaction, such as when supporting multiple outstanding requests in-flight through the RPU, or may be used to convey properties associated with messages or transactions, such as trace identifiers used for debugging and performance measurements, or authorization identifiers used for security. The RPU may include identifier pools and allocation mechanisms to prevent identifier exhaustion and may support identifier recycling upon transaction completion. The mapping structures may be optimized for fast lookup during high-frequency transaction processing and may utilize on-silicon SRAM, content-addressable memory (CAM) or Ternary Content-Addressable Memory (TCAM) structures.
[0163] Optionally, the RPU is further configured to: maintain a transaction tracking structure to monitor outstanding transactions from the entity, allocate coherent interconnect transaction identifiers for transactions initiated by the RPU, and release identifiers upon transaction completion. Optionally, the transaction tracking structure may be implemented using content-addressable memories, linked lists, or circular buffers optimized for the expected transaction rates. The RPU may include timeout logic to handle lost or excessively delayed transactions and may support error recovery procedures. The tracking structure may maintain additional transaction attributes such as timestamps, retry counts, or quality -of-service parameters.
[0164] Optionally, the RPU is further configured to enable bidirectional access by translating requests between messages conforming to the NVLink-based protocol and messages conforming to the protocol utilized by the coherent interconnect; whereby the entity accesses the memory according to the NVLink-based protocol, and the processing cores access resources attached to the entity via the coherent interconnect. Optionally, the bidirectional access capability may enable memory pooling and memory sharing architectures wherein system memory and entity -attached memory form a memory space accessible from both domains via translations. The RPU may maintain separate translation contexts for each direction and may apply different translation policies based on the initiator and target of each transaction. The bidirectional capability may support various computing paradigms including GPU -direct operations and peer-to-peer transfers. When processing cores access entity -attached resources, such as High-Bandwidth Memory (HBM) resources, the RPU may handle different memory attributes between the two domains.
[0165] Optionally, the entity comprises at least one of: high-bandwidth memory (HBM), Low-Power Double Data Rate (LPDDR) memory, or Graphics Double Data Rate (GDDR) memory; and wherein the RPU is further configured to map a portion of the entity memory into the PAS, enabling the processing cores to access the entity memory based on memory -mapped operations. Optionally, the mapping of entity memory such as HBM, LPDDR, or GDDR memory into PAS may include establishing memory windows with specific attributes optimized for the memory type. The RPU may handle differences in memory access granularity, bandwidth characteristics, and latency profiles between system memory and entity memory. The memory-mapped operations may be subject to caching policies and coherency protocols appropriate for cross-domain memory access.
[0166] Optionally, the RPU is further configured to provide access control by validating the physical addresses associated with the NVLink-based protocol against permitted address ranges for the entity, and blocking NVLink -based traffic targeting prohibited address ranges. Optionally, the permitted address ranges may be configured utilizing secure configuration registers or loaded from trusted firmware during system initialization. The RPU may support multiple access control contexts for different operational mode s or security domains. The blocking of prohibited traffic may generate error responses conforming to NVLink error reporting logic and may trigger security event logging.
[0167] Optionally, the RPU is further configured to evaluate transaction attributes associated with the NVLink-based protocol, including source identifiers and access types, and to apply security policies to allow or deny traffic based on preconfigured security rules. Optionally, the security policies may consider combinations of transaction attributes including source device identification, vendor-defined commands or fields, transaction type, address range, and temporal factors. The RPU may provide role-based access control wherein different entities have different access privileges. The security rules may be updateable utilizing authenticated channels and may support both static and dynamic security policy enforcement.
[0168] Optionally, the RPU is further configured to detect access patterns in NVLink-based traffic from the entity, and generates prefetch requests based on predicted future accesses; and wherein the prefetch requests are routed via the coherent interconnect and the memory controllers. Optionally, the access pattern detection may utilize algorithms such as stride detection, stream buffers, or correlation-based prediction algorithms. The RPU may maintain pattern history tables to track access behaviors and may adapt prefetching aggressiveness based on prefetch accuracy metrics. The prefetch requests may be tagged with lower priority to avoid interfering with demand requests and may be cancelled if subsequent access patterns diverge from predictions.
[0169] Optionally, the RPU is further configured to coalesce coherent interconnect transactions targeting contiguous or nearby addresses into fewer NVLink-based transactions; whereby the coalescing improves memory bandwidth utilization. Optionally, the request coalescing may consider factors including address proximity, request types, and timing windows when determining which transactions to combine. The RPU may include write combining buffers for write transactions and may support read coalescing for sequential read patterns. In one example, coherent interconnects may use up to 64 -byte transfers, that may reflect a nominal cacheline size utilized by the coherent interconnect, whereas NVLink may use larger transfers up to 256 bytes, making coalescing beneficial for bandwidth efficiency.
[0170] Optionally, the NVLink-based interface is configured to support virtual channels, and the RPU is further configured to map the virtual channels to quality -of-service (QoS) attributes in a protocol utilized by the coherent interconnect. Optionally, the virtual channel to QoS mapping may enable differentiated service levels for different traffic classes, such as bulk data transfers versus latency -sensitive communications. The RPU may include programmable mapping tables to allow flexible QoS policy configuration. The mapping may consider both NVLink virtual channel priorities and coherent interconnect QoS mechanisms to maintain end-to-end service level objectives.
[0171] Optionally, the memory comprises dynamic random-access memory (DRAM), and the entity comprises a graphics processing unit (GPU) or an accelerator coupled to the apparatus via the NVLink-based interface; and wherein the RPU enables the entity to access the DRAM with cache-line granularity. An entity, such as a GPU or an accelerator, may utilize the NVLink interface for memory access to memory resources attached to the processor. Optionally, when the entity is coupled through an NVLink switch, the RPU may handle switch-specific routinginformation and may support multiple entities sharing the NVLink interface through switch-based connectivity. The GPU or accelerator entity may utilize the NVLink interface for high-bandwidth memory access patterns characteristic of parallel computing workloads. The RPU may optimize translations for the specific access patterns and bandwidth requirements of GPU or accelerator workloads.
[0172] In one embodiment, a method for enabling an entity to access memory via an NVLink -based interface, comprising: operating a processor comprising processing cores, memory management units (MMUs), and coherent caches; wherein the processing cores respond to snoop requests that utilize physical addresses within a physical address space (PAS), and the MMUs translate virtual addresses to physical addresses within the PAS; communicating, via a coherent interconnect, between the processing cores and memory controllers that communicate with memory channels coupled to memory having a capacity of at least 64GB; executing, by the processing cores, an operating system (OS) that accesses the memory utilizing the physical addresses within the PAS; communicating according to an NVLink-based protocol with the entity via an NVLink-based interface; and translating physical addresses associated with the NVLink-based protocol to physical addresses within the PAS.
[0173] Optionally, the method further comprises translating from non-address fields conforming to the NVLink-based protocol message formats to corresponding fields conforming to message formats of a protocol utilized by the coherent interconnect; and wherein the translating of the physical addresses is performed by a resource provisioning unit (RPU) coupled between the NVLink-based interface and the coherent interconnect. Optionally, the protocol utilized by the coherent interconnect is based on Coherent Hub Interface (CHI-based protocol); and wherein the translating between non-address fields comprises translating NVLink-based protocol read commands to CHI -based protocol opcodes or commands comprising ReadOnce or ReadShared. Optionally, the method further includes translating CHI response opcodes to NVLink response opcodes, such as translating CHI responses carrying CompData to NVLink responses.
[0174] Optionally, the protocol utilized by the coherent interconnect is based on an Intel Coherent Processor Interconnect Protocol (ICPIP -based protocol) for scalable multiprocessors with a shared physical address space; and wherein the translating between non-address fields comprises translating NVLink-based protocol memory access commands to ICPIP -based protocol requests while maintaining coherency state tracking between domain of the NVLink-based protocol and domain of the ICPIP -based protocol. Optionally, the protocol utilized by the coherent interconnect is based on Infinity Fabric (IF -based); and wherein the translating between non-address fields comprises translating NVLink-based commands to IF -based commands while preserving memory ordering required by the entity. Optionally, the method further comprises translating NVLink-based commands to commands associated with a protocol utilized by the coherent interconnect, based on a mapping between NVLink-based transaction types and corresponding transaction types of the protocol utilized by the coherent interconnect. It is noted that in the context of such embodiments, NVLink-based commands and NVLink-based encodings may be used interchangeably.
[0175] Optionally, the translating of the physical addresses comprises utilizing an intermediate protocol selected from Peripheral Component Interconnect Express (PCIe) or Compute Express Link (CXL) as an intermediate stage between the NVLink-based protocol and a protocol utilized by the coherent interconnect. Optionally, the method further comprises translating transaction identifiers utilized by the NVLink-based protocol to transaction identifiers utilized by the coherent interconnect, maintaining a transaction tracking structure to monitor outstanding transactions from the entity, allocating coherent interconnect transaction identifiers for RPU -initiated transactions, and releasing identifiers upon transaction completion. Optionally, the method further comprises validating the physical addresses associated with the NVLink-based protocol against permitted address ranges for the entity, and blocking NVLink-based traffic targeting prohibited address ranges; and further comprising evaluating NVLink-based traffic attributes including source identifiers and access types, and applying security policies to allow or deny traffic based on preconfigured security rules. Optionally, the method further comprises detecting access patterns in NVLink-based traffic from the entity, and generating prefetch requests based on predicted future accesses, wherein the prefetch requests are routed via the coherent interconnect and the memory controllers.
[0176] In one embodiment, a system comprising: a host processor; a memory having a capacity of at least 64GB;a coherent interconnect architecture coupling processing elements to the memory, wherein the processing elements utilize a local physical address space to access the memory; and a resource provisioning unit (RPU) configured to translate physical addresses associated with an NVLink-based protocol, utilized by an entity coupled to the RPU via an NVLink-based interface, to physical addresses within the local physical address space; whereby the translate of the physical addresses enables the entity to utilize the memory as disaggregated memory accessed via the NVLink-based interface and the memory controllers.
[0177] FIG. 7A illustrates one embodiment of a system that may function as an NVLink memory switch appliance or an NVLink memory pool, and may include an MxPU, CPU, accelerator, or a memory switch ASIC, that is coupled to two entities denoted as Entity.1 / GPU.1 and Entity.2 / GPU.2. The MxPU includes processing cores and memory controllers coupled to a coherent interconnect that may be based on CHI. The MxPU utilizes translations, performed by the RPUs, between NVLink-based interfaces and an MxPU’s coherent interconnect. The first RPU (RPU.l) may enable Entity.1 / GPU.1 to access resources mapped to a physical address space utilized by the MxPU’s coherent interconnect, wherein the access is via the first NVLink interface and the MxPU’s coherent interconnect. Examples of resources mapped to the physical address space utilized by the MxPU’s coherent interconnect include DRAM or other memory resources of the MxPU. Correspondingly, the second RPU (RPU.2) may enable Entity.2 / GPU.2 to access, via the second NVLink interface and the MxPU’s coherent interconnect, resources mapped to a physical address space utilized by the MxPU’s coherent interconnect, such as memory resources of the MxPU.
[0178] FIG. 7B illustrates one embodiment of a TFD depicting a multi-entity memory access scenario wherein first and second entities / GPUs access memory mapped to one or more physical address spaces utilized by the coherent interconnect (CohlnterMappedMemory), through NVLink to ARM CHI translations. Entity.1 / GPU.1 initiates a first NVLink Request: Read with SourceID(a.1) to identify the source GPU, DestinationID(b.1) to identify the destination GPU, and Address(AS.2.1) representing an NVLink network address from a second physical address space. RPU.l translates the first NVLink Request to ARM CHI REQ carrying Opcode(ReadOnce), and Addr(AS.l.l) from a first physical address space utilized by the coherent interconnect. Concurrently or sequentially, Entity.2 / GPU.2 may initiate a second NVLink Request: Read with SourceID(a.2), DestinationID(b.2), and Address(AS.3.1) representing an NVLink network address optionally from a third physical address space or from the second physical address space. RPU.2 translates the second NVLink Request to ARM CHI REQ carrying Opcode(ReadOnce) and Addr(AS.1.2) from the first physical address space utilized by the coherent interconnect.
[0179] Both transactions flow through the coherent interconnect to one or more home nodes, which may send respective ARM CHI REQ messages to one or more memory controllers with Opcode(ReadNoSnp) and the addresses Addr(AS.l.l) and Addr(AS.1.2), respectively. The memory controller(s) retrieve the requested data from the CohlnterMappedMemory and send first and second ARM CHI RD AT messages with Opcode(CompData) carrying *Data.l* and *Data.2*, representing the data retrieved from the addresses AS.1.1 and AS.1.2, respectively. RPU.l translates the first ARM CHI RD AT message to NVLink Response with SourceID(b.l), DestinationID(a.1), and *Data.l* for Entity.1 / GPU.1. RPU.2 translates the second ARM CHI RD AT message to NVLink Response with SourceID(b.2), DestinationID(a.2), and *Data.2* for Entity.2 / GPU.2. The embodiment demonstrates how entities / GPUs may share access to the same CohlnterMappedMemory through different RPUs that translate between NVLink and ARM CHI, including physical address translations. Alternatively, the embodiment may be viewed as two separate NVLink transactions that utilize the same coherent interconnect infrastructure to access CohlnterMappedMemory, wherein the GPU entities may access the CohlnterMappedMemory via a shared or separate address spaces that are translated to the shared coherent interconnect physical address space. Still alternatively, the response and read data paths may be implemented according to other designs, such as wherein the memory controller(s) may send the data to the home node(s) that send it to the respective RPUs, or the home node(s) send responses to the RPUs while the memory controller(s) send the data directly to the RPUs.
[0180] Depending on system characteristics, such as implementation choices and platform configurations, different physical addresses, such as (AS.1.1) and (AS.1.2), within a physical address space utilized by the coherent interconnect, may be typically partitioned, such as via hashing or interleaving schemes, across a set of home nodes.Such partitioning is typically performed in order to reduce bottleneck effects in the system and spread the load of transaction processing across home nodes of the coherent interconnect, and may result in mapping the different physical addresses, such as (AS.1.1) and (AS.1.2), to the same home node, or to different home nodes. Similarly, different physical addresses may be associated with one memory controller, or with different memory controllers, such as according to a separate mapping scheme, which may be different from the mapping scheme utilized for selecting a home node for processing the request. Alternatively, other embodiments may co -locate the home node function with a specific memory controller, utilizing a unified mapping scheme that selects both a home node and a memory controller.
[0181] FIG. 8 illustrates one embodiment of a system comprising a processor (such as an MxPU) comprising processing cores, LLC, a CXL RP, and memory controllers coupled via memory channels to memory, such as DRAM. The CXL RP may be coupled to an on-chip coherent interconnect, such as a CHI ring or mesh interconnect, via a Ring-to-CXL (R2CXL) interconnect interface that may communicate with the coherent interconnect according to a protocol utilized by the coherent interconnect, such as ARM CHI, Intel IDI, Intel UPI, or AMD Infinity Fabric. An RPU, which may be included in the MxPU, performs physical address translations that may enable an entity such as a GPU to access the memory. The MxPU may expose to the entity, optionally via the RPU, an NVLink interface that may communicate with the entity according to NVLink. The RPU may further perform translations, such as from NVLink to a protocol utilized by the coherent interconnect, wherein the RPU may utilize an intermediate protocol, such as CXL (e.g., CXL.cache), for providing the translations. The RPU may expose to the processor, via a CXL RP that may be included in the RPU, a CXL device utilizing a CXL Endpoint (CXL EP), such as a Type-1 CXL device, or a Type-2 CXL device. The R2CXL interconnect interface, that may reside in the RPU, may couple the CXL RP to the coherent interconnect and complement the translation path from NVLink, via the intermediate protocol, such as CXL, to traffic based on the protocol utilized by the coherent interconnect. In some embodiments, the RPU, the NVLink interface, and the CXL device (e.g., CXL EP) may be implemented in a chiplet, such as an NVLink chiplet, or NVLink Fusion, inside an IC package of an MxPU, whereas in other embodiments, they may be implemented as functional blocks on the same die with the CXL RP of the processor, or split between silicon dies or chiplets inside the IC package of the MxPU.
[0182] FIG. 9 illustrates one embodiment of a TFD demonstrating an NVLink read request received from an entity (such as a consumer, GPU, accelerator, or a switch), wherein the RPU may translate a physical address (AS.2.1) carried in the NVLink request, to a physical address (AS.1.1) utilized for accessing the memory. The RPU may perform further translations, such as protocol translations, from NVLink traffic to traffic conforming to a protocol utilized by the processor’s coherent interconnect, possibly utilizing an intermediate protocol such as CXL.cache, wherein the RPU may perform further translations, such as opcode translations and Tag translations, e.g., of transaction Tags, such as translating from NVLink Tags to CXL.cache CQIDs. The CXL.cache D2H request, carrying the translated address (AS.1.1), is sent to the CXL RP for further processing and fetching of the requested data, such as from an LLC over the coherent interconnect, or from DRAM via the memory channels. The data may then return over the coherent interconnect to the RPU, wherein the RPU provides an NVLink response to the requesting entity.
[0183] Modem datacenters face unprecedented computational demands driven by generative Al (GenAI), Large Language Models (LLMs), distributed machine learning training, and real-time analytics workloads that require massive memory resources distributed across racks and pods within the datacenter fabric. These evolving workloads increasingly necessitate flexible memory disaggregation architectures that can operate at multiple scales, from memory sharing between adjacent servers within a rack to memory pooling across pods in large-scale datacenter deployments. The emergence of memory -intensive applications, such as those involving large models and distributed training systems, requires memory systems that can accommodate varying scales of deployment.
[0184] Compute Express Link (CXL) is establishing itself as a standard for memory expansion and pooling within datacenter environments, providing protocols including CXL.io, CXL.cache, and CXL.mem that enable coherent and non-coherent memory access between processors and memory devices. CXL facilitates memory disaggregation by allowing hosts to access remote memory resources through standardized interfaces, supporting both CXL type -2 andtype-3 devices for various memory expansion scenarios. Concurrently, IEEE 802.3 physical layer specifications provide robust, high-bandwidth physical medium attachment (PMA) capabilities that form the backbone of many datacenter networking infrastructures, from top-of-rack switches to spine-leaf architectures.
[0185] However, current CXL implementations rely on PCIe physical layers that may limit deployment flexibility within datacenters. While CXL switches enable memory pooling at the rack level using PCIe -based connections, extending CXL across the datacenter fabric to inter-pod distances or leveraging existing datacenter network infrastructure presents challenges. The separation between CXL interconnects and IEEE 802.3 -based datacenter networks may require organizations to deploy and maintain multiple infrastructure layers for memory disaggregation and networking. Additionally, the lack of mechanisms to bridge between different physical address spaces while maintaining CXL limits the scalability and flexibility of memory disaggregation solutions.
[0186] These limitations become apparent in scenarios requiring dynamic memory allocation across datacenter resources, distributed Al training that spans multiple racks or pods, flexible memory tiering within the datacenter, or cloud-native architectures that demand elastic memory provisioning. Solutions are needed that can extend CXL semantics over both short intra-rack distances and longer inter-pod connections using IEEE 802.3-based physical layers while maintaining the protocol's coherency and performance characteristics.
[0187] Some of the disclosed embodiments introduce novel architectural solutions that enable CXL semantics to traverse datacenter networks that utilize physical layers based on IEEE 802.3 PMA, facilitating memory disaggregation from rack-scale to pod-scale deployments. These embodiments leverage RPUs to translate between CXL data, optionally encapsulated, transmitted over physical layers based on IEEE 802.3 PMA, and CXL requests, enabling external entities to access memory resources across different physical address spaces whether within the same rack or across datacenter pods. By bridging CXL protocols with capabilities of physical layers based on IEEE 802.3 PMA, the embodiments provide memory pooling infrastructure that leverages datacenter network fabric for both intra-rack and inter-pod memory sharing. The embodiments optionally support various deployment scenarios including rack-level memory pooling, pod-scale distributed AI / ML training, datacenter-wide memory tiering, and / or cloud-native memory-as-a-service architectures.
[0188] In one embodiment, an apparatus comprises an IC package comprising processing cores comprising instruction caches, wherein the processing cores are coupled via a coherent interconnect to a memory controller, and are configured to respond to snoop requests that utilize physical addresses within a first physical address space. The apparatus further comprises a memory management unit (MMU), coupled to the processing cores, configured to translate virtual addresses to physical addresses within the first physical address space. Additionally, the apparatus comprises memory channels capable of supporting memory having a capacity of at least 64GB. The apparatus also comprises a physical layer, based on IEEE 802.3 physical medium attachment (PMA), configured to receive transmissions comprising data indicative of CXL opcodes and physical addresses within a second physical address space. Furthermore, the apparatus comprises a resource provisioning unit (RPU) configured to translate the data to CXL requests, whereby the translate of the data enables an entity external to the apparatus to read the memory via the physical layer based on IEEE 802.3 PMA, the memory controller, and the memory channels.
[0189] In another embodiment, a method comprises operating a processor comprising processing cores and instruction caches, wherein the processing cores communicate via a coherent interconnect with a memory controller, and respond to snoop requests that utilize physical addresses within a first physical address space. The method further comprises translating virtual addresses to physical addresses for the processing cores. Additionally, the method comprises communicating, via memory channels, with memory having a capacity of at least 64GB. The method also comprises receiving, via a physical layer based on IEEE 802.3 physical medium attachment (PMA), transmissions comprising data indicative of CXL opcodes and physical addresses within a second physical address space. Furthermore, the method comprises translating, by a resource provisioning unit (RPU), the data to CXL requests, whereby the translating enables an entity external to the processor to read the memory via the physical layer based on IEEE 802.3 PMA, the memory controller, and the memory channels.
[0190] In yet another embodiment, a system comprises an IC package comprising processing cores coupled via acoherent interconnect to memory controllers, wherein the processing cores respond to snoop requests utilizing physical addresses within a host physical address space. The system further comprises memory management units (MMUs) configured to translate virtual addresses to physical addresses within the host physical address space. Additionally, the system comprises memory channels coupled to memory having a capacity of at least 64GB accessible via the memory controllers. The system also comprises physical layers based on IEEE 802.3 physical medium attachment (PMA), configured to communicate with respective external entities, wherein the physical layers receive transmissions comprising data indicative of memory access requests comprising physical addresses. Furthermore, the system comprises at least one resource provisioning unit (RPU) configured to translate between (i) physical addresses associated with the transmissions and (ii) physical addresses within the host physical address space, whereby the translate enables the external entities to access the memory via the physical layers based on IEEE 802.3 PMA, the memory controllers, and the memory channels.
[0191] Datacenter workloads demand flexible memory architectures spanning from rack-level to pod-scale deployments. Embodiments herein disclose systems enabling CXL semantics over physical layers based on IEEE 802.3 PMA, facilitating memory disaggregation across datacenter fabric infrastructures. The embodiments comprise processing cores with coherent interconnects, MMUs for address translation, and memory channels supporting substantial memory capacities. Resource Provisioning Units (RPUs) translate between CXL data, optionally encapsulated, transmitted via physical layers based on IEEE 802.3 PMA, and CXL requests, enabling external entities to access memory across different physical address spaces. This architecture provides memory pooling using datacenter network infrastructure, supporting intra-rack memory sharing, inter-pod memory access, distributed Al training across datacenter resources, and elastic memory provisioning for cloud-native applications, overcoming physical layer limitations of traditional CXL implementations while maintaining protocol coherency suitable for GenAI, LLM inference, and HPC workloads.
[0192] In one embodiment, an apparatus comprising: an integrated circuit package (IC package) comprising processing cores comprising instruction caches; wherein the processing cores are coupled via a coherent interconnect to a memory controller, and are configured to respond to snoop requests that utilize physical addresses within a first physical address space; a memory management unit (MMU), coupled to the processing cores, configured to translate virtual addresses to physical addresses within the first physical address space; memory channels capable of supporting memory having a capacity of at least 64GB; a physical layer, based on IEEE 802.3 physical medium attachment (PMA), configured to receive transmissions comprising data indicative of Compute Express Link (CXL) opcodes and physical addresses within a second physical address space; and a resource provisioning unit (RPU) configured to translate the data to CXL requests; whereby the translate of the data enables an entity external to the apparatus to read the memory via the physical layer based on IEEE 802.3 PMA, the memory controller, and the memory channels.
[0193] Optionally, the translate of the data to CXL requests comprises translate the physical addresses within the second physical address space to the physical addresses within the first physical address space. Optionally, the physical layer based on IEEE 802.3 PMA comprises a UALink physical layer. Optionally, the physical layer based on IEEE 802.3 PMA comprises an NVLink physical layer. Optionally, the physical layer based on IEEE 802.3 PMA is configured to receive and transmit Ethernet frames. Optionally, the RPU is included in the IC package or is included in a second IC package coupled to the IC package. Optionally, the memory comprises dynamic random-access memory (DRAM), and the RPU coupled to at least one of a CXL device, a CXL endpoint, or a CXL port. Optionally, the CXL port comprises a switch port or a root port, and the entity comprises a GPU, a network device, or a storage device. Optionally, the CXL device comprises a Global Fabric -Attached Memory (G-FAM) Device (GFD), and the RPU communicates with the GFD according to at least one of CXL.mem or CXL.io. Optionally, the second physical address space is identical to the first physical address space, and the RPU is further configured to enable the entity to read memory-mapped I / O (MMIO) registers coupled to the apparatus. Optionally, the CXL opcode comprises a UIOMRd Transaction Layer Packet (TLP) type. Optionally, the second physical address space is a subset of the first physical address space, and the RPU is further configured to block the entity from reading at least one predetermined address region within the first physical address space. Optionally, the data indicative of CXL opcodes and physicaladdresses is encapsulated within messages conforming to a carrier protocol; and wherein the RPU is further configured to extract the data indicative of CXL opcodes and physical addresses within the messages conforming to the carrier protocol and to translate the extracted data to the CXL requests. Optionally, the physical layer based on IEEE 802.3 PMA is configured to receive the messages conforming to the carrier protocol; and wherein the RPU is further configured to translate the physical addresses within the second physical address space to physical addresses within the first physical address space when processing the CXL requests. Optionally, the RPU is further configured to encapsulate data from CXL responses into response messages conforming to the carrier protocol for transmission via the physical layer based on IEEE 802.3 PMA.
[0194] Optionally, the carrier protocol is based on Ethernet or based on IEEE 802.3; and wherein the data indicative of CXL opcodes and physical addresses is encapsulated within Ethernet frames or IEEE 802.3 frames, respectively. Optionally, the carrier protocol is based on Ultra Ethernet Transport (UET) protocol; and wherein the data indicative of CXL opcodes and physical addresses is encapsulated within Link Layer Retry eligible frames (LLR -eligible frames). Optionally, the carrier protocol is based on Scale Up Ethernet (SUE); and wherein the data indicative of CXL opcodes and physical addresses is encapsulated within an SUE -based Protocol Data Unit (PDU). Examples of SUE-based PDU may include SUE PDU, SUE Lite PDU, or PDUs based on future revisions of SUE. Optionally, the CXL requests comprise CXL.mem Master-to-Subordinate Request (M2S Req) messages; and wherein the RPU is further configured to receive CXL responses comprising CXL.mem Subordinate -to-Master Data Response (S2M DRS) messages. Optionally, the RPU is further configured to: translate Tags, associated with the data, from a first Tag space to a second Tag space; and wherein the first Tag space is utilized by the entity external to the apparatus and the second Tag space is utilized by a host coupled to the apparatus. Optionally, the RPU further translates Tags in CXL responses associated with the second Tag space to Tags associated with the first Tag space before encapsulation into the carrier protocol.
[0195] In one embodiment, a method comprising: operating a processor comprising processing cores and instruction caches; and wherein the processing cores communicate via a coherent interconnect with a memory controller, and respond to snoop requests that utilize physical addresses within a first physical address space; translating virtual addresses to physical addresses for the processing cores; communicating, via memory channels, with memory having a capacity ofat least 64GB; receiving, via a physical layer based on IEEE 802.3 physical medium attachment (PMA), transmissions comprising data indicative of Compute Express Link (CXL) opcodes and physical addresses within a second physical address space; and translating, by a resource provisioning unit (RPU), the data to CXL requests; whereby the translating enables an entity external to the processor to read the memory via the physical layer based on IEEE 802.3 PMA, the memory controller, and the memory channels.
[0196] Optionally, the translating of the data to CXL requests comprises translating the physical addresses within the second physical address space to the physical addresses within the first physical address space. Optionally, the second physical address space is identical to the first physical address space, and further comprising enabling, by the RPU, the entity to read memory-mapped I / O (MMIO) registers coupled to the processor. Optionally, the second physical address space is a subset of the first physical address space, and further comprising blocking, by the RPU, the entity from reading at least one predetermined address region within the first physical address space. Optionally, the data indicative of CXL opcodes and physical addresses is encapsulated within messages conforming to a carrier protocol, and further comprising extracting, by the RPU, the data indicative of CXL opcodes and physical addresses from the messages conforming to the carrier protocol, and translating the extracted data to the CXL requests. Optionally, the method further comprises receiving, via the physical layer based on IEEE 802.3 PMA, the messages conforming to the carrier protocol, translating, by the RPU, the physical addresses within the second physical address space to physical addresses within the first physical address space, and encapsulating, by the RPU, data from CXL responses into second messages conforming to the carrier protocol for transmission via the physical layer based on IEEE 802.3 PMA. Optionally, the method further comprises translating, by the RPU, Tags from a first Tag space to a second Tag space, wherein the first Tag space is utilized by the entity external to the processor, and the second Tag space is utilized by a host coupled to the processor; and further comprising translating, by the RPU, Tags associatedwith the second Tag space to Tags associated with the first Tag space before encapsulation into the carrier protocol.
[0197] In one embodiment, a system comprising: an integrated circuit package (IC package) comprising processing cores coupled via a coherent interconnect to memory controllers, wherein the processing cores respond to snoop requests utilizing physical addresses within a host physical address space; memory management units (MMUs) configured to translate virtual addresses to physical addresses within the host physical address space; memory channels coupled to memory, having a capacity of at least 64GB, accessible via the memory controllers; physical layers based on IEEE 802.3 physical medium attachment (PMA), configured to communicate with respective external entities, wherein the physical layers receive transmissions comprising data indicative of memory access requests comprising physical addresses; and at least one resource provisioning unit (RPU) configured to translate between (i) physical addresses associated with the transmissions and (ii) physical addresses within the host physical address space; whereby the translate enables the external entities to access the memory via the physical layers based on IEEE 802.3 PMA, the memory controllers, and the memory channels.
[0198] Optionally, the memory access requests conform to at least one protocol selected from Ultra Accelerator Link (UALink) requests, UALink Protocol Level Interface (UPLI) requests, or NVLink requests; and wherein at least one of the physical layers based on IEEE 802.3 PMA comprises a UALink physical layer, an NVLink physical layer, or an Ethernet physical layer operating at 100 Gbps or higher. Optionally, the at least one RPU comprises multiple RPUs distributed across the IC package or across multiple IC packages; and wherein the RPUs are further configured to maintain different address translation tables for different external entities, and to enforce access control policies defining permitted address ranges for corresponding external entities, thereby creating isolated security domains for memory access while sharing the same physical memory resources.
[0199] FIG. 10A illustrates one embodiment of a system wherein an entity is coupled via a Physical Layer based on IEEE 802.3 PMA to an RPU that includes or is coupled to a CXL device that is coupled to an ARM architecture processor through one or more CCG nodes and possibly one or more RN-D nodes. The memory controller (MC) within the system is coupled to SN-F nodes and interfaces with DRAM through DDR PHYs and memory channels. The embodiment includes crosspoints (XP) that function as routing elements within the ARM mesh interconnect, examining packet identifiers to determine appropriate routing paths and managing traffic flow between sources and destinations within the mesh structure.
[0200] FIG. 10B illustrates one embodiment of a TFD that may be executed on the embodiment described in FIG.10A, demonstrating a first CXL.mem to a second CXL.mem translation by an RPU, followed by a requester, which may be the CCG or the CXL device from FIG. 10A, initiating a CHI allocating ReadShared request to a Home Node, with the Memory Controller serving as the Subordinate node. The diagram illustrates the optimization technique of combined response from subordinate, wherein the Home Node sends a ReadNoSnp request to the Memory Controller, which then returns data directly to the Requester using a CompData opcode, reducing message count and potentially improving transaction latency by eliminating the need for data to flow back through the Home Node.
[0201] FIG. 11A illustrates one embodiment of a system comprising a processor having multiple interfaces that may utilize a physical layer based on IEEE 802.3 PMA. A first RPU includes or is coupled to a CXL device that is coupled to both a CCG node for handling coherent CXL.mem and / or CXL.cache transactions and an RN-D node for handling non-coherent CXL.io transactions. The system may optionally couple the first RPU to the CCG over a CXS interface, providing a path for coherent communications. A second RPU includes or is coupled to a Root Port that is coupled to both a fully coherent request nodes (RN-F) and to a fully coherent home nodes (HN-F) that may be included within a gateway or a bridge node stmcture, enabling bidirectional coherent access wherein an external entity, such as a GPU or a storage device may read from the processor's DRAM through the RN-F node; and in the opposite direction, the processor cores may read from the GPU's HBM or from buffers in the storage device through the HN-F node.
[0202] FIG. 1 IB illustrates one embodiment of a system comprising an XPU or a CPU, which may be a custom CPU design, incorporating accelerator cores and multiple interfaces that may utilize a physical layer based on IEEE 802.3 PMA. A Global Fabric -Attached Memory (G-FAM) Device (GFD), utilized by a first RPU, may operate as aspecialized CXL device, and may support only CXL.mem transactions, allowing it to service the external requests through CCGs that are optimized for handling CXL.mem traffic, thereby simplifying the design by eliminating the need for separate CXL.io handling paths typically managed by RN-D or RN-I nodes. The system further includes an optional second RPU that includes or is coupled to a Root Port, coupled to a coherent interconnect via a CCG and an I / O-Coherent Request Node with DVM support (RN-D), wherein the RN-D may handle CXL.io or PCIe traffic. It is noted that a line in a mesh drawing may denote more than one port, interface, or link. For example, a single line connecting a CCG to an XP may represent two ports, such as one port for a Request Agent (RA) proxy and another port for a Home Agent (HA) proxy.
[0203] FIG. 12 illustrates one embodiment of a multi-host memory pooling or sharing utilizing a switch-based topology with physical layers based on IEEE 802.3 PMA. The system includes at least two distinct paths for hosts to access memory resources through the same memory channels. The first path shows an entity, such as an xPU, with a host designated as A.1 that connects via CXL, carried on a carrier protocol interface, which transmits over a physical layer based on IEEE 802.3 PMA (which may utilize Ethernet PHY) to an RPU. RPUs are coupled to or included in a switch, which may be a PBR switch, enabling scalable connectivity. The switch includes downstream ports (DSPs) that are coupled to CXL Type-3 devices containing memory controllers and associated memory channels coupled to memories (denoted as A.2 and B.2), allowing for distributed memory resources. The second path shows the switch's upstream port (USP) coupled via a root port (RP) to a host system (denoted as B.l) that includes a coherent interconnect, processing cores with MMU functionality, and instruction caches. This dual-path coupling enables the host B .1 to access memory through its connection via the switch's U SP and RP, while simultaneously allowing external entities like the xPU (host A.l) to access the same memory resources through the RPU and switch infrastructure. A GPU may also connect to the switch through its own RPU and PHY based on IEEE 802.3 PMA, demonstrating the system's capability to integrate heterogeneous computing elements.
[0204] The explosive growth of artificial intelligence workloads, such as Large Language Models (LLMs) and Generative Al (GenAI) applications, has accelerated the adoption of heterogeneous computing architectures that combine diverse processing elements including CPUs, GPUs, and domain-specific accelerators. These computeintensive workloads demand unprecedented levels of memory bandwidth and capacity, driving the deployment of ARM-based server architectures in datacenters and edge computing environments. ARM-based processors have gained traction in cloud infrastructure and high-performance computing due to their power efficiency and scalability, with deployments ranging from hyperscale datacenters to edge Al inference systems. Using ARM-based platforms for Al training, real-time inference, and HPC workloads require efficient integration with accelerators, such as GPUs that handle the computational demands of neural network processing and parallel computing tasks.
[0205] ARM's Advanced Microcontroller Bus Architecture (AMBA) Coherent Hub Interface (CHI) has emerged as a standard for coherent interconnects in ARM -based systems, enabling cache-coherent communication between processing cores, memory controllers, and system components. CHI provides mechanisms for maintaining cache coherency across multiple processors and accelerators while supporting high-bandwidth memory access through its distributed architecture. Concurrently, NVLink has established itself as a high-bandwidth interconnect technology for GPU-to-GPU and GPU-to-CPU communication to achieve data rates that exceed traditional PCIe capabilities. NVLink enables direct memory access between connected devices, supporting the memory pooling and sharing requirements of modem AI / ML workloads.
[0206] However, existing architectures face limitations when attempting to integrate NVLink -connected accelerators with ARM CHI -based coherent interconnects. The protocol differences between NVLink and CHI create barriers to efficient memory sharing and resource access. While NVLink provides high-bandwidth connectivity between GPUs and supported processors, the lack of native interoperability with CHI-based systems prevents accelerators from directly accessing the coherent memory fabric of ARM-based servers. This limitation is apparent in edge computing scenarios where powerful ARM-based processors require access to GPU memory resources that exceed the local memory capacity of individual accelerators, and in datacenter deployments where memory disaggregation could improve resource utilization across heterogeneous compute elements. There is a need forsolutions that translate between NVLink and CHI -based interconnects, allowing accelerators to efficiently access memory resources within ARM-based coherent domains while maintaining the high-bandwidth advantages of both interconnect technologies.
[0207] Some of the disclosed embodiments introduce novel system-level architectural solutions leveraging RPUs to enable NVLink-connected accelerators to access memory and resources within ARM CHI-based coherent interconnect fabrics. These embodiments provide hardware-accelerated translation between NVLink and CHI, enabling GPUs and other NVLink-capable accelerators to efficiently access memory resources managed by ARMbased processors through coherent interconnects. By implementing RPUs that bridge the protocol gap between NVLink and CHI-based messaging, the embodiments enable memory sharing across heterogeneous compute environments in datacenter and / or edge deployments. The embodiments address memory disaggregation challenges for AI / ML training and inference workloads, enabling accelerators to access memory pools beyond their local capacity while potentially maintaining cache coherency through CHI. Some embodiments optionally support multiple NVLink-connected devices simultaneously through scalable RPU configurations, enabling flexible resource allocation for GenAI inference, distributed training, and HPC applications on ARM -based platforms.
[0208] In one embodiment, an apparatus comprises a coherent interconnect that utilizes a CHI -based protocol, comprising an interconnect component configured to receive CHI -based messages. The apparatus further comprises processing cores coupled via the coherent interconnect to memory controllers coupled to memory channels capable of supporting memory having a capacity of at least 64GB. Additionally, the apparatus comprises a resource provisioning unit (RPU) comprising an NVLink interface and a CHI interface, wherein the NVLink interface utilizes differential pairs and is capable of communicating according to an NVLink -based protocol with an entity external to the apparatus, and wherein the CHI interface is coupled to the interconnect component. The RPU is configured to translate between messages conforming to the NVLink -based protocol and messages conforming to the CHI -based protocol to enable the entity to access resources via the NVLink interface and the coherent interconnect.
[0209] In another embodiment, a method comprises operating a coherent interconnect that utilizes a CHI -based protocol, comprising an interconnect component that receives CHI -based messages. The method further comprises communicating, via the coherent interconnect, between processing cores and memory controllers that communicate with memory channels coupled memory having a capacity of at least 64GB. The method includes operating a resource provisioning unit (RPU) comprising an NVLink interface and a CHI interface, wherein the NVLink interface utilizes differential pairs and communicates according to an NVLink-based protocol with an entity external to the RPU, and wherein the CHI interface communicates with the interconnect component. Additionally, the method comprises translating, by the RPU, between messages conforming to the NVLink-based protocol and messages conforming to the CHI-based protocol to enable the entity to access resources via the NVLink interface and the coherent interconnect.
[0210] In yet another embodiment, a system comprises a coherent interconnect based on CHI, comprising interconnect components configured to route CHI -based messages. The system further comprises processing cores coupled via the coherent interconnect to memory controllers coupled to memory channels coupled to memory having a capacity of at least 64GB. Additionally, the system comprises RPUs comprising external interfaces and CHI interfaces, wherein at least one of the external interfaces comprises an NVLink interface utilizing differential pairs for communication according to an NVLink-based protocol with one or more external entities, and wherein the CHI interfaces are coupled to the interconnect components. The RPUs are configured to translate between protocols utilized by the external interfaces and the CHI -based protocol, whereby the translate enables the external entities to access system resources via the external interfaces and the coherent interconnect.
[0211] Modem AI / ML workloads demand efficient integration of GPUs and accelerators with ARM-based server architectures in datacenters and edge computing environments. Embodiments herein disclose systems incorporating RPUs that enable NVLink-connected accelerators to access memory resources within ARM CHI-based coherent interconnect fabrics. One embodiment comprises a CHI -based coherent interconnect with interconnect components routing CHI messages between processing cores and memory controllers supporting substantial memory capacities. The RPU bridges between NVLink and CHI, while translating between NVLink and CHI messaging, enabling GPUsand accelerators to access system memory through the coherent fabric. Multiple RPUs optionally support scalable configurations with multiple NVLink-connected devices accessing shared memory resources. The embodiments address memory disaggregation challenges for GenAI inference, LLM training, and distributed computing, enabling accelerators to leverage ARM-based system memory beyond local device capacity while potentially maintaining cache coherency, suitable for heterogeneous computing deployments.
[0212] In one embodiment, an apparatus comprising: a coherent interconnect based on Coherent Hub Interface (CHI) protocol (CHI-based protocol), comprising an interconnect component configured to receive CHI-based messages; processing cores coupled via the coherent interconnect to memory controllers coupled to memory channels capable of supporting memory having a capacity of at least 64GB; a resource provisioning unit (RPU) comprising an NVLink interface and a CHI interface; wherein the NVLink interface utilizes differential pairs and is capable of communicating according to an NVLink-based protocol with an entity external to the apparatus; wherein the CHI interface is coupled to the interconnect component; and wherein the RPU is configured to translate between messages conforming to the NVLink-based protocol and messages conforming to the CHI -based protocol to enable the entity to access resources via the NVLink interface and the coherent interconnect.
[0213] Optionally, the RPU is further configured to: translate first physical addresses associated with the NVLink-based protocol to second physical addresses associated with the CHI -based protocol, and translate NVLink command encodings to corresponding CHI opcodes. Optionally, the RPU may perform address translation from the NVLink domain to the CHI domain. The address translation may support different memory mapping schemes between the NVLink and CHI domains, while the command translation may preserve the intent of the transaction. For example, when translating an NVLink read request transaction, received from a GPU, to a CHI request transaction, targeting an xPU coherent interconnect, wherein the CHI transaction carries ReadOnce for obtaining a non-cacheable snapshot of the data, satisfying the intent of the I / O coherent NVLink read request. The RPU may preserve the ordering requirements of the original NVLink traffic within the CHI -based protocol framework. Optionally, the resources are selected from at least one of: registers within the apparatus, SRAM or HBM within the apparatus, at least some of the 64GB of memory, network devices coupled to the apparatus, or storage devices coupled to the apparatus. Optionally, the RPU further comprises a request node which does not include a hardware -coherent cache, and wherein the request node is configured to communicate with the interconnect component according to the CHI -based protocol. Optionally, the request node is coupled to the interconnect component and is further configured to expose registers accessible utilizing memory-mapped I / O (MMIO) operations, to enable the entity to detect at least one of: node type, node configuration, or connection topology based on register inspection. Optionally, the request node is configured to expose the registers via Advanced Microcontroller Bus Architecture (AMBA) Advanced Peripheral Bus (APB) interface, to enable the entity to read the registers via the NVLink interface.
[0214] Optionally, the request node comprises an I / O-Coherent Request Node (RN-I) or an I / O-Coherent Request Node with Distributed Virtual Memory (DVM) support (RN-D); and the RPU is configured to translate NVLink read requests to CHI read requests. Optionally, the integration with ARM mesh architecture allows the NVLink-coupled entity to participate in the broader system interconnect fabric, with interconnect components, such as crosspoints, providing routing decisions based on transaction addresses and types. The MMIO-accessible registers enable system firmware or diagnostic software to discover the structure of the coherent interconnect, the presence of request nodes and home nodes included in the RPU, verify correct node connections, detect NVLink translation capabilities in the RPU via additional register inspections, and configure operational parameters for the translation path.
[0215] Optionally, the RPU further comprises a home node which does not include a Point of Coherence (PoC) and is not capable of processing snoopable requests, and wherein the home node is configured to communicate with the interconnect component according to the CHI-based protocol. Optionally, the home node comprises a Noncoherent Home Node (HN-I), enabling the processing cores to access resources via the NVLink interface.
[0216] Optionally, the RPU further comprises a request node and a home node, the request node couples the NVLink interface to the interconnect component, and the home node couples the NVLink interface to a second interconnect component. Optionally, the RPU may implement routing decisions based on transaction types, directingmemory access transactions from the NVLink domain through a request node, such as an RN-I node, while receiving, from a home node, such as an HN-I node, transactions targeting the NVLink domain. The apparatus may enable entities communicating according to NVLink-based protocol to perform I / O-coherent accesses to resources within a CHI-based system through appropriate non-coherent or I / O-coherent nodes. A request node, such as an RN-D node, may receive DVM transactions and generate a subset of CHI transactions without maintaining a hardware-coherent cache. The home node, such as an HN-I node, may process a limited subset of request types and manage ordering between I / O requests targeting the I / O subsystem without maintaining coherency utilizing snooping. The RPU may perform protocol-specific translations including command mapping, address formatting, address translations, orchestration and tracking of transaction IDs, and transaction sequencing between the NVLink and CHI domains.
[0217] Optionally, the RPU further comprises an interconnect gateway configured to communicate with the interconnect component according to the CHI -based protocol, wherein the RPU is further configured to utilize a streaming interface protocol to enable connectivity between the NVLink interface and the coherent interconnect via the interconnect gateway. Optionally, the streaming interface protocol transports packets of an intermediate protocol; and wherein the RPU is further configured to translate between messages conforming to the intermediate protocol and messages conforming to the CHI -based protocol. Optionally, the intermedia protocol conforms to PCIe, and the RPU is further configured to translate a PCIe UIO memory read request utilizing a UIOMRd TLP type to a CHI REQ comprising ReadOnce.
[0218] Optionally, the streaming interface protocol is based on Advanced Microcontroller Bus Architecture (AMBA) Credited extensible Stream (CXS); and wherein the interconnect gateway provides credit -based flowcontrol and supports bi-directional connectivity between the NVLink interface and the coherent interconnect. Optionally, the interconnect gateway comprises CMN multi-Chip Gateway (CCG) comprising a link agent that supports the streaming interface protocol, providing flit packing and unpacking, end-to-end data integrity, and a flitretry mechanism for reliability, availability and serviceability (RAS) containment when data corruption is detected. Optionally, the interconnect gateway comprises at least one of Coherent Multi -Chip Link (CML) or Cache Coherent Interconnect for Accelerators (CCIX) Gateway (CXG); and wherein the gateway is configured to utilize a 32 -bit cyclic -redundancy check (CRC-32) to protect transactions conforming to the streaming interface protocol. Optionally, the RPU comprises a request agent (RA) proxy configured to communicate with the interconnect component according to the CHI-based protocol, enabling the entity to access, via the NVLink interface, resources coupled to the coherent interconnect. Optionally, the RPU comprises a home agent (HA) proxy configured to communicate with the interconnect component according to the CHI -based protocol, enabling the processing cores to access resources via the NVLink interface.
[0219] Optionally, the interconnect component comprises a crosspoint comprising at least four mesh ports and at least two device ports; and wherein the RPU is coupled to a device port of the at least two device ports. Optionally, the coherent interconnect comprises a scalable coherent fabric (SCF), the interconnect component comprises a Cache Switch Node (CSN), and the RPU is coupled to the CSN via the CHI interface. In one embodiment, the xPU may be based on an NVIDIA SCF coherent interconnect that includes CSNs as a crosspoint, and an NVLink-C2C for connecting to an external entity, such as a GPU, via an NVLink interface. Optionally, the SCF comprises an SCF Cache partition (SCC); and wherein the RPU and the SCC are coupled to the CSN, providing the entity, via the NVLink interface, with low-latency access to caching resources of the apparatus. Optionally, the memory comprises dynamic random-access memory (DRAM), and the entity comprises an NVLink Switch, a GPU, or an accelerator.
[0220] In one embodiment, a method comprising: operating a coherent interconnect that utilizes a protocol based on Coherent Hub Interface (CHI-based protocol), comprising an interconnect component that receives CHI-based messages; communicating, via the coherent interconnect, between processing cores and memory controllers that communicate with memory channels coupled to memory having a capacity of at least 64GB; operating a resource provisioning unit (RPU) comprising an NVLink interface and a CHI interface, wherein the NVLink interface utilizes differential pairs and communicates according to an NVLink-based protocol with an entity external to the RPU, and wherein the CHI interface communicates with the interconnect component; and translating, by the RPU, betweenmessages conforming to the NVLink-based protocol and messages conforming to the CHI -based protocol to enable the entity to access resources via the NVLink interface and the coherent interconnect.
[0221] Optionally, the method further comprises translating, by the RPU, first physical addresses associated with the NVLink-based protocol to second physical addresses associated with the CHI -based protocol, and translating NVLink command encodings to corresponding CHI opcodes. Optionally, the RPU comprises a request agent (RA) proxy, and further comprising communicating, by the RA proxy, with the interconnect component according to the CHI-based protocol, enabling the entity to access, via the NVLink interface, resources coupled to the coherent interconnect. Optionally, the RPU comprises a home agent (HA) proxy, and further comprising communicating, by the HA proxy, with the interconnect component according to the CHI -based protocol, enabling the processing cores to access resources via the NVLink interface.
[0222] In one embodiment, a system comprises: a coherent interconnect based on Coherent Hub Interface (CHI) protocol (CHI-based protocol), comprising interconnect components configured to route CHI-based messages; processing cores coupled via the coherent interconnect to memory controllers coupled to memory channels coupled to memory having a capacity of at least 64GB; resource provisioning units (RPUs) comprising external interfaces and CHI interfaces, wherein at least one of the external interfaces comprises an NVLink interface utilizing differential pairs for communication according to an NVLink-based protocol with one or more external entities; wherein the CHI interfaces are coupled to the interconnect components; and wherein the RPUs are configured to translate between protocols utilized by the external interfaces and the CHI -based protocol; whereby the translate enables the external entities to access system resources via the external interfaces and the coherent interconnect.
[0223] Optionally, the RPUs are configured to translate physical addresses from physical address spaces associated with their external interface protocol to addresses from physical address spaces associated with the CHI -based protocol, and to translate command encodings from the external interface protocol to command encodings from corresponding CHI opcodes.
[0224] Optionally, the RPUs comprise at least one of request agent (RA) proxies or home agent (HA) proxies configured to communicate with the interconnect components according to the CHI -based protocol; wherein the RA proxies enable external entities to access memory and I / O resources coupled to the coherent interconnect, and the HA proxies enable the processing cores to access external memory resources via the external interfaces, thereby implementing a distributed shared memory architecture.
[0225] Optionally, at least one of the RPUs comprises an interconnect gateway configured to communicate with a corresponding interconnect component according to the CHI-based protocol; wherein the interconnect gateway utilizes a streaming interface protocol to enable connectivity between the external interface associated with the at least one of the RPUs and the coherent interconnect via the at least one of the RPUs. Optionally, the external interfaces associated with the RPUs may implement various protocol bridging architectures to enable communication between external entities and the coherent interconnect. In one example, an RPU may utilize proxy -based mechanisms such as Request Agent (RA) proxy and Home Agent (HA) proxy for NVLink translations. In alternative implementations, the RPUs may employ direct translation engines that perform stateless or stateful conversion between external protocols and CHI-based messages, transaction queuing and reordering mechanisms that handle protocol -specific ordering requirements, or address remapping units that maintain translation tables for converting between addresses from different physical address spaces. The RPUs may implement credit -based flow control, transaction tracking structures, or protocol-specific state machines that manage the lifecycle of transactions as they traverse between domains. These various implementation approaches may enable external entities to access system memory while system components concurrently access resources attached to the external entities.
[0226] Optionally, the architectural flexibility of the RPUs may enable multiple protocols to co-exist within the system utilizing various mechanisms. Different RPUs in the system may support UALink through UPLI message processing engines, CXL protocol through CXL.mem and / or CXL.cache transaction handlers, PCIe protocol through TLP processing units, or proprietary interconnect protocols through custom translation logic. The system may include RPUs configured for multi-protocol operation, such as multi-protocol RPUs embedded in a Fabric Processing Unit(FPU) or in a software -defined fabric processor, wherein a single RPU implements protocol detection and routing logic, shared transaction buffers with protocol-specific handling, unified address translation units that support multiple addressing schemes, or configurable state machines that adapt to different protocol requirements. The streaming interface protocol utilized by the interconnect gateway may provide a common transport mechanism with protocol -agnostic packetization and framing, enabling these diverse protocols to efficiently communicate with the CHI -based coherent interconnect. The RPUs may implement protocol-specific optimizations such as transaction coalescing, speculative prefetching, or latency hiding techniques while maintaining protocol semantics and coherency requirements utilizing appropriate translation and synchronization mechanisms.
[0227] FIG. 13 A illustrates one embodiment of a system that translates between NVLink-based traffic and coherent interconnect CHI -based traffic. The NVLink connections are coupled via an RPU to an interconnect component such as a crosspoint (e.g., XP), which may serve as a fundamental building block of a coherent interconnect, providing switching or routing of CHI messages between participating elements such as request nodes, home nodes, gateways, protocol bridges, or other elements that connect to the coherent interconnect. The RPU may translate directly between NVLink traffic utilized by an entity, such as a GPU or a CPU, to CHI -based traffic utilized by the interconnect component, possibly eliminating intermediate protocol translations. Alternatively, the RPU may translate between an NVLink traffic and CHI traffic by utilizing intermediate protocols such as Advance Extensible Interface (AXI), or AXI Coherency Extensions Lite (ACE-Lite), or by utilizing streaming interface protocols such as Credited extensible Stream (CXS). Direct translation from NVLink to CHI may provide high-performance connectivity between a GPU coupled to the NVLink interface and memory coupled to the coherent interconnect, a performance gain that may be reflected via lower-latency accesses to memory and higher-bandwidth of reads and writes.
[0228] FIG. 13B illustrates one embodiment of a transaction flow diagram (TFD) showing the translation of NVLink traffic to CHI traffic. An entity, such as a GPU or a CPU, initiates an NVLink read request, that is received by the RPU via the NVLink interface. The RPU translates the NVLink request to a CHI request carrying ReadOnce, optionally translating the physical address (AS.1.1) associated with NVLink to a physical address (AS.2.1) associated with CHI. The RPU may capture identification information associated with the NVLink request, such as source identifier of the requesting entity, and transaction Tag identifier, and may record the information together with identification information associated with the CHI request generated, such as the transaction ID (TxnlD), in order to support the generation of an NVLink response for the NVLink request received from the entity. The RPU sends the CHI request, via the CHI interface, to an interconnect component, such as a crosspoint (e.g., an XP on a CHI coherent interconnect), that forwards the request to a home node. The home node processes the request and issues a CHI request carrying ReadNoSnp to a memory controller coupled to the coherent interconnect. The memory controller may read the requested data from memory, and may send the data directly to the RPU, or alternatively the memory controller may send the data to the home node, wherein the home node is responsible for sending the data to the RPU. When the RPU receives the data via the CHI interface, the RPU may issue an NVLink response with the data to the requesting entity, utilizing the identification information the RPU captured when processing and translating the NVLink request.
[0229] FIG. 14A illustrates one embodiment of a system that translates between NVLink-based traffic and CHIbased traffic. The NVLink connections are coupled via an RPU to crosspoint (e.g., XP) interconnect components of the CHI coherent interconnect. The RPU may include request nodes (e.g., RNs), such as I / O-cohcrcnt RN-I nodes and / or RN-D, and / or home nodes (e.g., HNs), such as non-coherent HN-I nodes. This embodiment enables external entities, such as GPUs, CPUs, or accelerators, communicating utilizing NVLink traffic, to access resources within the ARM-based processor's coherent domain utilizing appropriate translations and routing, such as by an RPU translating from NVLink traffic utilized by a GPU entity, to CHI traffic, utilized by a crosspoint (XP) component of the CHI interconnect, wherein a request node or a home node provides the CHI interface for connecting to the XP.
[0230] FIG. 14B illustrates one embodiment of an RPU that translates between NVLink traffic and CHI traffic, utilizing an intermediate protocol based on ARM Advanced Microcontroller Bus Architecture (AMBA) Advance Extensible Interface (AXI) Coherency Extensions Lite (ACE-Lite). The RPU may further translate physical addresses associated with NVLink to physical addresses associated with CHI. The RPU may process and translate the NVLinktraffic, received from an NVLink interface, to ACE -Lite traffic for further processing, and send the ACE -Lite traffic to a request node (e.g., RN). The request node translates the ACE -Lite traffic to CHI traffic and provides a CHI interface for connecting to the coherent interconnect. In this embodiment, The RPU receives from an entity, such as a GPU or a CPU, NVLink traffic that includes a read request. The RPU translates the NVLink traffic to an intermediate ACE-Lite ReadOnce, that is further translated by a request node to a CHI ReadOnce destined to a home node (e.g., HN). The home node processes the CHI ReadOnce and may issue a ReadNoSnp to a memory controller, for servicing the original read request received from the entity via the NVLink interface. The memory controller reads the requested data from memory, and may send the data via the coherent interconnect to the CHI interface of the RPU for delivery to the entity over the NVLink interface.
[0231] FIG. 15A illustrates one embodiment of a system that translates between an interface based on NVLink, and interconnect components that communicate according to a protocol based on ARM CHI. The system enables entities, such as GPUs or CPUs, to access, via an optional NVLink switch, and an NVLink interface, resources coupled to the coherent interconnect. The NVLink connections are coupled, via an RPU, to crosspoint (e.g., XP) interconnect components of the coherent interconnect. The RPU may include a gateway or interface logic (marked GW in the figure), such as CMN multi-Chip Gateway (CCG), Coherent Multi-Chip Link (CML), Cache Coherent Interconnect for Accelerators (CCIX) Gateway (CXG), CHI C2C, or NVLink -C2C, that may include a CHI interface coupled to the coherent interconnect, enabling connectivity between the NVLink interface and the coherent interconnect, via the RPU. The gateway or interface logic may utilize a streaming interface protocol, such as Credited extensible Stream (CXS), to provide packing and un-packing of CHI C2C or an intermediate protocol over the streaming interface. The RPU may further include one or more request nodes (e.g., RN-I), home nodes (e.g., HN-I), optionally placed in the GW, that may enable DMA engines in the RPU to move blocks of data between the NVLink interface and the CHI interface. Examples of the gateway or interface logic include CCG, CML, CXG, CHI C2C, or NVLink-C2C.
[0232] FIG. 15B illustrates one embodiment of an RPU that translates between NVLink traffic and CHI traffic. The RPU may further translate NVLink physical addresses to CHI physical addresses. The RPU utilizes a streaming interface protocol that may be based on ARM Advanced Microcontroller Bus Architecture (AMBA) Credited extensible Stream (CXS). Optionally, the RPU may utilize an intermediate protocol, such as CCIX, PCIe, or CXL, over the streaming interface protocol, and may translate from NVLink to intermediate protocol, and / or from the intermediate protocol to CHI. Optionally or alternatively, the RPU may include interfacing logic such as CHI C2C or NVLink-C2C, that may utilize a streaming interface protocol based on CXS.
[0233] FIG. 16A illustrates one embodiment of a TFD showing a read transaction from an entity such as a GPU to memory resources of an xPU or a memory pool, wherein an RPU provides translations between NVLink traffic, such as traffic based on a protocol utilizing NVLink5, and CHI traffic that may be utilized by the coherent interconnect of the xPU or the memory pool. The RPU may further translate physical addresses associated with NVLink to physical addresses associated with CHI, such as when translating from (AS.1.1) to (AS.2.1), optionally utilizing one stage of address translation. The RPU may utilize a streaming interface protocol, such as CXS, and may optionally utilize PCIe as an intermediate protocol over the CXS streaming interface protocol, translating from NVLink to the PCIe intermediate protocol, and / or from the PCIe intermediate protocol to CHI.
[0234] The entity / GPU initiates the transaction by sending an NVLink Read Request carrying a physical address (AS.1.1) to the RPU, which translates the NVLink Read Request to a PCIe UIO Memory Read Request utilizing a UIOMRd TLP type, optionally translating the physical address (AS.1.1) carried in the NVLink Read Request to a different physical address (AS.2.1) carried in the UIOMRd TLP. The RPU further translates the PCIe UIO Memory Read Request to an ARM CHI REQ carrying ReadOnce and a physical address (AS.2.1 in the illustrated embodiment), which is sent via the coherent interconnect to the Home Node (HN). The Home Node processes the request and sends a subsequent ARM CHI REQ with ReadNoSnp and the physical address (AS.2.1), to the Memory Controller (MC) for retrieving the requested data from memory. The Memory Controller accesses the memory and returns the data via an ARM CHI RD AT message carrying CompData and the requested data. The RPU receives the CHI response and translates it to the intermediate protocol, such as to PCIe UIO Read Completion with Data, utilizing a UIORdCplDTLP type, and further translates from the intermediate protocol to an NVLink response carrying the data, which is sent back to the entity / GPU via the NVLink interface, completing the read transaction.
[0235] When the RPU provides address translations, these address translations may take place during a stage wherein the RPU translates from NVLink to an intermediate protocol, such as PCIe or CXL. Additionally or alternatively, address translations may take place during a stage wherein the RPU translates from the intermediate protocol, such as PCIe or CXL, to CHI. In some embodiments, the RPU may perform address translations in stages, such as from a physical address (AS.1.1) in an NVLink request, to physical address (AS.2.1) in a PCIe request or a CXL request, and to physical address (AS.3.1) in a CHI request, optionally providing physical address space isolation between the NVLink domain, the intermediate protocol domain, and the CHI domain. Opcodes, TLP types, or intermediate protocols shown in this embodiment, serve as an example. Other embodiments may utilize other TLP types such as MRd for a PCIe or CXL request, CplD for PCIe or CXL response, and other intermediate protocols such as CXL. mem or CXL. io.
[0236] FIG. 16B illustrates one embodiment of a TFD showing a read transaction from an entity such as a GPU to memory resources of an xPU or a memory pool, wherein an RPU translates between NVLink traffic, such as traffic based on a protocol utilizing NVLink5, and CHI traffic that may be utilized by the coherent interconnect of the xPU or the memory pool. The RPU may further translate physical addresses associated with NVLink to physical addresses associated with CHI, such as when translating from (AS.1.1) to (AS.3.1), optionally utilizing two stages of address translation with an intermediate address (AS.2.1) that may be associated with an intermediate protocol. The RPU utilizes a streaming interface protocol, such as CXS, and may optionally utilize CXL as an intermediate protocol over the CXS streaming interface protocol, translating from NVLink to the CXL intermediate protocol, and / or from the CXL intermediate protocol to CHI.
[0237] The entity / GPU initiates the transaction by sending an NVLink Read Request carrying a physical address (AS.1.1) to the RPU, which translates the NVLink Read Request to a CXL.cache D2H Request comprising RdCurr, optionally translating the physical address (AS.1.1) carried in the NVLink Read Request to a different physical address (AS.2.1) carried in the CXL.cache D2H Request, wherein (AS.2.1) may be an intermediate address associated with the intermediate protocol. The RPU further translates the CXL.cache D2H Request to an ARM CHI REQ carrying ReadOnce, optionally translating the physical address (AS.2.1) carried in the CXL.cache D2H Request to a different physical address (AS.3.1), carried in the ARM CHI REQ, which is sent via the coherent interconnect to the Home Node (HN). The Home Node processes the request and sends a subsequent ARM CHI REQ withReadNoSnp and the physical address (AS.3.1), to the Memory Controller (MC) for retrieving the requested data from memory. The Memory Controller accesses the memory and returns the data via an ARM CHI RD AT message carrying CompData and the requested data. The RPU receives the CHI response and translates it to the intermediate protocol, such as to CXL.cache H2D Data, and further translates from the intermediate protocol to an NVLink response carrying the data, which is sent back to the entity / GPU via the NVLink interface, completing the read transaction.
[0238] When the RPU provides address translations, these address translations may take place during a stage wherein the RPU translates from NVLink to an intermediate protocol, such as PCIe or CXL. Additionally or alternatively, address translations may take place during a stage wherein the RPU translates from the intermediate protocol, such as PCIe or CXL, to CHI. In some embodiments, the RPU may perform address translations in stages, such as from a physical address (AS.1.1) in an NVLink request, to physical address (AS.2.1) in a PCIe request or a CXL request, and to physical address (AS.3.1) in a CHI request, optionally providing physical address space isolation between the NVLink domain, the intermediate protocol domain, and the CHI domain. Opcodes, TLP types, or intermediate protocols shown in this embodiment, serve as an example. Other embodiments may utilize other opcodes, such as CXL.cache RdShared or CXL.cache RdAny, other TLP types such as MRd for a PCIe or CXL request, CplD for PCIe or CXL response, and other intermediate protocols such as CXL.mem or CXL.io.
[0239] FIG. 17A illustrates one embodiment of a system comprising an external entity coupled to an optional NVLink switch coupled to a processor comprising (such as an xPU) comprising an RPU comprising an NVLink interface, a Request Agent (RA) Proxy, and a Home Agent (HA) Proxy. The RPU may further comprise an NVLinkcontroller, wherein the NVLink controller may include the NVLink interface. The RPU may be coupled to an interconnect component, such as a crosspoint (e.g., XP), optionally via the RA Proxy and / or the HA Proxy, wherein the RPU may communicate with the interconnect component according to a CHI-based protocol. The RPU may be further coupled, via the NVLink interface, and optionally via an NVLink switch, to an external entity, such as a GPU, wherein the RPU may communicate with the external entity according to an NVLink-based protocol. The RPU may translate between messages conforming to the NVLink-based protocol and messages conforming to the CHI -based protocol, possibly enabling the external entity to access resources of the xPU, such as xPU local memory (e.g., DRAM), and / or enabling the xPU to access resources of the external entity, such as remote memory coupled to the entity. The Request Agent (RA) proxy may receive requests that originate outside of the coherent interconnect, such as from remote agents, from the NVLink interface, from the NVLink controller, from an attached accelerator die, or from a remote chip, wherein the RA proxy may represent such remote initiators as a proxy when communicating with the coherent interconnect, e.g., by utilizing a Source ID (SrcID) namespace and a Transaction ID (TxnlD) namespace associated with the coherent interconnect. The Home Agent (HA) proxy may own an address window backed by memory that may be placed on another chip or silicon die, such as on the external entity, wherein the HA proxy may enable processing cores of the xPU to access resources coupled to the external entity, such as memory (e.g., HBM).
[0240] FIG. 17B illustrates one embodiment of a system comprising an xPU, such as a custom accelerator, that may utilize translations between NVLink and CHI, wherein the xPU may utilize NVLink for communicating with a first entity and with a second entity, which may each be a GPU external to the xPU, and wherein the xPU may further utilize CHI for intra-xPU communications between xPU resources coupled to a coherent interconnect of the xPU. The xPU may include first and second NVLink chiplets, or silicon dies, such as NVLink Fusion, coupled to the first and second entities, respectively. The first and second NVLink chiplets may be further coupled to first and second RPUs, respectively, via first and second physical layers (PHYs), respectively. The first and second RPUs may each include a Die-to-Die (D2D) adapter, a Request Agent (RA) Proxy, and / or a Home Agent (HA) proxy, wherein each RPU may communicate with the coherent interconnect, via the RA Proxy and / or the HA Proxy. The first and second PHYs may each include a UCIe PHY, an NVLink-C2C PHY, or a custom PHY.
[0241] The translations between NVLink and CHI may enable the first and / or the second entity to access resources coupled to the coherent interconnect of the xPU; and may further enable processing cores of the xPU to access resources coupled to the first and / or second entity. The translations between NVLink and CHI may further enable the xPU to perform as a switch, such as an NVLink switch, that may utilize NVLink to enable communication between the first entity and the second entity. The first entity may communicate with the second entity via the xPU, such as via the first NVLink chiplet, the first RPU, the coherent interconnect, the second RPU, and the second NVLink chiplet. Similarly, the second entity may communicate with the first entity via the xPU, such as via the second NVLink chiplet, the second RPU, the coherent interconnect, the first RPU, and the first NVLink chiplet
[0242] FIG. 18A illustrates one embodiment of a system comprising an xPU comprising an RPU that translates between NVLink traffic and CHI traffic. The RPU may include a die-to-die (D2D) adapter, such as UCIe D2D adapter or NVLink-C2C adapter, which may perform at least one of: (1) Serve as an interfacing logic coupling the coherent interconnect and a die-to-die link; (2) Packetize CHI C2C into flits that can be streamed out to another chip or die, and correspondingly, handle de -packetization in the reverse direction; (3) Provide a CHI interface for connecting to an interconnect component such as a crosspoint (e.g., XP); or (4) Couple to a PHY such as a UCIe PHY, an NVLink -C2C PHY, or a PCIe PHY, for connecting to an NVLink chiplet, such as NVLink Fusion
[0243] FIG. 18B illustrates one embodiment of a system comprising a third entity (Entity.3), such as a semiconductor device, a CPU, an MxPU, an accelerator, or a memory switch, wherein the third entity may be coupled to a memory, such as DRAM, optionally via memory channels. The third entity may include a coherent interconnect, a first RPU (RPU.l) comprising an NVLink port and a first CHI interface (CHI Interface.!), and a second RPU (RPU.2) comprising a CXL port and a second CHI interface (CHI Interface.2). The third entity may be coupled, via the NVLink port and optionally via a first switch (Switch.1), such as an NVLink switch or an NVSwitch, to a first entity (Entity.1), such as a GPU, wherein the third entity may be further coupled, via the CXL port and optionally viaa second switch (Switch.2), which may be a CXL switch, to a second entity (Entity.2), such as a CXL device (e.g., CXL memory). The third entity may utilize translations between NVLink and CHI that may enable the first entity to access the memory of the third entity, wherein the third entity may further utilize translations between CXL and CHI that may enable the second entity to access the memory of the third entity.
[0244] In some embodiments, the translations between NVLink and CHI, and the translations between CXL and CHI, may enable the third entity to perform as a switch, such as a multi -protocol switch, a hybrid switch, or an NVLink / CXL switch, enabling communication between the first entity and the second entity, such as enabling the GPU to utilize the CXL memory. For example, the first entity may communicate with the second entity via the third entity, such as via the first RPU comprising the NVLink port and the first CHI interface (CHI Interface.!), via the coherent interconnect, and via the second RPU that includes the CXL port and the second CHI interface (CHI Inetrface.2). In another example, the second entity may communicate with the first entity via the third entity, such as via the second RPU, the coherent interconnect, and the first RPU.
[0245] In some embodiments, the third entity may enable communication between the NVLink domain and the CXL domain, such as communication between NVLink ports and CXL ports, or communication between NVLink interfaces and CXL ports, whereas in other embodiments, the communication between the NVLink domain and the CXL domain may be restricted, optionally by an access control list (ACL), such as to a subset of the NVLink ports and / or to a subset of the CXL ports. Additionally or alternatively, communication between the NVLink domain and the CXL domain may be restricted to a subset of allowed address regions associated with one or more address spaces, or may be restricted to a subset of allowed protocols, such as CXL.mem (e.g., not allowing CXL.cache transactions).
[0246] FIG. 19A illustrates one embodiment of a system comprising an xPU or a custom accelerator, coupled to an entity such as a GPU, optionally via an NVLink switch. The xPU includes an RPU which translates between an NVLink traffic and CHI traffic. The RPU includes an NVLink chiplet, such as NVLink Fusion, that provides an NVLink interface for coupling to the external entity. The RPU further includes an NVLink-C2C for coupling the NVLink chiplet to the coherent interconnect, wherein the NVLink -C2C utilizes a CHI interface for connecting to at least one crosspoint of the coherent interconnect. The RPU may provide bi-directional memory access between the xPU and the GPU, enabling the xPU to read from the GPU’s HBM, and enabling the GPU to read from DRAM coupled to the xPU. Alternatively, the RPU may provide unidirectional memory access, enabling the GPU to access xPU memory but not vice-versa, such as by exposing at least some of the xPU resources as a memory expander or a memory pool for use by the GPU.
[0247] FIG. 19B illustrates one embodiment of a system comprising an xPU coupled to an entity such as a GPU. The xPU includes an RPU which translates between NVLink traffic and CHI -based traffic, wherein the RPU includes a CHI interface for coupling to a coherent interconnect, an NVLink-C2C logic, optionally integrated into an NVLink-C2C controller that includes a transactional layer, a data link layer and a physical layer. The RPU further includes an NVLink chiplet, such as NVLink Fusion, for coupling to the GPU, wherein the NVLink chiplet is further coupled to the coherent interconnect via the NVLink-C2C logic, optionally communicating with at least one crosspoint interconnect component according to a protocol based on ARM CHI.
[0248] FIG. 20A illustrates one embodiment of an xPU coupled to an entity such as a CPU or a GPU. The xPU includes an RPU which translates between NVLink traffic protocol and CHI -based traffic. The xPU further includes at least two silicon dies, wherein the first die includes a CHI interface of the RPU, and the second die includes an NVLink interface of the RPU. The second die may further include an optional PCIe PHY to communicate according to PCIe with a device external to the xPU. The first die and the second die are coupled by at least one C2C interface, utilizing chip-to-chip or die-to-die protocols such as CHI C2C or NVLink-C2C. The RPU may enable coherent memory access from the entity to the xPU, and optionally, from the device to the xPU.
[0249] FIG. 20B illustrates one embodiment of an xPU coupled to an entity such as an NVIDIA Blackwell GPU. The xPU includes processing cores, acceleration cores, memory controllers, a coherent interconnect, and an NVLink chiplet, such as NVLink Fusion, that is coupled to the coherent interconnect via a first NVLink-C2C. The NVLink chiplet includes a second NVLink-C2C, and an RPU that translates between NVLink traffic and CHI -based traffic.The RPU includes an NVLink interface for coupling to the entity, and a CHI interface for coupling to the second NVLink-C2C. The NVLink-C2C interfaces are optionally integrated into NVLink-C2C controllers that includes transactional layers, data link layers and physical layers. The RPU may enable the GPU to access, via the NVLink interface, resources mapped to the physical address space utilized by the xPU coherent interconnect. Correspondingly, the RPU may enable the processing cores of the xPU to access, via the NVLink interface, resources of the GPU, such as HBM or GDDR memory.
[0250] FIG. 21 A illustrates one embodiment of a system that may function as a multi -protocol memory switch appliance or a multi -protocol memory pool, and may include an MxPU, CPU, accelerator, or a memory switch ASIC, that may be coupled to two entities, optionally via switches: (1) Entity.1 / GPU via an optional first switch (Switch.1), such as an NVLink switch or NVSwitch, and (2) Entity.2 / Accelerator via an optional second switch (Switch.2), such as a UALink switch. The MxPU includes processing cores and memory controllers coupled to a coherent interconnect that may be based on CHI. The MxPU may utilize different translations for the external interfaces, performed by different RPUs, such as between NVLink-based interfaces and the MxPU coherent interconnect, or between UALink-based interfaces and the MxPU coherent interconnect. The first RPU (RPU.l) may enable Entity.1 / GPU to access resources mapped to a physical address space utilized by the MxPU coherent interconnect, wherein the access is via the optional first switch, the NVLink interface and the MxPU coherent interconnect. Examples of resources mapped to the physical address space utilized by the MxPU coherent interconnect include DRAM or other memory resources of the MxPU. Correspondingly, the second RPU (RPU.2) may enable Entity.2 / Accelerator to access, via the optional second switch, the UALink interface and the MxPU’s coherent interconnect, resources mapped to a physical address space utilized by the MxPU’s coherent interconnect, such as DRAM or other memory resources of the MxPU.
[0251] FIG. 2 IB illustrates one embodiment of a TFD depicting a multi -entity memory access scenario wherein a GPU / first entity and an accelerator / second entity access memory mapped to one or more address spaces utilized by the coherent interconnect (CohlnterMappedMemory) utilizing heterogeneous protocol message translations. Entity.1 / GPU.1 initiates an NVLink Request: Read with SourceID(a.1) to identify the source GPU, DestinationID(b.1) to identify the destination, and Address(AS.1.1) representing a physical address, such as an NVLink network address. RPU.1 translates the NVLink Request to ARM CHI REQ carrying Opcode(ReadOnce) while preserving Addr(AS.1.1) unchanged. Concurrently or sequentially, Entity.2 / Accelerator may initiate a UALink UPLI Request (Req) with ReqCmd(Read), ReqSrcPhysAccID(a.2) to identify the source accelerator, ReqDstPhysAccID(b.2) to identify the destination, and ReqAddr(AS.1.2) representing a request address, such as a network physical address (NPA). RPU.2 translates the UALink UPLI Request to ARM CHI REQ carrying Opcode(ReadOnce) while preserving Addr(AS.1.2) unchanged. Both transactions flow through the coherent interconnect to one or more home nodes, which may send respective ARM CHI REQ messages to one or more memory controllers with Opcode(ReadNoSnp) and the addresses Addr(AS.l.l) and Addr(AS.1.2), respectively. The memory controller(s) retrieve the requested data from the CohlnterMappedMemory and send first and second ARM CHI RD AT messages with Opcode(CompData) carrying *Data.l* and *Data.2*, representing the data retrieved from the addresses AS.1.1 and AS.1.2, respectively. RPU.l translates the first ARM CHI RD AT message to NVLink Response with SourceID(b.l), DestinationID(a.1), and *Data.l* for Entity.1 / GPU. RPU.2 translates the second ARM CHI RD AT message to UALink UPLI Read Response / Data (RdRsp) with RdRspSrcPhysAccID(b.2), RdRspDstPhysAccID(a.2), and RdRspData(*Data.2*) for Entity.2 / Accelerator.
[0252] This embodiment demonstrates how heterogeneous entities utilizing different protocols may share access to the same CohlnterMappedMemory through different RPUs that translate messages between different protocols while preserving the physical addresses unchanged. Alternatively, the embodiment may be viewed as separate NVLink and UALink transactions that utilize the same coherent interconnect infrastructure to access the CohlnterMappedMemory. Still alternatively, the response and read data paths may be implemented according to other designs, such as wherein the memory controller(s) may send the data to the home node(s) that send it to the respective RPUs, or the home node(s) send responses to the RPUs while the memory controllers) send the data directly to the RPUs.
[0253] FIG. 22 illustrates one embodiment of a heterogeneous computing system comprising an xPU or custom accelerator that utilizes an ARM-based mesh architecture with protocol interconnections. The xPU comprises a coherent interconnect implemented as a mesh topology with crosspoints (XP) that route transactions between various system components. Processing cores (C) are distributed throughout the mesh architecture and coupled to the coherent interconnect via the crosspoints. Home nodes are positioned within the mesh, optionally including HN-I nodes that may handle I / O coherent transactions and HN-F nodes that may manage fully coherent transactions. System Node Fully-coherent (SN-F) nodes are coupled to memory controllers (MC) which interface with external memory via physical layers (PHYs). The memory may be DRAM accessible through the memory channels. An entity comprising an NVIDIA Rubin GPU with integrated HBM is coupled to the xPU coherent interconnect via an NVLink chiplet. The NVLink chiplet, which may be an NVLink Fusion chiplet or custom PHY, is coupled utilizing a first physical layer (PHY.1, such as a UCIe PHY) to a die-to-die (D2D) adapter, which may be a CHI D2D Adapter or an NVLink-C2C Adapter, that enables communication between the NVLink chiplet and the coherent interconnect. The NVLink chiplet may provide the NVLink physical layer interface and may additionally provide higher protocol layers including the NVLink data link layer and transaction layer functionality.
[0254] Moreover, a CXL device, which may be a memory expander, may be coupled to the xPU coherent interconnect via a second physical layer (PHY.2) and a Root Port. The Root Port provides the interface between the CXL device and the coherent interconnect, enabling the CXL device to be discovered and configured by the system. The xPU architecture may enable the GPU to access memory resources of the CXL memory expander utilizing translations performed by the RPU and the coherent interconnect. The transaction path denoted as A.l to A.2 in the figure illustrates a memory access flow that may represent an NVLink read transaction initiated by the GPU. The transaction may traverse from the GPU through the NVLink chiplet to the ARM mesh interconnect, wherein the RPU may translate the NVLink read request to a CHI transaction compatible with the ARM mesh interconnect. The CHI transaction may then be routed through the coherent interconnect to the appropriate home node and subsequently to the Root Port, wherein it may be further translated to a CXL.mem MemRd transaction for delivery to the CXL memory expander (A.2). The xPU may additionally comprise accelerator cores that may perform specialized computation tasks and may access both the GPU -attached HBM and the CXL-attached memory through the coherent interconnect.
[0255] FIG. 23 A illustrates one embodiment of a memory switch configured to provide memory to entities coupled to it. Entity.1 is coupled to the memory switch wherein the entity may utilize the memory coupled to the coherent interconnect. The memory switch may function as an NVLink -based switch or an NVLink memory pool, providing switching capabilities between entities while also enabling access to memory resources.
[0256] FIG. 23B illustrates one embodiment of a TFD demonstrating an NVLink request from Entity.1 to access memory. RPU.1 receives an NVLink request carrying a read request comprising an address, and translates it to an ARM CHI request comprising ReadOnce, potentially with a different address due to address translation. The request flows through the coherent interconnect to a home node (HN), which may translate it to a ReadNoSnp transaction destined to a memory controller (MC). The MC retrieves the data from memory and may return the data directly to RPU.l without routing through the HN, or alternatively may send the data through the HN to RPU.l. Then RPU.l generates the NVLink response with the data to Entity.1.
[0257] Modem datacenters face unprecedented computational demands driven by generative Al (GenAI), Large Language Models (LLMs), and distributed machine learning workloads that require massive parallelization across heterogeneous computing resources. These evolving workloads, alongside High-Performance Computing (HPC) applications such as real-time analytics, genomics research, and climate modeling, demand flexible architectures that can efficiently share memory and computational resources across multiple processing elements. The convergence of Al training, inference at scale, and memory -intensive applications has created scenarios where traditional boundaries between compute domains limit system efficiency and scalability.
[0258] The proliferation of diverse interconnect technologies, including Compute Express Link (CXL), NVLink, PCIe, and Ultra Accelerator Link (UALink), enable high-bandwidth communication between processors, accelerators, and memory devices. Each protocol operates within its own addressing scheme and communication paradigm, withdevices utilizing distinct physical address spaces for memory access and resource allocation. As datacenters adopt heterogeneous computing models combining CPUs, GPUs, and domain-specific accelerators, the need for interoperability between different interconnect protocols and address spaces becomes increasingly apparent.
[0259] Current interconnect solutions face limitations when attempting to bridge different domains or translate between addresses from disparate physical address spaces. Devices operating under one protocol cannot directly access resources managed by another protocol, and the rigid coupling between specific form factors and their associated functionality restricts deployment flexibility, preventing organizations from leveraging existing infrastructure while adopting new interconnect technologies. These challenges become acute in memory disaggregation scenarios, multi-tenant cloud environments, and edge computing deployments where resources had better be dynamically allocated across different domains.
[0260] Some of the disclosed embodiments introduce novel solutions that leverage standardized retimer form factors to enable physical address translation and protocol interoperability in datacenter environments. These embodiments implement address translation capabilities within semiconductor devices that utilize the PCIe Retimer Supplemental Features and Standard BGA Footprint Specification, enabling protocol bridging and translations of addresses from different address space in a form factor designed for high-speed signal conditioning. By incorporating computational resources for address translation within retimer-compatible packages, the embodiments enable flexible deployment across existing PCIe infrastructure while providing translation capabilities that extend beyond traditional retimer functionality. The embodiments support various interconnect protocols and address translation scenarios, optionally enabling memory sharing between hosts, protocol conversion for datacenter interconnects, memory expansion and pooling, and accelerator integration across heterogeneous computing environments.
[0261] The implementation within a PCIe Retimer BGA footprint provides physical compatibility with existing retimer infrastructure, allowing organizations to deploy translation capabilities essentially without redesigning board layouts or modifying mechanical designs. The standardized high-speed signaling capabilities inherent to the retimer specification enable low -latency translation suitable for memory -intensive workloads. The drop-in replacement capability allows datacenter operators to upgrade existing retimer deployments with intelligent translation functionality, transforming a secondary signal conditioning infrastructure into intelligent translation resources that enable cross-vendor and cross-architecture scaling of datacenter infrastructures.
[0262] In one embodiment, a semiconductor device comprises an IC package comprising high-speed differential I / O balls positioned according to a ball grid array layout defined by a PCIe Retimer Supplemental Features and Standard BGA Footprint Specification. The semiconductor device further comprises a first interface configured to communicate according to a first protocol, and a second interface configured to communicate according to a second protocol. Additionally, the semiconductor device comprises a computer configured to extract physical addresses carried in messages received from the first interface, translate the physical addresses, generate messages that carry the translated physical addresses, and send the messages via the second interface.
[0263] In another embodiment, a method of operating a semiconductor device comprises transmitting, according to a first protocol, via a first interface of an IC package, wherein the IC package comprises high-speed differential I / O balls positioned according to a ball grid array layout defined by a PCIe Retimer Supplemental Features and Standard BGA Footprint Specification. The method further comprises transmitting, according to a second protocol, via a second interface of the IC package. Additionally, the method comprises extracting, by a computer located in the IC package, first physical addresses carried in first messages received from the first interface, translating, by the computer, the first physical addresses to second physical addresses, generating second messages that carry the second physical addresses, and sending the second messages via the second interface.
[0264] Modem datacenters require flexible interconnect solutions that bridge diverse domains while maintaining compatibility with existing infrastructure. Embodiments herein disclose semiconductor devices that implement translations within IC packages conforming to the PCIe Retimer Supplemental Features and Standard BGA Footprint Specification. The devices comprise first and second interfaces communicating according to first and second protocols respectively, with an embedded computer that extracts physical addresses from messages received via the firstinterface, translates these addresses, and generates messages carrying the translated addresses for transmission via the second interface. This retimer-compatible form factor essentially enables drop-in deployment within existing PCIe and cabling infrastructures while providing protocol bridging and address translation capabilities. The standardized BGA layout provides high-speed differential signaling suitable for low-latency address translation, optionally supporting memory disaggregation, host-to-host memory sharing, accelerator integration, and protocol conversion between CXL, UALink, NVLink, and / or PCIe domains, addressing interoperability challenges.
[0265] The semiconductor device described in the embodiment below performs physical address translation within a standardized PCIe retimer form factor. By extracting physical addresses within messages received from a first interface operating according to a first protocol, translating those addresses, and then generating messages carrying the translated addresses for transmission via a second interface operating according to a second protocol, the device enables communication between otherwise possibly incompatible entities, allowing for memory provisioning from one host to another, creation of shared memory regions between hosts, and / or abstraction of memory device resources, while maintaining compatibility with industry -standard physical packaging specifications.
[0266] In one embodiment, a semiconductor device comprising: an integrated circuit package (IC package) comprising high-speed differential I / O balls positioned according to a ball grid array layout defined by a PCIe Retimer Supplemental Features and Standard BGA Footprint Specification; a first interface configured to communicate according to a first protocol; a second interface configured to communicate according to a second protocol; and a computer configured to extract physical addresses carried in messages received from the first interface, translate the physical addresses, generate messages that carry the translated physical addresses, and send the messages via the second interface.
[0267] The IC package houses the functional components of the device, including the first and second interfaces as well as the computer configured to perform the address translation operations. The IC package uses a standardized ball grid array layout as defined by one or the current of future versions of the PCIe Retimer Supplemental Features and Standard BGA Footprint Specification, such as the PCIe 5.0, 6.0, or 7.0 Retimer Supplemental Features and Standard BGA Footprint Specification, which provides a familiar and compatible form factor for system designers. The High-Speed Differential I / Os are positioned according to this standardized ball grid array layout, allowing the semiconductor device to be readily integrated into system designs that conform to the PCIe 5.0, 6.0, or 7.0 specifications. These differential I / Os support the high-speed communication capabilities required for efficient data transfer between the coupled interfaces.
[0268] The first interface may communicate according to various protocols, such as PCIe, CXL -based protocol (that may include CXL.mem, CXL.cache, or CXL.io), or UALink -based protocols, such as UPLI, depending on the specific application requirements. Similarly, the second interface may communicate according to a different protocol than the first interface, or in some embodiments, the same protocol but with different addressing requirements. The computer within the semiconductor device may be implemented using various processing elements, such as ASICs, FPGAs, microprocessors, or other suitable processing architectures. The computer is configured to extract physical addresses within messages received via the first interface, translate these addresses according to predetermined mapping rules or dynamic translation tables, generate new messages incorporating the translated addresses, and send these translated messages via the second interface. The translation of physical addresses may involve various operations such as address space remapping, offset adjustments, or more complex transformations based on the requirements of the coupled systems. The translation process may be predefined, configurable utilizing management interfaces, or may adapt dynamically based on system conditions.
[0269] Optionally, the ball grid array layout is defined by PCIe 5.0, 6.0, or 7.0 Retimer Supplemental Features and Standard BGA Footprint Specification. The semiconductor device may be compatible with the ball grid array layout defined by the PCIe 5.0, 6.0, or 7.0 Retimer Supplemental Features and Standard BGA Footprint Specification. Complete compatibility with this layout specification means that the pin placements, pin functions, and electrical characteristics are in accordance with the specification, allowing system designers to utilize compatible PCB layouts when incorporating the device.
[0270] Optionally, at least one of Pin Name VD_1 to VD_6 or VD_9 to VD 15 are incompatible with the ball grid array layout defined by PCIe 5.0, 6.0, or 7.0 Retimer Supplemental Features and Standard BGA Footprint Specification. Optionally, these specific pin variations may enable additional functionality not available in standard retimer embodiments, such as specialized signaling for address translations, enhanced debugging capabilities, or additional configuration options. This partial deviation from the standard may allow for enhanced functionality while maintaining backward compatibility with system designs in other aspects.
[0271] Optionally, the messages received from the first interface and sent via the second interface may include additional messages that do not carry physical addresses, wherein such messages may be processed by the computer without performing physical address translations. Additionally or alternatively, the computer may further process additional messages that carry virtual addresses instead of physical addresses, and the messages carrying physical addresses may coexist with other types of messages that may be processed differently by the computer, such that the description of messages carrying physical addresses does not limit the presence or processing of other types of messages that may be communicated between the first interface and the second interface. Furthermore, the computer may apply different processing methods to different types of messages according to their content and / or requirements, which may include forwarding messages without modification, modifying message contents without performing address translations, or performing other types of translations or modifications that may differ from the above described physical address translations.
[0272] It is noted that PCIe 5.0 Retimer Supplemental Features and Standard BGA Footprint Specification and PCIe 6.0 Retimer Supplemental Features and Standard BGA Footprint Specification share a comparable ball grid array layout, which are sometimes referred to herein in the context of IC package comprising high-speed differential I / O balls positioned according to ball grid array layout defined by the PCIe 5.0 or 6.0 Retimer Supplemental Features and Standard BGA Footprint Specification.
[0273] Optionally, the first protocol is based on CXL.mem, the first interface is configured to expose a CXL type-2 device or a CXL type-3 device, the second protocol is based on CXL.cache, and the second interface is configmed to expose a CXL type-1 device or a CXL type-2 device. Optionally, the CXL.mem utilizes messages comprising first physical addresses (PAs) from a first Host Physical Address (HP A) space utilized by a first host coupled to the first interface; and wherein the CXL.cache utilizes messages comprising second PAs from a second HPA space utilized by a second host coupled to the second interface. Optionally, the first protocol conforms to CXL.cache, the first interface is configured to expose a CXL type-1 device or a CXL type-2 device, the second protocol conforms to CXL.cache, and the second interface is configured to expose a CXL type-1 device or a CXL type-2 device. Optionally, the first protocol utilizes messages comprising first physical addresses (PAs) from a first Host Physical Address (HPA) space utilized by a first host coupled to the first interface, and the second protocol utilizes messages comprising second PAs from a second HPA space utilized by a second host coupled to the second interface.
[0274] Optionally, the first protocol conforms to CXL.mem, the first interface is configured to expose a CXL type-2 device or a CXL type-3 device, the second protocol conforms to CXL.mem, and the second interface is configured to expose a CXL Root Port. Optionally, the first protocol utilizes messages comprising first physical addresses (PAs) from a first physical address space utilized by a first host coupled to the first interface, and the second protocol utilizes messages comprising second PAs from a second physical address space exposed by the Computer over the second interface.
[0275] Optionally, the first protocol conforms to PCIe or CXL.io, the first interface is configured to expose an Endpoint, the second protocol conforms to PCIe or CXL.io, and the second interface is configured to expose an Endpoint. Optionally, the first protocol utilizes TLPs or Protocol Data Units (PDUs) comprising first physical addresses (PAs) from a first Host Physical Address (HPA) space utilized by a first host coupled to the first interface, and the second protocol utilizes TLPs or PDUs comprising second PAs from a second HPA space utilized by a second host coupled to the second interface. Optionally, the semiconductor device implements non-transparent bridging functionality between two CXL.io domains, between a PCIe domain and a CXL.io domain, or between two PCIe domains, wherein both the first and second interfaces expose PCIe or CXL.io Endpoints, effectively creating separatePCIe / CXL.io domains that can communicate utilizing the address translation provided by the computer. The first interface may be coupled to a first host and expose a PCIe or CXL.io Endpoint to the first host, while the second interface may be coupled to a second host and expose a PCIe or CXL.io Endpoint to the second host. TLPs or Protocol Data Units (PDUs) from the first host containing physical addresses within its Host Physical Address (HP A) space are received by the first interface, processed by the computer to translate the addresses, and then sent through the second interface to the second host using physical addresses within the second host's HPA space. This non-transparent bridging functionality may allow two hosts, each with its own address space, to communicate with each other without requiring either host to have direct visibility into the other host's memory space. The computer may implement address windows or translation regions that map portions of one host's address space to portions of the other host's address space, allowing controlled access to specific memory regions. The address translation may be configured utilizing various means, such as configuration registers accessible via PCIe configuration space, management interfaces, or other control mechanisms. The translation may be static or may be dynamically updated based on system requirements.
[0276] Optionally, the first protocol utilizes at least one Unordered IO (UIO) Transaction Layer Packet (TLP) Type and at least 50% of Memory Requests transmitted by the first interface utilize Flit Mode (FM) TLP formats, and / or wherein the second protocol utilizes at least one of the TLP Types UIOMRd, UIOMWr, UIORdCpl, UlOWrCpl, UIORdCplD. The computer may translate between Non-Flit Mode (NFM) TLP formats, which may be associated with PCIe or with CXL.io, and may be received on one interface, and FM TLP formats, such as UIO TLPs, which may be associated with PCIe or with CXL.io, and may be transmitted on the other interface, or may perform translations between different UIO TLP formats as required by the coupled hosts or devices.
[0277] Optionally, the computer is further configured to translate the physical addresses to enable access to memory regions storing Large Language Model (LLM) parameters distributed across multiple memory domains, whereby the translated physical addresses enable the first interface to access LLM model weights and activation data residing in a physical address space associated with the second interface.
[0278] Optionally, the IC package is further compatible with power and ground pin distribution defined by PCIe 7.0 Retimer Supplemental Features and Standard BGA Footprint Specification. This compatibility allows the semiconductor device to be physically integrated into system designs that follow the PCIe Retimer specification, possibly minimizing modifications to the PCB layout or mechanical aspects of the system. The package dimensions to be defined in future PCIe Retimer specification (such as the anticipated PCIe 7.0 Retimer Supplemental Features and Standard BGA Footprint Specification) may include specifications for package size, height, ball pitch, and other physical characteristics that affect the mechanical integration of the component into a system. By adhering to these specifications, the semiconductor device may be used as a drop-in replacement for a future standard PCIe retimer in systems designed to accommodate such components. Additionally, the semiconductor device may be compatible with power and ground pin distribution as defined by the PCIe Retimer Specification, enabling proper power delivery and signal integrity when the device is integrated into a system designed according to this specification. The power and ground pin distribution may follow specific patterns designed to minimize noise, crosstalk, or other electrical issues that could affect the performance of high-speed differential signals.
[0279] Optionally, the first and second interfaces are configured to support a x4 lane configuration, comprising high-speed differential pairs arranged to support two sets of four lanes for data transmission; and wherein the IC package is further compatible with package dimensions defined for the x4 retimer by PCIe 6.0 Retimer Supplemental Features and Standard BGA Footprint Specification.
[0280] Optionally, the first and second interfaces are configured to support a x8 lane configuration, comprising high-speed differential pairs arranged to support two sets of eight lanes for data transmission; and wherein the IC package is further compatible with package dimensions defined for the x8 retimer by PCIe 6.0 Retimer Supplemental Features and Standard BGA Footprint Specification.
[0281] Optionally, the IC package is configured to support a xl6 lane configuration, comprising high-speed differential pairs arranged to support two sets of sixteen lanes for data transmission; and wherein the IC package is further compatible with package dimensions defined for the xl6 retimer by PCIe 6.0 Retimer Supplemental Featuresand Standard BGA Footprint Specification. The semiconductor device may support different lane configurations to accommodate various system requirements and bandwidths. When configured to support a x4 lane configuration, the device includes high-speed differential pairs arranged to support two sets of four lanes for data transmission, with the IC package being compatible with package dimensions defined for the x4 retimer by the PCIe Retimer Specification. Similarly, when configured to support a x8 lane configuration, the device includes high-speed differential pairs arranged to support two sets of eight lanes for data transmission, with the IC package being compatible with package dimensions defined for the x8 retimer by the PCIe Retimer Specification. For higher bandwidth applications, the device may support a xl6 lane configuration, including high-speed differential pairs arranged to support two sets of sixteen lanes for data transmission, with the IC package being compatible with package dimensions defined for the xl6 retimer by the PCIe Retimer Specification. The compatibility with the PCIe 5.0, 6.0, or 7.0 Retimer Specification ensures that each configuration can be integrated into systems designed for the corresponding lane width with minimal or no custom design accommodations.
[0282] Optionally, the first and second interfaces support operation at data rates of 32.0 GT / s using Non-Retum-to-Zero (NRZ) signaling as defined in PCIe 5.0 specification, and support operation at 64.0 GT / s using PAM4 signaling as defined in PCIe 6.0 specification, thereby providing compatibility with PCIe 5.0 systems and with PCIe 6.0 systems. The semiconductor device may support various data rates and signaling methods to ensure compatibility with different PCIe generations. For example, the first and second interfaces may support operation at data rates of 32.0 GT / s using NRZ signaling as defined in the PCIe 5.0 specification, and support operation at 64.0 GT / s using PAM4 signaling as defined in the PCIe 6.0 specification. This dual capability ensures that the device can operate in systems designed for either PCIe 5.0 or PCIe 6.0, providing backward compatibility while enabling the higher performance of newer systems.
[0283] Optionally, the first and second interfaces are compatible with PCIe 6.0 Retimer Supplemental Features and Standard BGA Footprint Specification, supporting data rates up to 64.0 GT / s and PAM4 signaling in addition to Non-Retum-to-Zero (NRZ) signaling, such that the semiconductor device is operable to negotiate data rates and signaling methods for both PCIe 5.0 and PCIe 6.0 systems.
[0284] Optionally, the first and second interfaces further comprises ground pins distributed according to PCIe 6.0 Retimer Supplemental Features and Standard BGA Footprint Specification, enabling improved signal integrity at data rates exceeding 32.0 GT / s. The compatibility with both PCIe 5.0 and PCIe 6.0 specifications may extend beyond data rates and signaling methods to include other aspects of the protocols, such as link training sequences, error handling mechanisms, and power management features. The device may automatically negotiate the appropriate data rate and signaling method based on the capabilities of the coupled components, selecting the highest performance mode that is supported by components in the link. The semiconductor device may include ground pins distributed according to the PCIe 6.0 Retimer Specification, which may include optional pin reassignment updates designed to reduce signal crosstalk and improve signal integrity at data rates exceeding 32.0 GT / s. These pin reassignments may involve placement of ground pins near high-speed differential pairs, or other layout optimizations that help maintain signal quality at the higher frequencies associated with 64.0 GT / s operation.
[0285] Optionally, the IC package is further compatible with PCIe 4.0 Retimer Supplemental Features and Standard BGA Footprint Specification, supporting data rates up to 16.0 GT / s. This backward compatibility enlarges the device’s deployment options, from legacy installations to cutting-edge designs.
[0286] Optionally, the first interface is coupled to a first host via one or more switches, and / or the second interface is coupled to a second host via one or more switches. The switches may be PCIe switches, CXL switches, or other compatible switching devices that provide connectivity between hosts and devices. The use of switches allows for more complex system topologies, may provide additional functionality such as flow control, quality -of-service, or traffic management that complements the address translation functionality of the semiconductor device. The semiconductor device may be designed to work transparently with these switches, such that the switches are unaware of the address translation occurring within the device. In configurations where both the first and second interfaces are coupled to hosts through switches, the semiconductor device may effectively serve as a bridge between two separateswitch domains, allowing communication between hosts that would otherwise be isolated from each other by the switch.
[0287] In one embodiment, a method of operating a semiconductor device, comprising: transmitting, according to a first protocol, via a first interface of an integrated circuit package (IC package), wherein the IC package comprises high-speed differential I / O balls positioned according to a ball grid array layout defined by a PCIe Retimer Supplemental Features and Standard BGA Footprint Specification; transmitting, according to a second protocol, via a second interface of the IC package; extracting, by a computer located in the IC package, first physical addresses carried in first messages received from the first interface; translating, by the computer, the first physical addresses to second physical addresses; generating second messages that carry the second physical addresses; and sending the second messages via the second interface.
[0288] Optionally, operating the first interface comprises operating according to CXL.mem utilizing the first messages comprising the first physical addresses within a Host Physical Address (HP A) space utilized by a first host communicating with the first interface; and wherein operating the second interface comprises operating according to CXL.cache utilizing the second messages comprising the second physical addresses within an HPA space utilized by a second host communicating with the second interface.
[0289] Optionally, operating the first interface comprises operating according to first CXL.cache utilizing the first messages comprising the first physical addresses within a Host Physical Address (HPA) space utilized by a first host communicating with the first interface; and wherein operating the second interface comprises operating according to second CXL.cache utilizing the second messages comprising the second physical addresses within an HPA space utilized by a second host communicating with the second interface.
[0290] Optionally, operating the first interface comprises operating according to first CXL.mem utilizing the first messages comprising the first physical addresses within a Host Physical Address (HPA) space utilized by a first host communicating with the first interface; and wherein operating the second interface comprises operating according to second CXL.mem utilizing the second messages comprising the second physical addresses within a physical address space exposed by the computer over the second interface.
[0291] Optionally, operating the first interface comprises operating according to first PCIe or CXL.io utilizing the first messages comprising the first physical addresses within a Host Physical Address (HPA) space utilized by a first host communicating with the first interface; and wherein operating the second interface comprises operating according to second PCIe or CXL.io utilizing the second messages comprising the second physical addresses within an HPA space utilized by a second host communicating with the second interface.
[0292] Optionally, operating the first and second interfaces comprises operating at data rates of 32.0 GT / s using Non-Retum-to-Zero (NRZ) signaling as defined in PCIe 5.0 specification, or operating at 64.0 GT / s using Pulse Amplitude Modulation 4-level (PAM4) signaling as defined in PCIe 6.0 specification, thereby providing compatibility with PCIe 5.0 systems and with PCIe 6.0 systems.
[0293] Optionally, operating the first and second interfaces comprises operating according to PCIe 6.0 Retimer Supplemental Features and Standard BGA Footprint Specification, supporting data rates up to 64.0 GT / s and Pulse Amplitude Modulation 4-level (PAM4) signaling in addition to Non-Retum-to-Zero (NRZ) signaling, such that the semiconductor device negotiates data rates and signaling methods for both PCIe 5.0 and PCIe 6.0 systems. Optionally, a non-transitory computer-readable medium comprises instructions that, when executed by a processor, cause the processor to perform the method of claim 23.
[0294] FIG. 24A illustrates one embodiment of a system comprising a computer coupled between a first interface (Interface.!) that may communicate according to CXL.mem and a second interface (Interface.2) that may communicate according to CXL.cache. The computer may be implemented in an IC package having high-speed differential I / O balls positioned according to a ball grid array layout defined by the PCIe 5.0, 6.0, or 7.0 Retimer Supplemental Features and Standard BGA Footprint Specification. The first interface may expose a CXL type-2 device or a CXL type-3 device and may communicate according to CXL.mem with a first entity (Entity.1), such as a first host (Host.1), optionally via a first Root Port (RP.1) of the first entity. The second interface may expose a CXLtype-1 device or a CXL type-2 device and may communicate according to CXL.cache with a second entity (Entity.2), such as second host (Host.2), optionally via a second Root Port (RP.2) of the second entity. The computer may extract physical addresses within messages received via the first interface, wherein these addresses may refer to a first HPA space utilized by the first entity; translate these addresses; and generate messages carrying the translated physical addresses for transmission via the second interface; wherein these translated addresses may correspond to a second HPA space utilized by the second entity. Optional CXL switch(es) may be positioned between the first interface and the first entity, and / or between the second interface and the second entity.
[0295] FIG. 24B illustrates one embodiment of a TFD demonstrating translations between CXL.mem messages received from a first entity (Entity.1), such as a first host (Host.1), and CXL.cache messages sent to a second entity (Entity.2), such as a second host (Host.2), possibly enabling the first entity to access resources mapped to the address space utilized by the second entity. The first entity may initiate a CXL.mem transaction that includes a CXL.mem M2S Request comprising MemOpcode(MemRd), Tag(p.l.l), and Address(AS.l.l). The computer may translate the CXL.mem transaction to a CXL.cache transaction that includes a CXL.cache D2H Request comprising Opcode(RdCurr), Command Queue ID CQID(q.2.1), and Address(AS.2.1), and may send the CXL.cache D2H Request to the second entity. Upon receiving a response from the second entity, which may include a CXL.cache H2D Data comprising CQID(q.2.1) and Data(*Data.l*), the computer may translate the CXL.cache H2D Data to a CXL.mem S2M DRS comprising Opcode(MemData), Tag(p.l.l), and Data(*Data.l*). The computer may perform further translations, such as opcode translations, e.g., translating between CXL.mem M2S Request opcodes, such as MemRd*, and CXL.cache D2H Request opcodes, such as RdCurr, RdOwn, RdShared, RdAny, RdOwnNoData, ItoMWr, WrCur, CLFlush, CleanEvict, DirtyEvict, CleanEvictNoData, WOWrlnv, WOWrlnvF, Wrlnv, or CacheFlushed. The computer may further perform other translations, such as translations between messages conforming to CXL.mem and messages conforming to CXL.cache, translations between CXL.mem Tags and CXL.cache CQIDs, translations between reserved fields, and / or translations between reserved and non-reserved fields. In some embodiments, the computer may act as a protocol endpoint, such as a first CXL device (e.g., a CXL type-3 device or CXL type-2 device), and terminate the CXL.mem transaction. The computer may issue the CXL.cache transaction, optionally acting as an independent protocol initiator, such as a second CXL device (e.g., a CXL type-1 device or a CXL type-2 device), and may utilize translated fields from the CXL.mem transaction for constmcting the CXL.cache transaction. In other embodiments, the computer may maintain, at least partly, an end-to-end transaction context along the path between the first entity and the second entity, optionally without terminating the CXL.mem transactions, such as by preserving, at least partly, transaction-related identification fields. In one example, the first entity may utilize 12-bit CXL.mem Tags when communicating with the computer, and the computer may reuse the 12-bit CXL.mem Tags for constructing 12-bit CXL.cache CQIDs when communicating with the second entity, hence optionally preserving, at least partly, a 12-bit transaction identifier over the path between the first entity and the second entity, for maintaining, at least partly, an end-to-end transaction context along that path.
[0296] FIG. 24C illustrates one embodiment of a TFD demonstrating translations between CXL.mem messages received from a first entity (Entity.1), such as a first host (Host.1), and CXL.cache messages sent to a second entity (Entity.2), such as a second host (Host.2), possibly enabling the first entity to access resources mapped to the address space utilized by the second entity, wherein such accesses from the first entity may affect cacheline states maintained in the second entity, such as by marking cachelines in shared state. The first entity may initiate a CXL.mem transaction that may include a CXL.mem M2S Request comprising MemOpcode(MemRd*), Tag(p.4.1), and Address(AS.4.1). A computer may translate the CXL.mem transaction to a CXL.cache transaction that may include a CXL.cache D2H Request comprising Opcode(RdShared), Command Queue ID CQID(q.3.1), and Address(AS.3.1), and may send the CXL.cache D2H Request to the second entity. Upon receiving one or more responses from the second entity, which may include a CXL.cache H2D Data comprising CQID(q.3.1) and Data(*Data.2*), and may further include a CXL.cache H2D Response comprising CQID(q.3.1) and GO-S, the computer may translate the one or more responses to a CXL.mem S2M DRS comprising Opcode(MemData), Tag(p.4.1), and Data(*Data.2*).
[0297] FIG. 25A illustrates one embodiment of a system comprising a computer coupled between a first interface(Interface.1) and a second interface (Interface.2). The first interface may expose a CXL type-1 device or a CXL type-2 device, and may communicate according to a first CXL.cache with a first entity, such as a first host (Host.l), optionally via a first CXL Root Port (CXL RP.1) of the first entity. Similarly, the second interface may expose a CXL type-1 device or a CXL type-2 device, and may communicate according to a second CXL.cache with a second entity (Entity.2), such as a second host (Host.2), optionally via a second CXL Root Port (CXL RP.2) of the second entity. The computer may : extract physical addresses within messages received via the first interface, wherein these addresses may be from a first HPA space utilized by the first entity; translate these addresses; and generate messages carrying the translated physical addresses for transmission via the second interface; wherein these translated addresses may correspond to a second HPA space utilized by the second entity. Optionally, the interfaces may be implemented as internal and / or external interfaces of a semiconductor device comprising the computer.
[0298] FIG. 25B illustrates one embodiment of a TFD demonstrating translations, performed by a computer, between first CXL.cache messages received from a first entity (Entity.1), such as a first host (Host.l), and second CXL.cache messages sent to a second entity (Entity.2), such as a second host (Host.2), possibly enabling the first entity to maintain, at least partly, memory sharing and / or memory coherency with the second entity, such as by enabling the first entity to invalidate cachelines in the second entity. The first entity may initiate a first CXL.cache transaction that includes a CXL.cache H2D Request comprising Opcode(SnpInv), UQID(t.1.1), and Address(AS.1.1). The computer may translate the first CXL.cache transaction to a second CXL.cache transaction that includes a CXL.cache D2H Request comprising Opcode(CLFlush), CQID(q.2.1), and Address(AS.2.1), and may send the CXL.cache D2H Request to the second entity. Upon receiving a response from the second entity, which may include a CXL.cache H2D Response comprising Opcode(GO-I) and CQID(q.2.1), the computer may translate the CXL.cache H2D Response to a CXL.cache D2H Response comprising Opcode(RspIHitl), and UQID(t.l.l). The computer may perform further translations, such as opcode translations, e.g., translating between CXL.cache H2D Request opcodes, such as Snp* (e.g., SnpData, Snplnv, and SnpCur), and CXL.cache D2H Request opcodes, such as RdCurr, RdOwn, RdShared, RdAny, RdOwnNoData, ItoMWr, WrCur, CLFlush, CleanEvict, DirtyEvict, CleanEvictNoData, WOWrlnv, WOWrlnvF, Wrlnv, or CacheFlushed. The computer may further perform other translations, such as field translations between messages conforming to the first CXL.cache transaction and messages conforming to the second CXL.cache transaction, such as translations between UQIDs and CQIDs, translations between reserved fields, and / or translations between reserved and non-reserved fields.
[0299] FIG. 25C illustrates one embodiment of a TFD demonstrating translations performed by a computer between CXL.cache transactions, wherein the transactions may include messages such as requests, responses, and optionally data messages. A first entity (Entity.1), such as a first host (Host.1), may send to the computer a CXL.cache H2D Request comprising Opcode(SnpCur), UQID(t.1.1), and Address(AS.1.1), wherein the CXL.cache H2D Request may indicate a snoop request for the current version of a cacheline. The computer may translate the CXL.cache H2D Request to a CXL.cache D2H Request comprising Opcode(RdCurr), CQID(q.2.1), and Address(AS.2.1), wherein the CXL.cache D2H Request may indicate a read request from the computer to the second entity for the current version of the cacheline. The computer may provide intent -based translations, such as by identifying intents in CXL.cache H2D Requests received from the first entity, such as intents to get the current version of a cacheline, and utilizing the identified intents for translating between the CXL.cache H2D Requests and the CXL.cache D2H requests. The second entity may respond to the CXL.cache D2H Request comprising the RdCurr opcode with a CXL.cache H2D Data comprising CQID(q.2.1) and Data(*Data.1*). The computer may translate the CXL.cache H2D Data to a CXL.cache D2H Response comprising Opcode(RspVFwdV) and UQID(t.1.1) and may send the CXL.cache D2H Response to the first entity. The computer may further translate the CXL.cache H2D Data to a CXL.cache D2H Data comprising UQID(t.1.1) and Data(*Data.1*) and may send the CXL.cache D2H Data to the first entity.
[0300] FIG. 26A illustrates one embodiment of a system comprising a computer coupled between a first interface (Interface.!) and a second interface (Interface.2), wherein both the first and second interfaces may communicate according to CXL.mem. The first interface may expose resources associated with a second device (Device.2), such as a CXL type-2 device or a CXL type-3 device, optionally comprising a second endpoint (EP.2), and may communicateaccording to CXL.mem with a first entity (Entity.1), such as a first host (Host.1), possibly via a first Root Port (RP.1) of the first host. The second interface may expose a second host (Host.2), optionally comprising a second Root Port (RP.2) and may communicate according to CXL.mem with a second entity (Entity.2), such as a first CXL device (Device.1), that may include a first Endpoint (EP.l). Additionally or alternatively, the first CXL device may include a Global Fabric -Attached Memory (G-FAM) Device (GFD). The computer may extract physical addresses from messages received via the first interface, wherein these addresses may be from a first HPA space utilized by the first host; translate these addresses; and generate messages carrying the translated physical addresses for transmission via the second interface; wherein these translated addresses may correspond to a physical address space exposed by the Computer over the second interface. Optional CXL switch(es) may be positioned between the first interface and the first entity, and / or between the second interface and the second entity. In some embodiments, the computer and at least one of the first entity and the second entity may be included within the same IC package, optionally coupled via one or more UCIe links.
[0301] FIG. 26B illustrates one embodiment of a transaction flow diagram (TFD) demonstrating translations, optionally performed by a computer, between first CXL.mem messages received from a first entity (Entity.1), such as a first host (Host.l), that may utilize a first CXL.mem, and second CXL.mem messages, sent to a second entity (Entity.2), such as a first CXL device (Device.1), that may utilize a second CXL.mem, possibly enabling the computer to abstract resources of the second entity, and possibly enabling the first entity to access resources of the second entity utilizing different memory flow types, such as utilizing optimized type-3 memory flows, instead of elaborated type-2 memory flows that may be utilized by the second entity. Additionally or alternatively, the computer may further initiate speculative memory reads targeting the second entity, and may handle memory prefetching on behalf of the first entity, possibly acting as a proxy of the first entity when communicating with the second entity. The first entity may initiate a first CXL.mem transaction that may include a first CXL.mem M2S Req comprising MemOpcode(MemRdData), SnpType(No-Op), MetaField(No-Op), MetaValue(N / A), Tag(p.2.1), and Address(AS.2.1). The computer may translate the first CXL.mem transaction to a second CXL.mem transaction that may include a second CXL.mem M2S Req comprising MemOpcode(MemRd*), SnpType(SnpCur), MetaField(MSO), MetaValue(I), Tag(p.l.l), and Address(AS.l.l), and may send the second CXL.mem M2S Req to the second entity. Upon receiving one or more responses from the second entity, that may include a CXL.mem S2M NDR comprising Opcode(Cmp), MetaField(No-Op), MetaValue(NA), and Tag(p.1.1), and may further include a first CXL.mem S2M DRS comprising Opcode(MemData), MetaField(No-Op), MetaValue(NA), Tag(p.l.l), and Data(*Data*), the computer may translate the one or more responses from the second entity to a second CXL.mem S2M DRS comprising Opcode(MemData), Tag(p.2.1), and Data(*Data*), and may send the second CXL.mem S2M DRS to the first entity.
[0302] One example of a speculative memory read targeting the second entity includes a CXL.mem M2S Req comprising MemOpcode(MemSpecRd) and Address(AS.1.2), which may utilize the speculative memory reads, optionally on behalf of the first entity, to facilitate data prefetches and potentially reduce read latency from the second entity. When utilizing MemSpecRd, some of the CXL.mem M2S Req fields, such as Tag, MetaField, MetaValue, and SnpType, may be reserved. The computer may perform further translations, such as opcode translations, e.g., translating between a first CXL.mem M2S Req opcode, such as MemRdData, and a second CXL.mem M2S Req opcode, such as MemRd. The computer may further perform other translations, such as field translations between messages conforming to the first CXL.mem and messages conforming to the second CXL.mem, such as translations between CXL.mem Tags of the two protocols, translations between values of reserved fields of the two protocols, and translations between values of reserved and non-reserved fields of the two protocols. In some embodiments, the computer may perform translations between protocols conforming to different CXL revisions, such as translating between transactions of the first CXL.mem conforming to CXL 1.1, which may be utilized by the first entity, and transactions of the second CXL.mem conforming to CXL 2.0, which may be utilized by the second entity.
[0303] In still some embodiments, the computer may act as a second device (Device.2), such as a CXL type-3 device or CXL type-2 device optionally comprising a protocol endpoint, and terminate the first CXL.mem transaction. The computer may then issue the second CXL.mem transaction, optionally acting as an independent protocol initiator,such as a second host (Host.2), and may utilize translated fields from the first CXL.mem transaction for constructing the second CXL.mem transaction. In other embodiments, the computer may maintain, at least partly, an end-to-end transaction context along the path between the first entity and the second entity, optionally without terminating CXL.mem transactions received from the first entity, such as by preserving, at least partly, transaction-related identification fields. In one example, the computer may reuse CXL.mem Tags received from the first entity for constructing CXL.mem Tags sent to the second entity, hence optionally preserving, at least partly, a transaction identifier over the path between the first entity and the second entity, for maintaining, at least partly, an end-to-end transaction context along that path.
[0304] FIG. 27 A illustrates one embodiment of a system comprising a computer coupled between first and second interfaces. The first interface (Interface.!) may communicate according to a UALink -based protocol, such as UPLI, with a first entity (Entity.1), which may be an accelerator. The second interface (Interface.2) may communicate according to a PCIe-based protocol, such as a protocol conforming to PCI Express Base Specification Revision 6.2, with a second entity (Entity.2), which may be a PCIe host or a PCIe device. The computer may be implemented in an IC package having high-speed differential I / O balls positioned according to a ball grid array layout defined by a retimer specification, such as the PCIe 5.0, 6.0, or 7.0 Retimer Supplemental Features and Standard BGA Footprint Specification. The computer may extract physical addresses from requests received via the first interface, wherein these addresses may refer to a Network Physical Address (NPA) space utilized by the first entity. The computer may further translate these addresses, and generate requests carrying the translated physical addresses for transmission via the second interface; wherein these translated addresses may correspond to a Host Physical Address (HP A) space utilized by the second entity. Optional UALink switch(es) may be positioned between the first interface and the first entity. Similarly, optional PCIe switch(es) may be positioned between the second interface and the second entity.
[0305] FIG. 27B illustrates one embodiment of a TFD demonstrating translations between UALink-based requests, such as UPLI requests, received from a first entity (Entity.1), such as an accelerator, and PCIe TLPs sent to a second entity (Entity.2), such as a PCIe host or a PCIe device, possibly enabling the first entity to access resources mapped to an address space utilized by the second entity. The first entity may initiate a UPLI transaction that includes a UPLI Request (Req) comprising Request Command (e.g. ReqCmd(Read)), Request Source Physical Accelerator ID (e.g., ReqSrcPhysAccID(a.l)), Request Destination Physical Accelerator ID (e.g., ReqDstPhysAccID(b.l)), Request Address (e.g., physical addresses, such as NPAs ReqAddr(AS.l.l)), Request Tag (e.g., ReqTag(c.l.l)), and Request Length (e.g., ReqLen(d.1.1)). The computer may translate the UPLI transaction to a PCIe transaction that includes a PCIe Memory Read Request (MRd) comprising physical addresses, such as Host Physical Addresses (HP As), Address(AS.3.1), Tag(w.3.1), and Length(d.3.1), and may send the PCIe MRd to the second entity. Upon receiving a completion from the second entity, which may include a PCIe Completion with Data (CplD) comprising Tag(w.3.1) and DataPayload(*Data.l*), the computer may translate the PCIe CplD to a UPLI Read Response / Data (RdRsp) comprising Read Response Source Physical Accelerator ID (e.g., RdRspSrcPhysAccID(b.l)), Read Response Destination Physical Accelerator ID (e.g., RdRspDstPhysAccID(a.l)), Read Response Transaction Tag (e.g., RdRspTag(c.1.1)), and Read Response Data (e.g., RdRspData(*Data.1*)), and send the UPLI RdRsp to the first entity.
[0306] The computer may perform further translations, such as protocol translations, opcode translations, command translations, TLP type translations, and translations between messages conforming to the UALink-based protocol (e.g., UPLI messages) and protocol data units (PDUs) of the PCIe-based Protocol (e.g., PCIe TLPs), Tag translations, traffic class (TC) translations, and / or cross-field translations; wherein the computer may maintain tracking between Tags associated with the UALink-based protocol and Tags associated with the PCIe-based protocol, such as in order to associate responses with their corresponding requests. In some embodiments, the computer may issue more than one PCIe transaction in response to receiving a UPLI request from the first entity, such as when splitting a UPLI read request for a large block of data to multiple smaller PCIe memory read requests, or when prefetching data from the second entity.
[0307] In one embodiment, PCIe MRd and PCIe CplD TLPs may be utilized by legacy PCIe hosts or devices, whereas recent PCIe hosts or PCIe devices may utilize PCIe UIO Memory Read Request (UIOMRd) and PCIe UIORead Completion with Data (UIORdCplD) TLPs, leveraging the PCIe Unordered IO (UIO) optional capability, that is intended to address the limitations of the PCI / PCIe fabric -based ordering rules, and enables fabrics with multiple paths between a source and destination to be supported, optionally enabling higher-bandwidth communication. The computer may perform translations of requests or transactions initiated from the UALink-based domain to the PCIe domain, may perform additional translations of requests or transactions initiated from the PCIe domain to the UALink-based domain, or may perform translations of requests or transactions initiated from both the UALink-based domain and the PCI domain.
[0308] FIG. 27C illustrates one embodiment of a TFD demonstrating translations between physical addresses carried in UALink-based requests, such as UPLI requests, received from a first entity (Entity.1), and physical addresses carried in PCIe UIO TLPs sent to a second entity (Entity.2), possibly enabling the first entity to access resources mapped to an address space utilized by the second entity. The first entity may initiate a UPLI transaction that includes a UPLI Request (Req) comprising ReqCmd(Read), ReqSrcPhysAccID(a.l), ReqDstPhysAccID(b.l), ReqAddr(AS.2.1), ReqTag(c.2.1), and ReqLen(d.2.1). The computer may translate the UPLI transaction to a PCIe UIO transaction that includes a PCIe UIOMRd comprising Address(AS.4.1), Tag(w.4.1), and Length(d.4.1), and may send the PCIe UIOMRd to the second entity. Upon receiving a completion from the second entity, which may include a PCIe UIORdCplD comprising Tag(w.4.1) and DataPayload(*Data.2*), the computer may translate the PCIe UIORdCplD to a UPLI RdRsp comprising RdRspSrcPhysAccID(b.1), RdRspDstPhysAccID(a.1), RdRspTag(c.2.1), and RdRspData(*Data.2*), and may send the UPLI RdRsp to the first entity. In some embodiments, the computer may issue more than one PCIe UIO transaction in response to receiving a UPLI request from the first entity, such as when splitting a UPLI read request for a large block of data to multiple smaller PCIe UIO memory read requests, or when prefetching data from the second entity. The computer may perform translations of requests or transactions initiated from the UALink-based domain to the PCIe domain, may perform additional translations of requests or transactions initiated from the PCIe domain to the UALink-based domain, or may perform translations of requests or transactions initiated from both the UALink-based domain and the PCI domain.
[0309] When referring to fields, operations, or operation types associated with communication protocols, the terms "opcode", "command", “TLP type”, "request", “request type”, “transaction”, and “transaction type” may be used herein interchangeably as long as they refer to the same operation, and unless a particular context specifies otherwise. This interchangeable usage may apply to data indicative of operation types (such as a field or a set of fields) within messages, packets (such as TLPs), flits, phits, frames, protocol data units (PDUs), or other protocol data structures, as well as to descriptions of protocol operations, requests, transactions, or communications across different communication protocols. For example, a “CXL.cache DirtyEvict opcode”, a “CXL.cache DirtyEvict command”, and a “CXL.cache DirtyEvict request” may refer to the same operation where a device communicates with a host, such as via a D2H Request message, asking the host to evict a full 64 -byte modified cacheline from the device. Likewise, an "ARM CHI ReadOnce opcode", an "ARM CHI ReadOnce command", an "ARM CHI ReadOnce request", and an “ARM CHI ReadOnce transaction” may refer to the same operation that specifies a read within the CHI framework, whether referring to the actual field within a CHI message or to the operation itself. Similarly, a "UPLI read command", a "UPLI read opcode", a "UPLI read request", and a “UPLI read transaction” may refer to the same operation, field, or set of fields within a UPLI message that indicates a read within the UPLI framework.
[0310] Asterisks (*) may be utilized as wildcard notations within the context of a specific embodiment and / or example, such as for representing a subset of relevant operations within a broader set of operations that may be indicated by opcodes, TLP types, commands, requests, request types, transaction, or transaction types, collectively referred to in this specific paragraph as “operation types”. The subset of relevant operations may include operation types that are relevant to the revisions or standards being discussed, encompassing both existing operation types and potential future operation types that may be introduced in subsequent versions of the applicable interconnect standards, including CXL, UALink, SUE, PCIe, UCIe, ARM CHI, ARM AXI, or protocol implementations based on NVLink technology, provided they are applicable and relevant to the embodiment in question. For example, the wildcard operation type ReadOnce* may represent a subset of relevant requests or transactions within the ARM CHIspecifications, which may include, but is not limited to: ReadOnce, ReadOnceCleanlnvalid, and ReadOnceMakelnvalid. Similarly, the wildcard operation type MemRd* may represent a subset of relevant opcodes within the CXL standard, which may include, but is not limited to: MemRd, MemRdData, MemRdFwd, MemRdTEE, MemRdDataTEE, or other opcodes that may be introduced in future CXL standard revisions, provided they are relevant to the specific embodiment under consideration. Likewise, the wildcard operation type *Rd* may represent an even broader subset of relevant operations across different protocols or different standards, which may encompass, but is not limited to: (1) ReadNoSnp, ReadOnce, ReadClean, ReadShared, ReadUnique and MakeReadUnique commands in ARM CHI; (2) UIOMRd and MRd TLP types in CXL.io; (3) RdCurr, RdOwn, RdShared, RdAny, and RdOwnNoData opcodes in CXL.cache; (4) MemRd, MemRdData, MemRdFwd, MemRdTEE, MemRdDataTEE, MemSpecRd, or MemSpecRdTEE opcodes in CXL.mem; (5) read commands in UALink UPLI; (6) memory read TLP types in PCIe; (7) read-class operations in SUE; or (8) read request types in NVLink-based protocol implementations, provided these operation types are applicable to the specific embodiment being described. It is noted that the wildcard notation does not extend to operation types that are irrelevant to the embodiment in question, even if such operation types exist within the broader specifications of the respective standards.
[0311] The wildcard form "*Data*" may be utilized for denoting essentially the same underlying information (“the Data”) irrespective of its representation or state (at rest, in transit, or in use). *Data* may encompass functionally equivalent forms, transformations and reverse -transformations of “the Data”, such as encoding / decoding, packetization / framing, encapsulation, serialization, mapping, scrambling, compression, encryption, segmentation / reassembly, distribution / replication, or splitting / merging, represented in a suitable stmcture, manner, form, or format that may be carried by or interoperate with the applicable interconnect standard specifications, such as CXL, UALink, SUE, PCIe, UCIe, ARM CHI, ARM AXI, or NVLink-based protocol implementations. Encryption of the Data may include but is not limited to: CXL Integrity and Data Encryption (CXL IDE), UALink encryption mechanisms, SUE security features, PCIe Data Object Exchange (DOE) encryption, or when using different encryption keys on different interconnect links or channels. Moreover, *Data* may further encompass an equivalent representations of “the Data”, such as when carried in PDUs that may be associated with the same protocol or associated with different protocols, wherein PDUs may refer to: (1) messages, such as CXL.cache H2D Data messages; (2) requests, such as CXL.mem M2S Request with Data (RwD), or NVLink write request with data; (3) responses, such as CXL.mem S2M Data Response (DRS); (4) completions, such as PCIe Completion with Data (CplD), or PCIe UIO Read Completion with Data (UIORdCplD); or (5) beats, such as UALink UPLI Data Beats carrying Read Response Data.
[0312] *Data* may also denote PDUs having collectively essentially the same payload, such as when splitting a 64B cacheline received over a single CXL.mem S2M DRS message into 2x32B smaller transfers carried in two CXL.cache H2D Data messages, or when an RPU may split a UPLI read request for a large block of data (e.g., 256B) into multiple smaller requests, such as when the RPU translates between the UPLI request and a request associated with another protocol, such as CXL.mem, that may respond with no more than 64B per each request. Additionally, *Data* is intended to cover all forms of data transmissions and references to data defined in the applicable interconnect standard specifications, such as in the case of CXL.mem S2M DRS wherein the opcode MemData is followed by “the Data” itself, CXL.cache H2D Data transfer wherein the CXL Specification refers to “the Data” as "Data", UALink UPLI data payloads, SUE data units, PCIe TLP payloads, or NVLink-based data transmissions. Moreover, *Data* may also encompass metadata associated with the primary data payload, and may also include trimmed variants of “the Data” such as when responding to a 64B read from a CPU that uses 128B cachelines.
[0313] Depending on the context, each line, arrow, label, and / or box illustrated in the figures may represent one or more lines, arrows, labels, and / or boxes. For example, *Rd* M2S request in CXL, *Rd* read command in UALink UPLI, *Rd* read transaction in SUE, *Rd* memory read TLP in PCIe, or *Rd* read request in an NVLink-based protocol may encompass one or more *Rd* or data messages (which are relevant to the specific embodiment and applicable standard), even though each may be represented by a single arrow. Additionally, optional messages, such as *Cmp* S2M NDR message in CXL, completion messages in UALink, acknowledgment messages in SUE,completion TLPs in PCIe, or response messages in NVLink -based protocols, may be explicitly depicted or implicitly included within the mandatory messages or their equivalents in the respective standards.
[0314] It is specifically noted that the transaction flow diagrams (TFDs) presented herein are schematic representations, which means that the number, order, timings, dimensions, and other properties of the information illustrated in the TFDs are non-limiting examples. Every modification, variation, or alternative allowed by a current or future Specification mentioned in the TFD (such as CXL, UALink, SUE, PCIe, UCIe, CHI, AXI, etc.) that is relevant to a diagram, is also intended to be included within the scope of said diagrams. Furthermore, the scope of these diagrams extends to encompass implementations that may deviate from the strict specifications mentioned in the TFDs due to factors such as hardware bugs, relaxed designs, or implementation-specific optimizations.
[0315] Sentences in the form of "a port / interface configured to communicate with a host / device" encompass "a port / interface configured to support communication with a host / device", which refer to direct coupling between the port / interface and the host / device, or to indirect coupling between the port / interface and the host / device, such as via one or more intermediate components including but not limited to switches, retimers, redrivers, bridges, and / or protocol translators.
[0316] Herein, terms such as send / sending, receive / receiving, communicate / communicating, or exchange / exchanging when used to describe elements (e.g., computer, RPU, MxPU, processor, semiconductor device, switch, port, interface) involved in data, message, packet, or other information exchanges, may refer to direct or indirect operation(s) that facilitate information transfer to / from / between such elements. When a first element is said to send information to a second element, it is not required to directly transmit the information from the first element to the second element; similarly, when a first element is said to receive information from a second element, the first element is not required to directly obtain the information from the second element. Instead, the elements may initiate, cause, make available, control, direct, participate in, or otherwise facilitate such transfer. The information transfer may occur directly or indirectly utilizing one or more intermediary components, and may include routing, forwarding, or other suitable data transfer mechanisms over suitable communication path and / or connection.
[0317] In a similar manner, when a port / interface is said to send / receive / exchange / communicate information to / from / with another entity (which may be for example a computer, a host, a device, a switch, a port, an interface, an RPU, or a retimer), it is not required to directly send / receive / exchange information with the other entity. Instead, the port / interface may communicate through a suitable intermediate medium, component, or entity that facilitates transfer of the information. Such communication may involve one or more intermediary components, protocols, or mechanisms that encrypt, process, convert, buffer, route, or otherwise handle the information between the port / interface and the other entity.
[0318] Additionally, the terms "port" and "interface" may be used herein interchangeably unless the context requires distinction between them. Depending on the context, the term "port" may refer to physical or logical interface, connection point, access point, or termination point that is configured to support communication with or within components, devices, or systems in a network or computing architecture. A port may include, be included in, or be coupled to various interface types and may support one or more communication protocols. Still depending on the context, the term port may refer to various specialized port types including but not limited to a switch port (e.g., a UALink port may refer to a UALink switch port), a downstream port, an upstream port, a root port, an endpoint port, a device port, a mesh port, a fabric port, or an ISoL port. Depending on the implementation and context, a port may be integrated within a device, may comprise a device interface, or may function as a standalone entity. For example, the following pairs may be used herein interchangeably unless a particular context specifies otherwise: CHI interface and CHI port, CHI-based interface and CHI-based port, NVLink interface and NVLink port, and NVLink-based interface and NVLink-based port.
[0319] Claims in the form of “A non-transitory computer-readable medium comprising instructions which, when executed by a processor, cause the processor to perform the computer-implemented method of claim X” are intended to encompass physical storage media capable of storing instructions, including but not limited to semiconductor memory, magnetic storage, optical storage, and other persistent storage technologies. The instructions may be in anyform capable of directing a processor to perform the method, including but not limited to compiled code, interpreted code, bytecode, firmware, as well as other forms of directives such as natural language directives, declarative specifications, model parameters or configurations, and symbolic representations, among other formats that may be suitable for processing by processors, Al modules, neural processing units, or other current or future processing architectures. The processor may include any processing unit capable of executing or interpreting stored instructions, including but not limited to CPUs, microprocessors, microcontrollers, DSPs, GPUs, neural processing units, Al accelerators, and quantum processing units. The stored instructions may cause a single processor to perform the method, or may cause the processor to coordinate with one or more additional processors to collectively perform the method in a distributed manner.
[0320] Claims in the form of "One or more integrated circuits configured to perform the method of claim X, wherein the one or more integrated circuits comprise at least one of: (i) a general -purpose processing unit, comprising or connected to communication interfaces, configured to perform the method via software and / or firmware execution, (ii) circuitry comprising firmware and / or hardware logic integrated into an electronic device, wherein the circuitry utilizes operations that benefit from hardware acceleration and / or specialized processing capabilities not typically provided by a general-purpose processing unit, or (iii) one or more chiplets within one or more integrated circuit packages" are intended to encompass hardware implementations that execute, implement, realize, or carry out method steps through circuitry, programmable circuitry, stored instructions executed by processing elements, or distributed across multiple chiplets. The first alternative covers implementations based on processing units designed to execute arbitrary software instructions, including but not limited to CPUs, microprocessors, and application processors, that execute software or firmware to perform the method, with communication interfaces enabling data exchange with other system components. The second alternative covers implementations where specialized circuitry provides hardware acceleration or dedicated processing capabilities, including but not limited to ASICs, FPGAs, PLDs, and SoC devices, wherein the functionality is implemented using electronic and / or photonic components, as well as programmable logic. The third alternative covers chiplet -based implementations where the method is performed by one or more semiconductor dies designed for integration within multi-chip modules or system-in-package configurations. These chiplets may reside within a single package or across multiple packages, communicating via inter-chiplet protocols such as UCIe, AIB, CHI-C2C, or other die-to-die interfaces when within the same package, or via package-to-package interfaces when distributed across different packages. The packages may utilize various integration technologies, including but not limited to 2.5D silicon interposers, 3D stacking, organic substrates, and embedded bridge technologies. The method may be partitioned across multiple chiplets with different chiplets implementing different portions, or a single chiplet may implement the complete method.
[0321] Claims in the form of "An active cable comprising first and second pluggable modules coupled by a physical medium; wherein the active cable further comprises hardware circuitry, integrated into the active cable, configured to perform the method of claim X" are intended to encompass cable assemblies that include active electronic components capable of processing and modifying signals during transmission. Such claims cover cables having connectors at each end designed for insertion into corresponding receptacles, connected by a transmission medium that may include copper conductors, optical fibers, or other signal -carrying media. The electronic components performing the method may be incorporated anywhere within the cable assembly, including within either or both of the pluggable connectors, or positioned along the cable between segments of the physical medium. The implementation may utilize fixed circuit arrangements, programmable logic, firmware, or combinations thereof. The electronic components may perform the entire method within the cable or may work in conjunction with other processing elements to implement the complete functionality.
[0322] Claims in the form of "An apparatus configured to operate as a switch, wherein the apparatus comprises switching circuitry and is configured to perform the method of claim X" are intended to encompass apparatus that selectively routes signals, data, or communications between ports while also performing the method. Such claims cover traditional switching devices with dedicated switch ports as well as processor-based switches and other architectures that achieve switching functions through alternative port configurations. The ports through which dataenters or exits the switching function may include physical ports, logical ports, virtual ports, upstream ports, downstream ports, root ports, endpoint ports, host ports, device ports, ingress ports, egress ports, memory -mapped interfaces, coherent interconnect attachment points, and software -defined interfaces. The internal routing mechanisms may include crossbar fabrics, mesh networks, buffering systems, ring or mesh topologies, routing tables, coherent interconnects, and other switching implementations. The apparatus may include homogeneous ports supporting a single protocol or heterogeneous ports supporting different protocols, speeds, or functionalities. The switching may be implemented using store-and-forward, cut-through, adaptive routing, source routing, or other methodologies. The method operations are performed as part of the switching functionality through hardware, firmware, and / or logic contained within the apparatus.
[0323] Accordingly, this disclosure is intended to embrace all such alternatives, modifications, and variations that fall within the spirit and scope of the appended claims and their equivalents.
Claims
WE CLAIM:
1. A system, comprising:a processor comprising a coherent interconnect; the processor is coupled to at least 64GB of memory and is configured to utilize physical addresses within a Host Physical Address (HP A) space to access the memory, and to execute an operating system (OS) that utilizes a virtual address space;a memory management unit (MMU) configured to enable access to the memory based on mapping addresses within the virtual address space to physical addresses within the HPA space; a resource provisioning unit (RPU) comprising a Compute Express Link (CXL) device configured to communicate with an entity according to a protocol based on CXL; and wherein the RPU is further coupled to the coherent interconnect and configured to perform host- to-host physical address translations, whereby the host-to-host physical address translations enable the entity to access the memory via the CXL device.
2. The system of claim 1, wherein the entity utilizes a second HPA space, and the host-to-host physical address translations translate physical addresses within the second HPA space to physical addresses within the HPA space.
3. The system of claim 2, further comprising a CXL Root Port configured to communicate with a CXL memory expander that utilizes a Device Physical Address (DPA) space; and wherein at least one of the operating system, system firmware, or the memory expander is configured to map between physical addresses within the HPA space and physical addresses within the DPA space, which enable the entity to utilize the memory and / or the CXL memory expander.
4. The system of claim 3, wherein the RPU further comprises a second CXL device configured to communicate with a second entity utilizing a second protocol based on CXL, whereby the second entity utilizes a third HPA space; and wherein the RPU is further configured to translate physical addresses within the third HPA space to physical addresses within the HPA space, which enable the second entity to utilize the CXL memory expander.
5. The system of claim 2, wherein the RPU further comprises a second CXL device configured to communicate with a second entity utilizing a second protocol based on CXL, whereby the second entity utilizes a third HPA space, and the RPU is further configured to translate physical addresses within the third HPA space to physical addresses within the HPA space, which enable the second entity to utilize the memory.
6. The system of claim 5, wherein the entity comprises a host coupled to the processor via at least one of a CXL root port or a CXL switch, and the second protocol based on CXL is different from the protocol based on CXL.
7. The system of claim 2, wherein the processor comprises a Modified CPU or GPU (MxPU), the memory comprises dynamic random-access memory (DRAM), and the RPU enables the entity to utilize more than 250GB of the DRAM.
8. The system of claim 1, wherein the memory comprises dynamic random -access memory (DRAM) that is coupled via memory channels to the processor, and the CXU device comprises a Global Fabric-Attached Memory (G-FAM) Device (GFD).
9. The system of claim 1, wherein the protocol based on CXU utilizes CXU.mem semantics, and the CXU device exposes at least one Host-managed Device Memory (HDM) address region to the entity.
10. The system of claim 1, wherein the protocol based on CXU utilized CXU.io semantics, and the host-to-host physical address translation translates from physical addresses carried in CXU.io UIOMRd Transaction Uayer Packets (TUPs) received from the entity to physical addresses within the HPA space.
11. The system of claim 1, wherein the processor comprises multiple cores, from which at least one is a hidden core; and wherein the RPU is further configured to utilize the hidden core for internal tasks, wherein the internal tasks comprise at least one of internal firmware processing, CXU Fabric Manager (FM) API processing, processing in memory (PIM), near-memory processing, or housekeeping tasks.
12. The system of claim 11, wherein the hidden core is isolated from user access and visibility, providing user-infrastructure isolation.
13. The system of claim 1, wherein the processor comprises multiple cores, from which at least one is hidden and is utilized for collection of memory telemetry.
14. The system of claim 1, wherein the processor comprises multiple cores, from which at least one is a hidden core utilized for secure key storage and management for encrypting and decrypting data transmitted according to the protocol based on CXU, leveraging user-infrastructure isolation provided by the hidden core.
15. The system of claim 14, further comprising a hardware -accelerated cryptographic engine, wherein the hidden core is configured to utilize the hardware-accelerated cryptographic engine for performing at least part of the cryptographic operations on the data transmitted according to the protocol based on CXU.
16. The system of claim 14, wherein the hidden core enables support for confidential computing over memory exposed by the RPU via the CXU device; whereby confidential computing performs computation within a secure isolated environment to protect data in use.
17. The system of claim 1, wherein the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core forerror handling and / or correction tasks within a memory pool comprising the memory, enhancing data integrity and reliability.
18. The system of claim 17, wherein the error handling and / or correction tasks further comprise predictive failure analysis (PFA) operations, configured to predict and handle imminent failure of memory components within the memory pool, thereby preempting potential data loss and system downtime.
19. The system of claim 1, wherein the memory comprises dynamic random-access memory (DRAM), and the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core for controlling or managing memory access scheduling within a memory pool comprising the DRAM, to improve memory utilization and throughput.
20. The system of claim 1, wherein the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core for managing security protocols within a memory pool comprising the memory, including data encryption and / or access controls.
21. The system of claim 1, wherein the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core for configuration management tasks within a memory pool comprising the memory, including dynamic allocation and deallocation of memory resources.
22. The system of claim 1, wherein the processor comprises multiple cores, from which at least one core is a hidden core; and wherein the RPU is further configured to utilize the hidden core for memory tiering tasks.
23. The system of claim 22, wherein the memory tiering tasks further comprise migration of data between memory tiers based on hotness level of the data, thereby increasing performance of memory accesses from the entity to hot data.
24. The system of claim 22, further comprising a direct Memory Access (DMA) engine, wherein the hidden core is configured to utilize the DMA engine for migrating data between memory tiers.
25. A method, comprising:accessing memory coupled to a processor utilizing physical addresses within a Host Physical Address (HP A) space; wherein the processor comprises a coherent interconnect; mapping addresses within a virtual address space to physical addresses within the HPA space; whereby the addresses within the virtual address space are utilized by an operating system (OS) of an apparatus comprising the processor;communicating, by a Compute Express Link (CXL) device of a resource provisioning unit (RPU), with an entity coupled to the apparatus according to a protocol based on CXL; wherein the RPU is coupled to the coherent interconnect; andperforming, by the RPU, host-to-host physical address translations which enable the entity to access the memory via the CXL device.
26. The method of claim 25, wherein the entity comprises a second host that utilizes a second HPA space, and the host-to-host physical address translations are translating physical addresses within the second HPA space to physical addresses within the HPA space.
27. The method of claim 26, further comprising communicating, via a CXL Root Port, with a CXL memory expander that utilizes a Device Physical Address (DPA) space; and wherein at least one of the operating system or system firmware is mapping between physical addresses within the HPA space and physical addresses within the DPA space, whereby the mapping enables the second host to utilize the memory and / or the CXL memory expander.
28. An apparatus, comprising:a processor comprising a coherent interconnect; the processor is coupled to at least 64GB of memory and is configured to utilize physical addresses within a first Host Physical Address (HPA) space to access the memory, and to execute an operating system (OS) that utilizes a virtual address space;a memory management unit (MMU) configured to enable access to the memory, based on mapping addresses within the virtual address space to physical addresses within the first HPA space;a resource provisioning unit (RPU), coupled to a Compute Express Link (CXL) device configured to exchange messages conforming to a protocol based on CXL which utilizes a second HPA space; andwherein the RPU is further coupled to the coherent interconnect and configured to translate physical addresses within the second HPA space to physical addresses within the first HPA space.
29. A system designed to function as a Multi-Headed Device (MHD), comprising:a processor comprising a coherent interconnect; the processor is coupled to at least 32GB of dynamic random-access memory (DRAM), and is configured to utilize physical addresses within a Host Physical Address (HPA) space to access the DRAM, and to execute an operating system (OS) that utilizes a virtual address space;a memory management unit (MMU) configured to enable access to the DRAM, based on mapping addresses within the virtual address space to physical addresses within the HPA space; first and second Compute Express Link (CXL) Endpoints configured to communicate with hosts coupled to the system according to a protocol based on CXL; anda resource provisioning unit (RPU) configured to perform host-to-host physical address translations which enable the hosts to access the DRAM utilizing messages conforming to the protocol based on CXL.
30. The system of claim 29, wherein the DRAM is coupled via at least four memory channels to the processor; wherein the DRAM has a memory capacity exceeding 128 GB, 256 GB, 512 GB, or 1 TB; and wherein the DRAM comprises mainstream DRAM modules exhibiting an average unit price per gigabyte that does not exceed three times an average unit price per gigabyte of a lowest-cost DRAM module technology in volume production for servers in data centers.