Systems and methods for multi-link CXL switches for NUMA architectures

By generating virtual nodes and dynamically allocating memory through multi-link CXL switches, the problems of remote memory access latency and redundancy management in NUMA architecture are solved, achieving efficient memory resource utilization and seamless access, and improving system performance.

CN122073573APending Publication Date: 2026-05-22SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-11-19
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

In NUMA architectures, remote memory access leads to increased latency. Existing technologies cannot effectively manage redundant memory and efficiently allocate memory resources, resulting in inefficient resource utilization and potential conflicts.

Method used

By employing a multi-link CXL switch, virtual nodes are generated and the physical addresses of data in memory are identified based on the offset of the physical node address range. By combining virtual nodes and dynamic memory allocation, memory access and resource utilization are optimized.

Benefits of technology

It reduces latency in NUMA architecture, enables efficient resource utilization and seamless memory access, prevents redundancy and conflicts, and improves system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122073573A_ABST
    Figure CN122073573A_ABST
Patent Text Reader

Abstract

A system and method for managing memory in a computing system is disclosed. The method includes generating a virtual node by combining two or more physical nodes coupled to a compute express link (CXL) switch; and identifying a physical address of data stored in the memory based on an offset between address ranges of the two or more physical nodes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] priority

[0002] This application is based on and claims priority to U.S. Provisional Patent Application Serial No. 63 / 722,849, filed November 20, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to memory management in non-uniform memory access (NUMA) architectures, and more specifically, to systems and methods for optimizing memory access and resource allocation using multi-link compute fast link (CXL) switches. Background Technology

[0004] NUMA architectures can be used in high-performance computing systems to manage memory resources across multiple central processing unit (CPU) sockets. In such architectures, memory access latency can vary significantly depending on whether the memory being accessed is local to the CPU socket of the executing process or resides in a remote socket. To address this variability, techniques such as CXL have been developed to facilitate high-speed, consistent access to memory resources across distributed systems.

[0005] The CXL host adapter can connect to the CPU socket and interface with the CXL memory expander via the CXL switch. While this configuration provides scalability and efficient resource sharing, it can lead to increased latency when the CPU socket accesses memory via a remote adapter. This latency variation is particularly noticeable in workloads requiring frequent memory access, as the time spent accessing remote memory can impact overall system performance.

[0006] To optimize memory access in NUMA architectures, it may be necessary to address challenges such as redundant memory exposure and efficient allocation of memory resources. These challenges are further complicated when multiple CXL host adapters are used to connect to the same CXL memory expander, resulting in overlapping memory regions exposed across multiple physical nodes. Existing operating systems and memory management frameworks often do not adequately account for this redundancy, leading to inefficient resource utilization and potential conflicts.

[0007] It should be understood that the background section provided is for the purpose of describing the general motivation and background of the invention only. The discussion herein is intended to enhance understanding and should not be construed as an admission or endorsement of the prior art. Summary of the Invention

[0008] The embodiments disclosed herein enable the use of multi-link CXL switches to reduce latency in NUMA architectures. Virtual nodes and dynamic memory allocation provide efficient resource utilization, while migration between nodes maintains seamless memory access.

[0009] According to an embodiment, a method for managing memory in a computing system includes: generating virtual nodes by combining two or more physical nodes coupled to a CXL switch; and identifying the physical address of data stored in the memory based on an offset between address ranges of the two or more physical nodes.

[0010] According to another embodiment, an apparatus for managing memory in a computing system includes a CXL switch configured to couple two or more physical nodes, and a processor. The processor is configured to generate virtual nodes by combining the two or more physical nodes; and to identify the physical address of data stored in the memory based on an offset between address ranges of the two or more physical nodes. Attached Figure Description

[0011] The above and other aspects, features and advantages of certain embodiments of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:

[0012] Figure 1 A CXL memory system according to an embodiment is shown;

[0013] Figure 2 An enhanced CXL memory system according to an embodiment is shown;

[0014] Figure 3A This is a memory allocator and node management design according to an embodiment;

[0015] Figure 3B This is an enhanced memory allocator and node management design according to an embodiment;

[0016] Figure 4 This is a flowchart illustrating the node initialization process for a multi-link CXL architecture according to an embodiment;

[0017] Figure 5A This is a memory allocator and node management design according to an embodiment, illustrating the allocation of pages from physical memory;

[0018] Figure 5B This is an enhanced memory allocator and node management design according to an embodiment, illustrating the allocation of pages from physical memory;

[0019] Figure 6A This is an enhanced memory allocator and node management design according to an embodiment, illustrating the allocation of pages from physical memory;

[0020] Figure 6B This is a flowchart illustrating the function of the memory allocator when a virtual node is used upon receiving a memory allocation request, according to an embodiment.

[0021] Figure 7 This is a diagram illustrating the implementation of a CXL memory expander using a unique address space according to an embodiment;

[0022] Figure 8 This is a diagram illustrating the use of multiple PTEs to manage memory allocation to support process migration according to an embodiment;

[0023] Figure 9 This is a flowchart illustrating a method for managing memory in a computing system according to an embodiment; and

[0024] Figure 10 This is a diagram illustrating a storage system according to an embodiment. Detailed Implementation

[0025] In the following description, embodiments of the present disclosure are described in detail with reference to the accompanying drawings. It should be noted that the same elements will be designated by the same reference numerals, although they are shown in different drawings. In the following description, only specific details such as detailed configurations and components are provided to aid in the overall understanding of the embodiments of the present disclosure. Therefore, those skilled in the art will understand that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Additionally, for clarity and brevity, descriptions of well-known functions and structures have been omitted. The terminology described below is defined in consideration of the functions in this disclosure and may vary depending on the user, the user's intent, or habits. Therefore, the definitions of the terms should be determined based on the entire contents of this specification.

[0026] This disclosure can have various modifications and embodiments, which are described in detail below with reference to the accompanying drawings. However, it should be understood that this disclosure is not limited to the embodiments, but includes all modifications, equivalents, and substitutions within the scope of this disclosure.

[0027] Although various elements may be described using ordinal terms such as first, second, etc., structural elements are not limited by these terms. These terms are used only to distinguish one element from another. For example, a first structural element may be referred to as a second structural element without departing from the scope of this disclosure. Similarly, a second structural element may also be referred to as a first structural element. As used herein, the term "and / or" includes any and all combinations of one or more related items.

[0028] The terminology used herein is for describing various embodiments of this disclosure only and is not intended to limit the disclosure. Unless the context clearly indicates otherwise, the singular form is intended to include the plural form. In this disclosure, it should be understood that the terms “comprising” or “having” indicate the presence of features, numbers, steps, operations, structural elements, components, or combinations thereof, and do not preclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, structural elements, components, or combinations thereof.

[0029] Unless otherwise defined, all terms used herein shall have the same meaning as understood by one of ordinary skill in the art to which this disclosure pertains. Terms such as those defined in common dictionaries shall be interpreted as having the same meaning as in the context of the relevant technical field and shall not be interpreted as having an ideal or overly formal meaning unless expressly defined in this disclosure.

[0030] According to one embodiment, the electronic device can be one of various types of electronic devices utilizing storage devices. The electronic device can use any suitable storage standard, such as Peripheral Component Interconnect Rapid (PCIe), Non-Volatile Memory Rapid (NVMe), Architectural NVMe (NVMeoF), Advanced Scalable Interface (AXI), Hyper Path Interconnect (UPI), Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), Remote Direct Memory Access (RDMA), RDMA over Converged Ethernet (ROCE), Fibre Channel (FC), Infinite Bandwidth (IB), Serial Advanced Technology Attachment (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), Internet Wide Area RDMA Protocol (iWARP), etc., or any combination thereof. In some embodiments, the interconnect interface can be implemented using one or more memory semantics and / or memory coherence interfaces and / or protocols, including one or more CXL protocols, such as CXL.mem, CXL.io and / or CXL.cache, Gen-Z, Coherent Accelerator Processor Interface (CAPI), Cache Coherent Interconnect for Accelerators (CCIX), etc., or any combination thereof. Any memory device can be implemented using one or more of any type of memory device interface, including Double Data Rate (DDR), DDR2, DDR3, DDR4, DDR5, Low Power DDR (LPDDRX), Open Memory Interface (OMI), NVlink High Bandwidth Memory (HBM), HBM2, HBM3, etc. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computers, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. However, electronic devices are not limited to those listed above.

[0031] The terminology used in this disclosure is not intended to limit the disclosure, but rather to include various changes, equivalents, or substitutions for the corresponding embodiments. Regarding the description of the drawings, similar reference numerals may be used to refer to similar or related elements. Unless the relevant context clearly indicates otherwise, the singular form of a noun corresponding to an item may include one or more things. As used herein, each of the phrases such as “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B, or C,” “at least one of A, B, and C,” and “at least one of A, B, or C” may include all possible combinations of items listed together in the corresponding phrase. As used herein, terms such as “first,” “second,” “first,” and “second” may be used to distinguish a corresponding component from another component, but are not intended to limit the component in other respects (e.g., importance or order). When the terms “operably” or “communically” are used, or when the terms “operably” or “communically” are not used, if an element (e.g., a first element) is referred to as “coupled to another element (e.g., a second element),” “coupled to another element (e.g., a second element),” “connected to another element (e.g., a second element),” or “connected to another element (e.g., a second element),” it means that the element can be directly (e.g., wired) connected to the other element, wirelessly connected to the other element, or connected to the other element via a third element.

[0032] As used herein, the term "module" can include units implemented in hardware, software, firmware, or a combination thereof, and is used interchangeably with other terms such as "logic," "logic block," "part," and "circuit." A module can be a single integrated component adapted to perform one or more functions, or its smallest unit or part. For example, according to one embodiment, a module can be implemented as an application-specific integrated circuit (ASIC), a coprocessor, or a field-programmable gate array (FPGA).

[0033] Traditional NUMA architectures suffer from increased latency when accessing memory from remote CPU sockets. This limitation arises from the lack of localized memory access paths and inefficient memory management across multiple nodes. CXL is an interconnect and protocol designed to provide high-speed, consistent access to memory and accelerators, thereby enabling improved performance in distributed computing systems.

[0034] Figure 1 A CXL memory system according to an embodiment is shown.

[0035] refer to Figure 1System 100 includes two CPU sockets 101 and 102, two PCIe switches 103 and 104, and a CXL host adapter 105. The CXL host adapter 105 connects to a shared CXL switch 106, which uses a Virtual CXL Switch (VCS) unit to facilitate access to one or more CXL memory expanders 107a, 107b, and / or 107n. The VCS unit can partition one or more CXL switches into virtual switches, allowing individual hosts to manage and access shared memory resources. This architecture enables both CPU sockets 101 and 102 to access one or more CXL memory expanders 107a, 107b, and / or 107n via local connections or via remote socket connections.

[0036] Figure 1 Distinguishing between local and remote memory access paths. Local memory access occurs when the CPU socket communicates directly with its corresponding CXL host adapter to access CXL memory expanders via a shared CXL switch, reducing latency. Conversely, remote access occurs when the CPU accesses memory via a CXL host adapter from another socket, resulting in higher latency due to additional interconnect traversal. Therefore, CPU 101 can locally access one or more CXL memory expanders 107a, 107b, and / or 107n via CXL switch 106 and CXL host adapter 105, as these components are downstream of CPU 101's PCIe switch 103. Meanwhile, CPU 102 can only remotely access those same resources, operated via CPU 101.

[0037] Figure 1 Some architectural limitations are shown, where processes executing on a single CPU socket (e.g., CPU 102) experience increased latency when accessing memory resources via a remote socket (e.g., CPU 101).

[0038] Figure 2 An enhanced CXL memory system according to an embodiment is shown.

[0039] refer to Figure 2System 200 includes two CPU sockets 201 and 202, two PCIe switches 203 and 204, and two CXL host adapters 205 and 206. CXL host adapters 205 and 206 are connected to a shared CXL switch 207, which facilitates access to one or more CXL memory expanders 208a, 208b, and / or 208n using VCS units. This architecture enables the two CPU sockets to access one or more CXL memory expanders 208a, 208b, and / or 208n via a low-latency local connection. Specifically, both CPUs 201 and 202 have their own corresponding CXL adapters 205 and 206, which are local to their respective PCIe switches 203 and 204.

[0040] Figure 2 The CXL memory system 200 is enhanced with multi-link capabilities, where each CPU socket is equipped with a dedicated CXL host adapter. This configuration ensures that both CPU 201 and CPU 202 can locally access one or more CXL memory expanders 208a, 208b, and / or 208n via a shared CXL switch 207. System 200 can allow processes running on either CPU to access memory with reduced latency by routing the memory access via the local CXL host adapter.

[0041] The boxes labeled “VCS 0”, “VCS 1”, “VCS n-1”, and “VCS n” in CXL switch 207, along with the lines labeled “Shared” connecting them to one or more CXL memory expanders 208a, 208b, and / or 208n, represent the concept of an enhanced VCS system that implements memory sharing in a multi-link CXL switch 207. A VCS unit can refer to a logical entity within a physical CXL switch that creates a separate memory tier for each connected host, allowing each VCS unit to independently access its assigned CXL host adapter (e.g., 205, 206) as if it were directly attached to the host, isolating its memory space and providing efficient memory management across multiple systems (hosts or applications).

[0042] as Figure 1 The VCS unit shown, Figure 2 These VCS units are logical entities within the CXL switch that facilitate the sharing of memory resources within one or more CXL memory expanders 208a, 208b, and / or 208n. However, with Figure 1 The VCS unit depicted in the text is the opposite. Figure 2The VCS unit in the VCS is designed to facilitate efficient memory sharing among multiple CPUs 201 and 202 using one or more CXL memory expanders 208a, 208b, and / or 208n. Figure 1 The difference lies in Figure 1 In this context, memory access may involve remote communication paths. Figure 2 The architecture shown includes dedicated CXL host adapters 205 and 206 for each CPU 201 and 202, respectively, allowing each CPU 201 and 202 to establish a direct connection to the VCS unit within the CXL switch 207. This localized access mechanism enables each CPU 201 and 202 to retrieve memory directly (locally) from one or more CXL memory expanders 208a, 208b, and / or 208n, without having to use a remote connection path via a remote CXL host adapter. As a result, Figure 2 The system 200 significantly reduces or eliminates latency associated with remote memory access.

[0043] Figure 2 The “shared” line indicates that the memory resources in one or more CXL memory expanders 208a, 208b, and / or 208n are not statically allocated, but dynamically shared among VCS units. This property allows multiple CPUs to directly (locally) access the same memory regions without conflicts or redundancy, using CXL’s cache coherence protocol to maintain data consistency.

[0044] By equipping both CPU sockets with CXL host adapters Figure 2 The architecture minimizes the CPU dependency on remote CXL host adapters for memory access, thus reducing the likelihood of remote memory access and improving overall latency. Additionally, System 200 introduces mechanisms at the software layer to manage shared memory resources and prevent redundant memory exposure. This configuration can scale beyond multiple CPUs and supports the allocation and migration of memory resources across nodes.

[0045] When two CXL host adapters (e.g., 205 and 206) are connected to the same CXL memory expander (e.g., 208a, 208b, or 208n), the memory can be redundantly exposed as multiple nodes with different physical addresses. This redundant exposure complicates memory management and increases the likelihood of resource conflicts. To address this issue, the memory allocator can operate on a per-virtual-node basis, merging redundant physical memory regions into a single virtual node. This virtual node abstraction allows multiple physical nodes referencing the same underlying memory medium to be managed as a unified entity. Therefore, the term "physical node" (e.g., Figure 3A Nodes 311a-314a in the middle Figure 3B Nodes 311b-314b in the middle Figure 5A Nodes 511a-514a in the middle Figure 5B Nodes 511b-514b in the and Figure 6A Nodes 611a-614a in the table refer to the mapping of physical address ranges associated with a specific CXL host adapter. See the reference below. Figures 3A-3B , Figures 5A-5B and Figure 6A The physical nodes discussed can be redundant representations of the same physical memory region, which may lead to inefficient memory allocation.

[0046] Another challenge arises during process migration between nodes. When a process migrates from one node to another, the system can update memory addresses to reflect the memory mapping of the local node. Failure to update addresses can cause processes to access memory via remote nodes, introducing unnecessary latency and negating the benefits of a multi-link architecture.

[0047] This disclosure describes a method for advanced node and memory management using virtual nodes. A virtual node can be a logical entity that manages memory resources by combining or partitioning physical nodes based on shared or overlapping memory regions. Virtual nodes can achieve efficient memory allocation and prevent redundancy by treating multiple physical nodes with overlapping memory as a unified node in a logical memory map.

[0048] This method addresses the challenges associated with redundant memory regions in multi-link CXL architectures, where multiple physical nodes may expose overlapping memory regions due to the presence of multiple CXL host adapters. By creating virtual nodes, the system can merge or partition physical nodes to logically manage memory resources and reduce redundancy.

[0049] Figure 3A This is a memory allocator and node management design based on an embodiment.

[0050] refer to Figure 3A In one method of node management, logical nodes 301a, 302a, 303a, and 304a can be directly mapped to physical nodes 311a, 312a, 313a, and 314a, regardless of redundant or overlapping memory regions. The mapping from logical nodes to physical nodes can occur on the CPU (e.g., ...). Figure 2 Inside CPU 201 or CPU 202.

[0051] As shown in the figure, the four physical nodes 311a, 312a, 313a, and 314a directly correspond to the logical nodes 301a, 302a, 303a, and 304a, respectively. The memory allocator 300a can operate independently for each logical node 301a, 302a, 303a, and 304a, which allows redundant memory regions to be managed individually.

[0052] However, physical nodes 313a and 314a both correspond to the same underlying CXL memory 306a (the term "CXL memory" can be used interchangeably with "CXL memory extender" and "CXL memory region"). Since logical nodes 303a and 304a are mapped 1:1 to physical nodes 313a and 314a, memory allocator 300a independently and redundantly tracks logical nodes 303a and 304a to manage what is physically a single shared CXL memory resource, represented by CXL memory 306a. For example, CXL memory 306a may include 64 GB of physical memory, but due to the redundant exposure of physical nodes 313a and 314a, it appears to memory allocator 300a as two separate 64 GB memory regions. As a result, the memory may appear to be 128 GB of total system memory even when only 64 GB of physical memory actually exists. Therefore, memory allocator 300a operating under this configuration may treat overlapping memory regions as memory regions with different physical addresses, leading to inefficient memory utilization.

[0053] Figure 3B This is an enhanced memory allocator and node management design based on an embodiment.

[0054] refer to Figure 3B Logical nodes 301b and 302b are directly mapped to physical nodes 311b and 312b, similar to... Figure 3A The mapping of logical nodes 301a and 302a and physical nodes 311a and 312a in the system. However, with Figure 3A The methods are different. Figure 3B The memory allocator 300b identifies overlapping memory regions between physical nodes 313b and 314b and merges them into a single virtual node 305b. The virtual node 305b enables the host to manage access to the same CXL memory region through different address spaces associated with different CXL host adapters.

[0055] Specifically, overlapping physical nodes 313b and 314b, redundantly mapped to the same CXL storage region 306b, are combined into a single virtual node 305b. By introducing the virtual node 305b, the memory allocator 300b manages the shared memory as a single, unified resource. This prevents redundant memory allocation that occurs when the same physical memory regions are managed independently, such as... Figure 3A As shown.

[0056] Therefore, in this embodiment, the memory allocator 300b can be reconfigured to manage virtual nodes (e.g., 305b) instead of logical nodes that directly correspond to physical nodes, which allows the system to treat overlapping memory regions as a unified entity.

[0057] Figure 4 This is a flowchart illustrating the node initialization process for a multi-link CXL architecture according to an embodiment.

[0058] refer to Figure 4 The process begins in step 401, where the system firmware (such as the Basic Input / Output System (BIOS) or Unified Extensible Firmware Interface (UEFI)) detects available memory nodes in the system. These firmware components can identify physical memory blocks and collect metadata about their configuration. In step 402, the system initializes the NUMA node table using platform-specific information provided by system tables (such as the System Resource Affinity Table (SRAT), System Locality Information Table (SLIT), and / or CXL Early Discovery Table (CEDT)). These tables can provide details about memory topology, locality, and interconnection relationships.

[0059] In step 403, the system constructs a set of memory blocks corresponding to physical nodes. Each block can represent a contiguous memory region belonging to a single physical node and can correspond to the entire memory device or a subdivided region of the memory region. In step 404, the first memory block is retrieved, and the system begins evaluating its status. In step 405, a check is performed to determine whether all detected memory blocks have been processed and registered in the NUMA node table. If all memory blocks are registered, the initialization process ends. However, if there are remaining unprocessed memory blocks, in step 406, the system evaluates whether the current memory block resides in an overlapping region. Overlapping regions can occur when two or more physical nodes are mapped to the same physical memory region due to redundant CXL host adapter connections.

[0060] If an overlapping region is detected, in step 407, the system further checks whether the memory block is completely contained within the overlapping region. For blocks that are not completely overlapping, in step 408, the system divides the block into smaller sub-blocks to handle the overlap more precisely. For completely overlapping blocks, in step 409, the system processes the block without further division and determines whether the memory block is already registered in the NUMA node table. If the block is already registered, in step 410, the system moves to the next unprocessed memory block. If not, in step 411, the system creates a NUMA node table to register the memory block (e.g., associating a virtual node with the overlapping physical node). The system retrieves the next unprocessed memory block in step 411 and repeats this sequence of steps until all blocks are registered ("Yes" in step 405). Once all memory blocks have been processed, the node initialization process ends.

[0061] Figure 5A This is a memory allocator and node management design according to one embodiment, illustrating the allocation of pages from physical memory.

[0062] refer to Figure 5A In this configuration, each logical node 501a, 502a, 503a, and 504a maintains an independent data structure to track free memory pages within its boundaries. For example, each of logical nodes 501a-504a corresponds to physical nodes 511a-514a, which can be mapped to the same physical memory (e.g., the same CXL memory region).

[0063] and Figure 3A The situation is very similar in that, since logical nodes 503a and 504a independently correspond to physical nodes 513a and 514a, physical nodes 513a and 514a are mapped to... Figure 5A The same CXL memory 506a in the memory leads to redundant exposure of identical or overlapping memory (memory pages in CXL memory 506a). In this case, memory allocator 500a can treat the same memory page as if it exists at two different physical addresses, because pages in CXL memory 506a can be accessed through different logical nodes that are each associated with different physical address mappings. Figure 5A This situation is illustrated by showing that the same page is exposed to both logical nodes 503a and 504a.

[0064] This redundancy presents a challenge in memory management architectures. Because the memory allocator 500a lacks visibility into the overlapping nature of mappings, it may allow two separate programs operating on different logical nodes to use the same physical memory pages under the flawed assumption that they are accessing different memory regions. Without any mechanism to detect or coordinate this overlap, programs may each write to the same underlying memory, leading to inconsistent states or memory corruption. The conflict arises because the same memory pages are reachable through different physical address ranges, and the memory allocator 500a interprets these physical address ranges as independent when they actually involve the same shared resource.

[0065] Figure 5B This is an enhanced memory allocator and node management design according to an embodiment, illustrating the allocation of pages from physical memory.

[0066] refer to Figure 5B Logical nodes 501b and 502b are directly mapped to physical nodes 511b and 512b, similar to... Figure 5A The mapping of logical nodes 501a and 502a and physical nodes 511a and 512a in the system. However, with Figure 5A The methods are different. Figure 5B The memory allocator 500b identifies overlapping memory regions between physical nodes 513b and 514b and merges them into a single virtual node 505b.

[0067] Specifically, overlapping physical nodes 513b and 514b, redundantly mapped to the same pages in CXL storage area 506b, are combined into a single virtual node 505b. By introducing the virtual node 505b, the memory allocator 500b manages the shared memory as a single, unified resource. This prevents redundant memory allocations that occur when the same pages are managed independently, such as... Figure 5A As shown.

[0068] This enhanced design offers several advantages. By merging overlapping regions into virtual nodes, the system can prevent conflicts and reduce the complexity of memory management. This approach can be beneficial in multi-link CXL systems where multiple CXL host adapters can expose overlapping regions of the CXL memory expander. The enhanced memory allocator can provide a scalable solution for high-performance computing systems, offering consistent and conflict-free memory allocation.

[0069] Figure 6A This is an enhanced memory allocator and node management design according to an embodiment, illustrating the allocation of pages from physical memory.

[0070] refer to Figure 6AThe reference numerals 600a, 601a, 602a, 605a, 611a, 612a, 613a, 614a, and 606a in the attached drawings can respectively correspond to Figure 5B The reference numerals 500b, 501b, 502b, 505b, 511b, 512b, 513b, 514b and 506b in the figures refer to these components, and similar descriptions and functions apply to them.

[0071] Unlike some memory allocators that can directly return physical addresses, embodiments of this disclosure, when allocating memory from a virtual node, may return an offset instead of a physical address. The offset can represent a location within the address space of the virtual node and allows the system to determine the physical memory address based on the physical node to which the process is running. For example, the memory allocator can add the base address of the physical node to the offset to calculate the final physical address. Therefore, this mechanism ensures that memory allocated from a virtual node can be accessed from more than one physical node mapped to the virtual node.

[0072] Figure 6B This is a flowchart illustrating the function of the memory allocator when a virtual node is used upon receiving a memory allocation request, according to an embodiment.

[0073] refer to Figure 6B The process begins in step 601b, where a memory allocation request is received, specifying a node identifier (ID) and size. In step 602b, the allocator retrieves the offset of the requested memory from the node identified by the node ID and size. The allocator can use offsets to manage free memory pages, meaning that a free list stores and returns offset values ​​relative to a base address instead of the full physical address. The free list can be used to track available memory blocks and is configured to return offset values ​​when allocating memory pages.

[0074] In step 603b, the allocator then determines whether the node ID matches a virtual node. If the node ID matches a virtual node, in step 604b, the physical memory address is calculated by adding an offset to the base address of the current virtual node (identified by the node ID and size). The node ID can represent the node where a process is running and therefore does not necessarily need to be stored. Instead, a metadata structure (e.g., struct node) can maintain information about virtual nodes, allowing the allocator to determine whether a given node ID corresponds to a virtual node.

[0075] If the node ID differs from the virtual node, the physical memory address is determined in step 605b by adding an offset to the base address of the physical node. In step 606b, the allocator sends the determined physical address to the requester, who can then update the process's page table entry (PTE). The stored data can then be retrieved using the determined physical address. Therefore, by returning an offset instead of a physical address, the system maintains compatibility with processes running on different physical nodes.

[0076] Figure 7 This is a diagram illustrating the implementation of a CXL memory expander using a unique address space according to an embodiment.

[0077] refer to Figure 7 Each node, CPU node 701 and CPU node 702, corresponds to a process running on the associated CPU, denoted as process A and process B, respectively. The memory management system can use virtual address 703 for process A and virtual address 704 for process B, each of which is mapped to physical address 705. These virtual addresses are resolved to physical addresses through a multi-level page table hierarchy managed by the memory management unit (MMU). This hierarchy, following a format such as that used in the x86-64 architecture, may include the Page Global Directory (PGD), Page Upper Directory (PUD), Page Middle Directory (PMD), and Page Entrance Table (PTE), which together resolve virtual addresses to physical addresses.

[0078] PGD, PUD, PMD, and PTE form a hierarchical translation mechanism that progressively narrows the virtual address range. When accessing a virtual address, the most significant bit of the virtual address is used to index the PGD to locate the correct PUD. The PGD points to the PUD, which divides the high-level virtual address space into manageable regions to help isolate large memory segments across different processes. The PUD stores a pointer to the PMD, which provides further granularity by allowing selection within smaller regions. The PMD determines which PTE table contains the final mapping of the virtual address. The PTE table consists of smaller memory regions than the PMD, further improving granularity. Additionally, the PMD can also serve as a control point for changing the path to physical memory resources (e.g., the CXL memory expander 708) without modifying the entire page table hierarchy (e.g., without modifying the PGD and PUD).

[0079] For example, when a process migrates from one CPU node to another, the underlying physical memory it accesses can remain the same (e.g., CXL memory extender 708), but the physical addresses used to reach that memory can differ depending on which CXL host adapter (CHA) is local to that node. Instead of rebuilding or rewriting the entire page table (PGD, PUD, PMD, and PTE), the system can redirect translations by modifying PMD entries to point to different page tables (different PTEs) that contain valid mappings to the new node's local CHA address space. Therefore, the system requires fewer page table rewrites to access the same physical memory regions across different CPU nodes. This redirection mechanism avoids address conflicts by ensuring that each CPU node accesses shared memory through PTE pages that reflect its local physical address space.

[0080] CHA 706 and CHA 707 can maintain a unique physical address space for the CXL memory expander 708. This allows overlapping memory regions in the CXL memory to be exposed differently to each node, since each CXL host adapter is local to that node (e.g., CHA 706 is local to node 701, and CHA 707 is local to node 702). For example, a memory region exposed to CHA 706 can be accessed via one physical address, while the same memory region exposed to CHA 707 can be accessed via a different physical address. This ensures that each node accesses memory through its local CXL host adapter, thereby minimizing latency and optimizing performance.

[0081] This memory configuration ensures that each process uses the appropriate physical address corresponding to its local CXL host adapter. For example, process A running on node 701 resolves its virtual address to the physical address exposed via CHA 706, while process B running on node 702 resolves its virtual address to the physical address exposed via CHA 707. This approach avoids conflicts and ensures efficient memory access across nodes.

[0082] in addition, Figure 7 The CR3 register is shown; it is a system control register that includes the physical address of the page directory and points to the base address of the paging hierarchy in each node, enabling the CPU to efficiently translate virtual addresses for each process.

[0083] therefore, Figure 7This scenario depicts two distinct processes, A and B, running independently on two separate CPU nodes, 701 and 702, each accessing a shared CXL memory extender 708. Processes A and B utilize different PTEs, pointing to different physical address ranges corresponding to the same underlying CXL memory extender 708. Even though processes A and B access the same physical memory, they do so using different addresses due to exposure via separate host adapters (706 and 707). This arrangement allows processes running on separate CPU nodes to independently manage and access memory via localized paths.

[0084] Figure 8 This is a diagram illustrating the use of PTE to manage memory allocation to support process migration according to an embodiment.

[0085] and Figure 7 compared to, Figure 7 This illustrates two independent processes accessing a shared CXL memory from separate CPU nodes 701 and 702, while Figure 8 The scenario illustrates a single process (process C) migrating from one CPU node 801 to another CPU node 802. To ensure that process C continues to access the same memory region via the local host adapter after the migration, the system uses dual PTEs corresponding to the same memory pages in the CXL memory expander 808, located at different physical addresses exposed by CHA 805 and CHA 806, respectively.

[0086] refer to Figure 8 The diagram illustrates a hierarchical paging structure consisting of PGD, PUD, PMD, and multiple PTEs. These tables are managed by the MMU, which translates virtual addresses (803) into physical addresses (804). The translation is performed in stages: PGD provides a high-level partition of the virtual address space, where each entry points to a PUD that further subdivides the address range. PUD, in turn, points to the PMD table. The PMD then points to the PTE table, which resolves virtual addresses to specific physical addresses.

[0087] exist Figure 8 In the example, the system maintains dual PTEs. The base address of each PTE is identified along paths 803a and 803b. Dual PTEs represent two separate address ranges corresponding to two different CXL host adapters, CHA 805 and CHA 806. These host adapters provide access to shared memory regions within the CXL memory expander 808. Although the PTEs pointed to by 803a and 803b consist of different physical address ranges, both ultimately map to the same physical memory page 807. CHA 805 is local to CPU node 801, while CHA 806 is local to CPU node 802.

[0088] To support CXL memory allocation, a pair of memory pages (e.g., 8KB in total) can be reserved for the last-level paging structure. This last-level structure may include a PTE for 4 KB pages of normal size, a PMD for 2 megabyte (MB) large pages, and a PUD for 1 gigabyte (GB) very large pages. When a program allocates memory within the CXL memory expander 808, the associated PTE is initialized such that one entry (e.g., a PTE from 803a) corresponds to the base address used by CHA 805, and a second entry (e.g., a PTE from 803b) corresponds to the base address used by CHA 806. The second entry can be calculated by applying a known offset between the two adapter address ranges.

[0089] During execution, if process C migrates from CPU node 801 to CPU node 802, the system can update the corresponding PMD entry to direct the PTE mapped to the local CXL host adapter (e.g., CHA 805). This update can be triggered by detecting a change in the executing CPU node and can be performed by adjusting the PMD entry to point to the new base address, for example by adding or subtracting a fixed offset (e.g., ±4KB) so that the PMD entry points to reference path 803a instead of 803b. This redirection ensures that subsequent memory accesses issued by process C occur through the local adapter by isolating the exposure of physical addresses at the PTE level.

[0090] therefore, Figure 8 A migration-aware memory conversion mechanism is demonstrated, which uses dual PTEs and dynamic PMD updates to maintain a valid access path to shared memory. By aligning memory accesses with the node's local host adapter (CHA 806 in the post-migration scenario), the system reduces interconnect traffic, avoids remote memory accesses, and maintains consistency between CPU nodes using CXL's cache coherence protocol.

[0091] Figure 9 This is a flowchart illustrating a method for managing memory in a computing system according to an embodiment.

[0092] Figure 9 The method shown can be implemented by a processor, memory controller, system-on-a-chip (SoC), or other processing unit capable of managing virtual memory. In some embodiments, the method is executed by system software (e.g., an operating system (OS)) running on a general-purpose CPU.

[0093] refer to Figure 9In step 901, the virtual node is generated by combining two or more physical nodes coupled to the CXL switch. For example, a processor implementing this method can identify overlapping memory regions exposed to two nodes and assign them to the virtual node.

[0094] In step 902, the physical address of the data stored in memory is identified based on the offset between the node address ranges. This can be achieved by maintaining dual PTE pages, where the base address for the second node mapping is derived by adding a fixed offset (e.g., ±4 KB) to the first node. The PMD entry can be updated during execution to select the PTE page based on the CPU node being executed.

[0095] Figure 10 This is a diagram illustrating a storage system according to an embodiment.

[0096] refer to Figure 10 Storage system 1000 includes a host 1001 and a storage device 1002. Although one host and one storage device are depicted, storage system 1000 may include multiple hosts and / or multiple storage devices. Storage device 1002 may be a solid-state drive (SSD), universal flash memory (UFS), hard disk drive (HDD), embedded multimedia card (eMMC), compressed flash memory (CF) card, secure digital card (SD) card, etc. Storage device 1002 may include a controller (processor) 1003 and a storage medium 1004 connected to the controller 1003. Host 1001 and / or storage device 1002 may include a CXL switch. Storage medium 1004 may include volatile memory, non-volatile memory, or both, and may include one or more flash memory chips (or other storage media). Controller 1003 may include one or more processors, one or more error correction circuits, one or more field-programmable gate arrays (FPGAs), one or more host interfaces, one or more flash bus interfaces, etc., or combinations thereof. Controller 1003 may be configured to facilitate the transfer of data / commands between host 1001 and storage medium 1004. Host 1001 can send data / commands to storage device 1002 for reception and processing by controller 1003 in conjunction with storage medium 1004. As described herein, methods, processes, and algorithms can be implemented on a storage device controller such as controller 1003.

[0097] Embodiments of the subject matter and operations described in this specification can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs (i.e., one or more modules of computer program instructions) encoded on a computer storage medium for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Additionally or alternatively, the program instructions may be encoded on artificially generated propagating signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiver device for execution by the data processing apparatus. The computer storage medium may be or is included in a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or a combination thereof. Furthermore, although the computer storage medium is not a propagating signal, it may be a source or destination of computer program instructions encoded in artificially generated propagating signals. The computer storage medium may also be or be included in one or more separate physical components or media (e.g., multiple optical discs (CDs), magnetic disks, or other storage devices). Furthermore, the operations described in this specification can be implemented as operations performed by a data processing apparatus on data stored on one or more computer-readable storage devices or received from other sources.

[0098] While this specification may contain numerous specific implementation details, these details should not be construed as limiting the scope of any claimed subject matter, but rather as descriptions of features specific to particular embodiments. Certain features described in the context of individual embodiments in this specification may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments. Furthermore, although features may be described above as functioning in certain combinations and even initially claimed in this way, in some cases one or more features from a claimed combination may be removed from the combination, and the claimed combination may be for sub-combinations or variations thereof.

[0099] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or to perform all the shown operations to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of the various system components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.

[0100] Therefore, specific embodiments of the subject matter have been described herein. Other embodiments are within the scope of the following claims. In some cases, the actions set forth in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired result. In some embodiments, multitasking and parallel processing may be advantageous.

[0101] As those skilled in the art will recognize, the innovative concepts described herein can be modified and varied across a wide range of applications. Therefore, the scope of the claimed subject matter should not be limited to any of the specific exemplary teachings discussed above, but is defined by the appended claims.

Claims

1. A method for managing memory in a computing system, comprising: Virtual nodes are generated by combining two or more physical nodes coupled to a compute fast link (CXL) switch; as well as The physical address of the data stored in the memory is identified based on the offset between the address ranges of the two or more physical nodes.

2. The method according to claim 1, wherein, The two or more physical nodes expose different address ranges corresponding to the shared memory regions in the CXL memory expander.

3. The method according to claim 1, further comprising: Maintain a memory allocation table that associates one or more pages in the shared memory region with the virtual node.

4. The method according to claim 1, wherein, The offset is determined based on the difference in the base addresses assigned to the two or more physical nodes.

5. The method according to claim 1, wherein, The two or more physical nodes are each coupled to the CXL switch via two or more CXL host adapters.

6. The method according to claim 1, further comprising: Use the physical address to retrieve the memory page.

7. The method according to claim 1, further comprising: When a process is migrated between central processing unit (CPU) nodes, the page intermediate directory (PMD) is updated to point to the page table entry (PTE).

8. The method according to claim 1, further comprising: When a process is migrated between central processing unit (CPU) nodes, the page intermediate directory (PMD) entries are updated to point to the page table entries (PTEs) associated with the base address corresponding to the local CXL host adapter.

9. The method according to claim 1, further comprising: Select the page table entry (PTE) associated with the base address corresponding to the local CXL host adapter to manage access to the shared memory region.

10. The method according to claim 1, wherein, Accessing the data stored in the memory with reduced latency compared to accessing the memory without the virtual node.

11. An apparatus for managing memory in a computing system, comprising: Compute Fast Link (CXL) switches are configured to couple two or more physical nodes; and The processor is configured as follows: Virtual nodes are generated by combining two or more physical nodes; as well as The physical address of the data stored in the memory is identified based on the offset between the address ranges of the two or more physical nodes.

12. The apparatus according to claim 11, wherein, The two or more physical nodes expose different address ranges corresponding to the shared memory regions in the CXL memory expander.

13. The apparatus according to claim 11, wherein, The processor is also configured to maintain a memory allocation table that associates one or more pages in the shared memory region with the virtual node.

14. The apparatus according to claim 11, wherein, The offset is determined based on the difference in the base addresses assigned to the two or more physical nodes.

15. The apparatus of claim 11, further comprising two or more CXL host adapters, in, The two or more physical nodes are each coupled to the CXL switch via the two or more CXL host adapters.

16. The apparatus according to claim 11, wherein, The processor is also configured to use the physical address to retrieve memory pages.

17. The apparatus according to claim 11, wherein, The processor is also configured to update the Page Intermediate Directory (PMD) to point to the Page Table Entries (PTE) when migrating processes between central processing unit (CPU) nodes.

18. The apparatus according to claim 11, wherein, The processor is also configured to update the Page Intermediate Directory (PMD) entries to point to the page table entries (PTEs) associated with the base address corresponding to the local CXL host adapter when migrating processes between central processing unit (CPU) nodes.

19. The apparatus according to claim 11, wherein, The processor is also configured to select a page table entry (PTE) associated with a base address corresponding to the local CXL host adapter to manage access to the shared memory region.

20. The apparatus according to claim 11, wherein, Accessing the data stored in the memory with reduced latency compared to accessing the memory without the virtual node.