Device and method for memory management using a multilink CXL switch for NUMA architecture

The multilink CXL switch in NUMA architectures addresses latency and resource inefficiencies by creating virtual nodes to manage overlapping memory areas, ensuring efficient and low-latency memory access.

JP2026089681APending Publication Date: 2026-06-01SAMSUNG ELECTRONICS CO LTD

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-11-18
Publication Date
2026-06-01

AI Technical Summary

Technical Problem

Existing NUMA architectures face increased latency and inefficient memory resource allocation due to redundant memory exposure and contention when multiple CXL host adapters connect to the same CXL memory expander, leading to overlapping memory areas and inefficient resource use.

Method used

A memory management method and device using a multilink CXL switch that combines physical nodes to create virtual nodes, identifies overlapping memory areas, and allocates resources efficiently through dynamic memory allocation, ensuring seamless access and reducing latency by routing memory access via local CXL host adapters.

Benefits of technology

This approach enhances memory access efficiency by minimizing latency and resource contention, allowing CPUs to access memory locally, maintaining seamless access and efficient resource utilization across nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026089681000001_ABST
    Figure 2026089681000001_ABST
Patent Text Reader

Abstract

This invention provides a device and method for memory management that utilize a multilink compute express link switch to optimize memory access and resource allocation. [Solution] The memory management method in a computing system according to the present invention comprises the steps of: generating a virtual node by combining two or more physical nodes coupled to a compute express link (CXL) switch; and identifying the physical address of data stored in memory based on the offset between the address ranges of the two or more physical nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to memory management in a Non-Uniform Memory Access (NUMA) architecture, and more particularly to an apparatus and a memory management method for memory management using a Multi-Link Compute Express Link (CXL) switch to optimize memory access and resource allocation.

Background Art

[0002] The NUMA architecture may be employed in high-performance computing systems to manage memory resources across multiple Central Processing Unit (CPU) sockets. In such an architecture, the latency of memory access may vary significantly depending on whether the memory being accessed is local to the CPU socket executing the process or exists in a remote socket. To address this variability, technologies such as CXL have been developed, enabling fast and coherent access to memory resources across the distributed system.

[0003] A CXL host adapter is connected to a CPU socket and communicates with a CXL memory expander via a CXL switch. While this configuration enables scalability and efficient resource sharing, there is a potential for increased latency when a CPU socket accesses memory via a remote adapter. Such latency variations can be particularly pronounced in workloads that require frequent memory access, as the time required to access remote memory can impact the performance of the entire system.

[0004] To optimize memory access in a NUMA architecture, it is necessary to address issues such as redundant memory exposure and efficient allocation of memory resources. These challenges become even more complex when multiple CXL host adapters are connected to the same CXL memory expander, resulting in overlapping memory areas being exposed across multiple physical nodes. Existing operating systems and memory management frameworks often fail to adequately account for such redundancy, which can result in inefficient resource use and potential contention. [Overview of the project] [Problems that the invention aims to solve]

[0005] The present invention has been made in view of the problems in the above-mentioned conventional high-performance computing systems, and the object of the present invention is to provide a device and a memory management method for memory management using a multilink compute express link (CXL) switch to optimize memory access and resource allocation. [Means for solving the problem]

[0006] To achieve the above objective, the memory management method in a computing system according to the present invention comprises the steps of generating a virtual node by combining two or more physical nodes coupled to a Compute Express Link (CXL) switch, The method is characterized by comprising the step of identifying the physical address of data stored in memory based on the offset between the address ranges of the two or more physical nodes.

[0007] To achieve the above objective, the present invention provides a device for memory management in a computing system comprising: a compute express link (CXL) switch configured to connect two or more physical nodes; and a processor, wherein the processor is configured to combine the two or more physical nodes to generate a virtual node and to identify the physical address of data stored in the memory based on the offset between the address ranges of the two or more physical nodes. [Effects of the Invention]

[0008] According to the memory management device and memory management method of the present invention, in a NUMA architecture using a multilink CXL switch, each CPU socket is enhanced by a multilink function equipped with a dedicated CXL host adapter. This configuration allows both CPUs to locally access one or more CXL memory expanders via a shared CXL switch, and by routing memory access via the local CXL host adapter, processes running on either CPU can reduce latency. Furthermore, resources can be used efficiently through virtual nodes and dynamic memory allocation, and seamless memory access is maintained through migration between nodes. [Brief explanation of the drawing]

[0009] [Figure 1] This is a block diagram showing a schematic configuration of a CXL memory system according to a conventional embodiment. [Figure 2] This block diagram shows a schematic configuration of an extended CXL memory system according to an embodiment of the present invention. [Figure 3] (a) is a diagram showing the configuration of a memory allocator and node management according to a conventional embodiment, and (b) is a diagram showing the configuration of an extended memory allocator and node management according to an embodiment of the present invention. [Figure 4] This is a flowchart illustrating the node initialization process in a multilink CXL architecture according to an embodiment of the present invention. [Figure 5] (a) is a diagram showing the configuration of a memory allocator and node management that illustrates page allocation from physical memory according to a conventional embodiment, and (b) is a diagram showing the configuration of an extended memory allocator and node management that illustrates page allocation from physical memory according to an embodiment of the present invention. [Figure 6](a) is a diagram showing the configuration of an extended memory allocator and node management that illustrates page allocation from physical memory according to an embodiment of the present invention, and (b) is a flowchart for explaining the operation of the memory allocator when using a virtual node upon receiving a memory allocation request according to an embodiment of the present invention. [Figure 7] This figure illustrates the use of a unique address space for implementing the CXL memory expander according to an embodiment of the present invention. [Figure 8] This figure illustrates the use of multiple page table entries (PTEs) to manage memory allocation for supporting process transitions according to embodiments of the present invention. [Figure 9] This is a flowchart illustrating a memory management method in a computing system according to an embodiment of the present invention. [Figure 10] This block diagram shows a schematic configuration of a storage system according to an embodiment of the present invention. [Modes for carrying out the invention]

[0010] Next, specific examples of embodiments for implementing the memory management device and memory management method according to the present invention will be described with reference to the drawings.

[0011] Note that even if they are shown in different drawings, the same elements are given the same reference numeral. In the following description, specific details such as detailed configurations and components are provided only to aid in the overall understanding of the embodiments of the present invention. Therefore, it will be obvious to those skilled in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present invention. Furthermore, in order to maintain clarity and conciseness, explanations of well-known functions and structures will be omitted. The terms described below are defined in consideration of the function of this invention, and may differ depending on the user, user intent, or custom. Therefore, the definitions of these terms should be interpreted based on the descriptions throughout this specification.

[0012] The present invention can have various modifications and other embodiments, and some of these embodiments will be described in detail below with reference to the accompanying drawings. However, it should be understood that the present invention is not limited to these embodiments, and includes all modifications, equivalents, and alternatives within the scope of the present invention. Terms including ordinal numbers such as first and second may be used to describe various elements, but the components are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the present invention, the first component can be referred to as the second component. Similarly, the second component can also be referred to as the first component. As used herein, the term "and / or" includes any combination of one or more of the related items. The terms used herein are merely used to describe various embodiments of the present invention and are not intended to limit the present invention.

[0013] The singular form shall include the plural form unless otherwise clearly defined in the context. The terms "include" or "have" used in the present invention indicate the existence of features, numerical values, steps, operations, components, parts, or combinations thereof, and it should be understood that they do not exclude the additional existence or possibility of one or more other features, numerical values, steps, operations, components, parts, or combinations thereof. Unless otherwise defined, all terms used herein shall have the same meaning as understood by those skilled in the art to which the present invention pertains. Terms as defined in commonly used dictionaries should be interpreted as having the same meaning as their contextual meaning in the art, and should not be interpreted as having an ideal or overly formal meaning unless explicitly defined in the present invention.

[0014] An electronic device according to an embodiment of the present invention may be any of various types of electronic devices that utilize a memory device. Electronic devices may use any suitable storage standard, or any combination thereof, such as Peripheral Component Interconnect Express (PCIe), Nonvolatile Memory Express (NVMe), NVMe-over-Fabric (NVMeoF), Advanced eXtensible Interface (AXI), Ultra Path Interconnect (UPI), Ethernet, Transmission Control Protocol / Internet Protocol (TCP / IP), Remote Direct Memory Access (RDMA), RDMA over Converged Ethernet (ROCE), Fibre Channel (FC), InfiniBand (IB), Serial Advanced Technology Attachment (SATA), Small Computer System Interface (SCSI), Serial Attached SCSI (SAS), Internet Wide-Area RDMA Protocol (iWARP), etc.

[0015] In one embodiment, the interconnect interface may be implemented by one or more memory semantic and / or memory coherent interfaces and / or protocols, which may include one or more CXL protocols such as CXL.mem, CXL.io, CXL.cache, Gen-Z, Coherent Accelerator Processor Interface (CAPI), Cache Coherent Interconnect for Accelerators (CCIX), or any combination thereof. Any memory device may be implemented using one or more of any type of memory device interface, including Double Data Rate (DDR), DDR2, DDR3, DDR4, DDR5, Low-Power DDR (LPDDRX), Open Memory Interface (OMI), NVLink High Bandwidth Memory (HBM), HBM2, HBM3, and others. Electronic devices may include, for example, mobile communication devices (such as smartphones), computers, portable multimedia devices, portable medical devices, cameras, wearable devices, or household electrical appliances. However, electronic devices are not limited to the examples given above.

[0016] The terms used in this invention are not intended to limit the disclosure, but rather to encompass various modifications, equivalents, or substitutions to the corresponding embodiments. With regard to the description of the attached drawings, similar reference numerals may be used to refer to similar or related elements. The singular form of a noun corresponding to an item shall encompass one or more objects unless the context clearly indicates otherwise. As used herein, the terms “A or B,” “at least one of A and B,” “at least one of A or B,” “A, B, or C,” “at least one of A, B, and C,” and “at least one of A, B, or C” may encompass any combination of the items listed together in the corresponding terms. The terms "first," "second," etc., used herein may be used to distinguish corresponding components from other components, but are not intended to limit the components in any other respect (e.g., importance or order).

[0017] In this specification, when an element (e.g., a first element) is described as being "coupled with," "coupled to," "connected with," or "connected to" another element (e.g., a second element), with or without the terms "operatively" or "communicatively," it is intended that the element may be coupled with other elements directly (e.g., wired) or wirelessly, or via a third element. As used herein, the term “module” may encompass a unit implemented by hardware, software, firmware, or a combination thereof, and may be used synonymously with other terms such as “logic,” “logic block,” “component,” or “circuit.” A module may be a single, integrated component, or a smallest unit or part thereof, configured to perform one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC), a coprocessor, or a field-programmable gate array (FPGA).

[0018] While this specification contains many specific implementation details, these implementation details should not be interpreted as limitations on the scope of the claimed subject matter, but rather as descriptions of features specific to particular embodiments. Certain features described herein in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment can also be implemented separately or in any appropriate combination of sub-embodiments in multiple embodiments. Furthermore, while the features are described above as acting in a specific combination, and may even be initially claimed as such, one or more features from the claimed combination may be removed from the combination in some cases, and the claimed combination may relate to a subcombination or a variation of a subcombination. Similarly, although the actions are depicted in a specific order in the drawings, this should not be understood as requiring that such actions be performed in a specific illustrated order or sequential order, or that all illustrated actions be performed, in order to achieve a desirable result. In some situations, multitasking or parallel processing can be advantageous. Furthermore, the separation of various system components in the embodiments described above should not be understood as necessary in all embodiments, and the described program components and systems can generally be integrated in a single software product or packaged in multiple software products.

[0019] Traditional NUMA architectures suffered from increased latency when accessing memory from remote CPU sockets. This limitation arises from the lack of local memory access paths and inefficient memory management across multiple nodes. CXL is an interconnect and protocol designed to provide high-speed, coherent access to memory and accelerators, enabling improved performance in distributed computing systems.

[0020] Figure 1 is a block diagram showing the schematic configuration of the CXL memory system. Referring to Figure 1, the system 100 includes two CPU sockets (101, 102), two PCIe switches (103, 104), and a CXL host adapter 105.

[0021] The CXL host adapter 105 is connected to the shared CXL switch 106 and uses a virtual CXL switch (VCS) unit to enable access to one or more CXL memory expanders (107a, 107b, and / or 107n). A VCS unit divides one or more CXL switches into virtual switches, allowing individual hosts to manage and access shared memory resources. This architecture allows both CPU sockets (101, 102) to access one or more CXL memory expanders (107a, 107b, and / or 107n) via local or remote socket connections.

[0022] Figure 1 shows a distinction between local memory access paths and remote memory access paths. Local memory access occurs when the CPU socket communicates directly with the corresponding CXL host adapter and accesses the CXL memory expander via a shared CXL switch, thereby reducing latency. On the other hand, remote access occurs when the CPU accesses memory via the CXL host adapter on the other socket, and the traversal of the additional interconnect increases latency. Therefore, the CPU 101 accesses one or more CXL memory expanders (107a, 107b, and / or 107n) locally via the CXL switch 106 and the CXL host adapter 105. This is because these components are located downstream of the PCIe switch 103 of the CPU 101. On the other hand, CPU102 operates via CPU101 and only performs remote access to the same resources. Figure 1 illustrates several architectural limitations that increase latency when a process running on one CPU socket (e.g., CPU102) accesses memory resources via a remote socket (e.g., CPU101).

[0023] Figure 2 is a block diagram showing a schematic configuration of an extended CXL memory system according to an embodiment of the present invention. Referring to Figure 2, System 200 includes two CPU sockets (201, 202), two PCIe switches (203, 204), and two CXL host adapters (205, 206).

[0024] The CXL host adapters (205, 206) are connected to the shared CXL switch 207 and use the VCS unit to enable access to one or more CXL memory expanders (208a, 208b, and / or 208n). This architecture allows both CPU sockets to access one or more CXL memory expanders (208a, 208b, and / or 208n) via a low-latency local connection. In other words, both CPU201 and CPU202 are equipped with their respective CXL adapters (205, 206) that are local to their respective PCIe switches (203, 204).

[0025] The CXL memory system 200 shown in Figure 2 is enhanced by a multilink function that provides each CPU socket with its own dedicated CXL host adapter. This configuration allows both CPU201 and CPU202 to locally access one or more CXL memory expanders (208a, 208b, and / or 208n) via the shared CXL switch 207. System 200 routes memory access via a local CXL host adapter, enabling processes running on any of the CPUs to access memory with reduced latency.

[0026] The boxes labeled "VCS 0," "VCS 1," "VCS n-1," and "VCS n" within the CXL switch 207, and the lines labeled "Shared" connecting them to one or more CXL memory expanders (208a, 208b, and / or 208n), represent the concept of an extended VCS system that enables memory sharing in the multilink CXL switch 207. A VCS unit refers to a logical entity within a physical CXL switch that generates a separate memory hierarchy for each connected host. This allows each VCS unit to independently access its assigned CXL host adapter (e.g., 205 and 206) as if it were directly connected to the host, providing memory space isolation and efficient memory management across multiple systems (hosts or applications).

[0027] Similar to the VCS units shown in Figure 1, these VCS units in Figure 2 are logical entities within a CXL switch that facilitate the sharing of memory resources within one or more CXL memory expanders (208a, 208b, and / or 208n). However, in contrast to the VCS unit shown in Figure 1, the VCS unit in Figure 2 is designed to facilitate efficient memory sharing between CPU 201 and CPU 202 via one or more CXL memory expanders (208a, 208b, and / or 208n). Unlike Figure 1, where memory access involves a remote communication path, the architecture shown in Figure 2 incorporates dedicated CXL host adapters (205, 206) into CPU 201 and CPU 202, respectively, enabling each CPU 201 and CPU 202 to establish a direct connection to the VCS unit in the CXL switch 207. This localized access mechanism allows each CPU201 and CPU202 to access memory directly (locally) from one or more CXL memory expanders (208a, 208b, and / or 208n) without unnecessarily using remote connection paths via a remote CXL host adapter. As a result, the system 200 shown in Figure 2 significantly reduces or eliminates latency associated with remote memory access.

[0028] The lines labeled "Shared" in Figure 2 indicate that memory resources within one or more CXL memory expanders (208a, 208b, and / or 208n) are dynamically shared among VCS units rather than being statically allocated. This characteristic allows multiple CPUs to directly (locally) access the same memory region without causing contention or redundancy, while maintaining data consistency using CXL's cache-coherent protocol. By providing CXL host adapters on both CPU sockets, the architecture shown in Figure 2 minimizes the CPU's dependence on accessing memory using a remote CXL host adapter, thereby reducing the likelihood of remote memory access and improving overall latency.

[0029] Furthermore, system 200 introduces a software layer mechanism to manage shared memory resources and prevent the exposure of redundant memory. This configuration is scalable across multiple CPUs and supports the allocation and migration of memory resources between nodes. When two CXL host adapters (e.g., 205 and 206) connect to the same CXL memory expander (e.g., 208a, 208b, or 208n), the memory may be redundantly exposed as multiple nodes with different physical addresses. This redundant exposure complicates memory management and increases the likelihood of resource contention. To address this, the memory allocator operates on a per-virtual node basis, consolidating redundant physical memory areas into a single virtual node. This abstraction of virtual nodes makes it possible to manage multiple physical nodes that reference the same memory medium as a unified entity.

[0030] Therefore, the term "physical node" (for example, nodes (311a-314a) in Figure 3(a), nodes (311b-314b) in Figure 3(b), nodes (511a-514a) in Figure 5(a), nodes (511b-514b) in Figure 5(b), and nodes (611a-614a) in Figure 6(a)) refers to the mapping of physical address ranges associated with a particular CXL host adapter. As will be discussed later with reference to Figures 3(a) and (b), Figures 5(a) and (b), and Figure 6(a), physical nodes may redundantly represent the same physical memory area, which may result in inefficient memory allocation.

[0031] Another challenge arises during process migration between nodes. When a process migrates from one node to another, the system updates its memory addresses to reflect the memory map of the local node. If the address is not updated, the process will access memory via the remote node, which can result in unnecessary latency and undermine the advantages of a multilink architecture. This invention provides a method for highly efficient management of nodes and memory using virtual nodes. A virtual node is a logical entity that manages memory resources by combining or dividing physical nodes based on shared or overlapping memory areas. Virtual nodes enable efficient memory allocation and prevent redundancy by treating multiple physical nodes with overlapping memory areas as a single node on a logical memory map. This method addresses the challenges associated with redundant memory regions in multilink CXL architectures, where the presence of multiple CXL host adapters can expose overlapping memory regions across multiple physical nodes. By creating virtual nodes, the system can consolidate or divide physical nodes to logically manage memory resources and reduce redundancy.

[0032] Figure 3(a) shows the configuration of the memory allocator and node management. Referring to Figure 3(a), in one node management method, logical nodes (301a, 302a, 303a, 304a) are directly mapped to physical nodes (311a, 312a, 313a, 314a) without considering redundant or overlapping memory areas. The mapping from logical nodes to physical nodes is performed within the CPU (for example, CPU201 or CPU202 in Figure 2). As shown in the diagram, the four physical nodes (311a, 312a, 313a, 314a) each directly correspond to a logical node (301a, 302a, 303a, 304a).

[0033] The memory allocator 300a may operate independently for each logical node (301a, 302a, 303a, 304a), resulting in redundant memory areas being managed individually. However, both physical nodes (313a, 314a) are compatible with the same underlying CXL memory 306a (the term "CXL memory" is used interchangeably with "CXL memory expander" and "CXL memory area"). Because the logical nodes (303a, 304a) are mapped one-to-one to the physical nodes (313a, 314a), the memory allocator 300a independently and redundantly tracks the logical nodes (303a, 304a) and manages what is physically a single shared CXL memory resource represented by the CXL memory 306a. For example, CXL memory 306a contains 64GB of physical memory, but due to redundant exposure by physical nodes (313a, 314a), it may be recognized by memory allocator 300a as two separate 64GB memory regions. As a result, even though there is actually only 64GB of physical memory, the system may recognize the memory as 128GB in total. As a result, the memory allocator 300a operating under this configuration may treat overlapping memory regions as separate memory regions with different physical addresses, which could lead to inefficient memory utilization.

[0034] Figure 3(b) shows the configuration of the extended memory allocator and node management according to an embodiment of the present invention. Referring to Figure 3(b), the logical nodes (301b, 302b) are directly mapped to the physical nodes (311b, 312b), respectively. This mapping relationship is similar to the correspondence between the logical nodes (301a, 302a) and the physical nodes (311a, 312a) in Figure 3(a). However, unlike the approach in Figure 3(a), the memory allocator 300b shown in Figure 3(b) identifies overlapping memory regions between physical nodes (313b, 314b) and consolidates them into a single virtual node 305b. Virtual node 305b allows the host to manage access to the same CXL memory region through multiple address spaces associated with different CXL host adapters.

[0035] Specifically, duplicate physical nodes (313b, 314b) that are redundantly mapped to the same CXL memory area 306b are consolidated into this single virtual node 305b. By introducing virtual node 305b, memory allocator 300b manages shared memory as a single, unified resource. This prevents redundant memory allocation that can occur when the same physical memory area is managed independently, as shown in Figure 3(a). Therefore, in this embodiment, the memory allocator 300b is reconfigured to manage virtual nodes (e.g., 305b) rather than managing logical nodes that directly correspond to physical nodes. This allows the system to treat overlapping memory regions as a unified entity.

[0036] Figure 4 is a flowchart illustrating the node initialization process in a multilink CXL architecture according to an embodiment of the present invention. Referring to Figure 4, the process begins in step S401. In step S401, system firmware such as the Basic Input / Output System (BIOS) or Unified Expansion Firmware Interface (UEFI) detects available memory nodes within the system. These firmware components identify physical memory blocks and collect metadata about their configuration.

[0037] In step S402, the system initializes the NUMA node table using platform-specific information provided by system tables such as the System Resource Affinity Table (SRAT), the System Locality Information Table (SLIT), and / or the CXL Early Discovery Table (CEDT). These tables provide detailed information about memory topology, locality, and interconnect relationships. In step S403, the system constructs a set of memory blocks corresponding to the physical nodes. Each block represents a contiguous memory region belonging to an individual physical node, and corresponds to an entire memory device or a portion of a memory region.

[0038] In step S404, the first memory block is acquired, and the system begins evaluating its state. In step S405, a check is performed to determine whether all detected memory blocks have already been processed and registered in the NUMA node table. Once all memory blocks are registered, the initialization process ends.

[0039] However, if there are unprocessed memory blocks remaining, in step S406 the system evaluates whether the current memory block is in a duplicate area. When redundant connections in a CXL host adapter map two or more physical nodes to the same physical memory area, overlapping memory areas may occur. If an overlapping area is detected, in step S407 the system further evaluates whether the memory block is completely contained within the overlapping area.

[0040] For blocks that do not completely overlap, in step S408, the system divides the block into smaller subblocks to allow for more precise handling of overlaps. For blocks that are completely duplicates, in step S409, the system processes the block without further division and determines whether the memory block is already registered in the NUMA node table.

[0041] If the block is already registered, in step S410 the system moves on to the next unprocessed memory block. Otherwise, in step S411, the system generates a NUMA node table and registers memory blocks (for example, associating a virtual node with a duplicate physical node). In step S411, the system retrieves the next unprocessed memory block and repeats this series of steps until all blocks have been registered (i.e., until "Yes" is determined in step S405). Once all memory blocks have been processed, the node initialization process ends.

[0042] Figure 5(a) shows the configuration of a memory allocator and node management that illustrates page allocation from physical memory according to a conventional embodiment. Referring to Figure 5(a), in this configuration, each logical node (501a, 502a, 503a, 504a) maintains an independent data structure for tracking free memory pages within its own node's management area. For example, each logical node (501a to 504a) corresponds to a physical node (511a to 514a), and these are mapped to the same physical memory (e.g., the same CXL memory area).

[0043] Similar to the case in Figure 3(a), in Figure 5(a), the physical nodes (513a, 514a) are mapped to the same CXL memory 506a. As a result, each logical node (503a, 504a) independently corresponds to a physical node (513a, 514a), thus redundantly exposing identical or duplicate memory (memory pages within CXL memory 506a). In this case, since the pages in CXL memory 506a are accessible through different logical nodes associated with different physical address mappings, the memory allocator 500a may treat the same memory page as if it were located at two different physical addresses. Figure 5(a) illustrates this state by showing that the same page is exposed to both logical nodes (503a, 504a).

[0044] This redundancy creates challenges in the memory management architecture. Because the memory allocator 500a cannot recognize overlapping mappings, it may incorrectly perceive two separate programs running on different logical nodes as accessing different memory regions and using the same physical memory page. If there is no mechanism to detect or adjust for this duplication, programs may write to the same underlying memory location, potentially resulting in inconsistent states or memory corruption. Conflicts occur because the same memory page is reachable through different physical address ranges, and the memory allocator 500a interprets them as independent, even though they actually refer to the same shared resource.

[0045] Figure 5(b) shows the configuration of an extended memory allocator and node management system illustrating page allocation from physical memory according to an embodiment of the present invention. Referring to Figure 5(b), the logical nodes (501b, 502b) are directly mapped to the physical nodes (511b, 512b), similar to the mapping between the logical nodes (501a, 502a) and the physical nodes (511a, 512a) in Figure 5(a). However, unlike the approach in Figure 5(a), the memory allocator 500b in Figure 5(b) identifies overlapping memory regions between the physical nodes (513b, 514b) and consolidates them into a single virtual node 505b.

[0046] Specifically, physical nodes (513b, 514b) that are duplicately mapped to the same page within the CXL memory area 506b are consolidated into this single virtual node 505b. By introducing virtual node 505b, memory allocator 500b manages shared memory as a single, unified resource. This prevents redundant memory allocation that occurs when identical pages are managed independently, as shown in Figure 5(a). This improved configuration has several advantages. By consolidating duplicate memory areas into virtual nodes, the system can prevent contention and reduce the complexity of memory management. This approach is effective in multilink CXL systems where multiple CXL host adapters may expose overlapping areas of the CXL memory expander. Enhanced memory allocators can provide a scalable solution for high-performance computing systems, ensuring consistent and contention-free memory allocation.

[0047] Figure 6(a) shows the configuration of an extended memory allocator and node management system illustrating page allocation from physical memory according to an embodiment of the present invention. Referring to Figure 6(a), the symbols (600a, 601a, 602a, 605a, 611a, 612a, 613a, 614a, 606a) correspond to the symbols (500b, 501b, 502b, 505b, 511b, 512b, 513b, 514b, 506b) in Figure 5(b), respectively, and the same explanation and function apply to these components.

[0048] According to embodiments of the present invention, unlike conventional memory allocators that may directly return a physical address, this memory allocator can return an offset instead of a physical address when allocating memory from a virtual node. The offset represents the virtual node's location in the address space, allowing the system to determine the physical memory address based on the physical node on which the process is running. For example, a memory allocator can calculate the final physical address by adding the base address of the physical node to the offset. Therefore, this mechanism ensures that the memory allocated from a virtual node is accessible from multiple physical nodes mapped to that virtual node.

[0049] Figure 6(b) is a flowchart illustrating the operation of the memory allocator when using a virtual node upon receiving a memory allocation request according to an embodiment of the present invention. Referring to Figure 6(b), the process begins in step S601b by receiving a memory allocation request that may specify a node identifier (ID) and size.

[0050] In step S602b, the allocator obtains the requested memory offset from the node identified by the node ID and size. The allocator manages free memory pages using offsets. In other words, the free list stores and returns an offset value relative to the base address, rather than the complete physical address. The free list is used to track available memory blocks and is configured to return an offset value when a memory page is allocated.

[0051] In step S603b, the allocator determines whether the node ID matches the virtual node. If the node ID matches a virtual node, in step S604b, the physical memory address is calculated by adding an offset to the base address of the current virtual node (the node identified by the node ID and size). The node ID may represent the node on which the process is running, so it does not necessarily need to be stored. Alternatively, a metadata structure (e.g., struct node) could hold information about the virtual node, allowing the allocator to determine whether a given node ID corresponds to a virtual node.

[0052] If the node ID is different from that of the virtual node, the physical memory address is determined in step S605b by adding an offset to the base address of the physical node. In step S606b, the allocator sends the determined physical address to the requester, which in turn allows the requester to update the process's page table entry (PTE). The stored data can be retrieved using the determined physical address. Therefore, by returning an offset rather than a physical address, compatibility with processes running on different physical nodes can be maintained.

[0053] Figure 7 illustrates the use of a unique address space for implementing the CXL memory expander according to an embodiment of the present invention. Referring to Figure 7, each node, namely CPU node 701 and CPU node 702, is shown as process A and process B, respectively, corresponding to the processes running on the corresponding CPU.

[0054] The memory management system uses virtual address 703 for process A and virtual address 704 for process B, with each being mapped to physical address 705. These virtual addresses are translated to physical addresses via a multi-level page table hierarchy managed by the Memory Management Unit (MMU). This hierarchy conforms to the format used in the x86-64 architecture and may include a Page Global Directory (PGD), Page Upper Directory (PUD), Page Middle Directory (PMD), and PTE, which work together to translate virtual addresses to physical addresses. PGD, PUD, PMD, and PTE constitute a hierarchical translation mechanism that gradually narrows the virtual address range.

[0055] When a virtual address is accessed, the most significant bit of the virtual address is used as the index for the PGD to identify the corresponding PUD. PGD ​​refers to PUD, which helps isolate large memory segments between different processes by dividing the high-level virtual address space into more manageable areas. PUD provides even finer granularity by storing a pointer to PMD and allowing the selection of smaller regions. PMD determines which PTE table contains the final mapping for the virtual address. PTE tables consist of smaller memory areas than PMDs, resulting in even greater granularity. Furthermore, the PMD also functions as a control point for changing the path to physical memory resources (e.g., the CXL memory expander 708) without altering the entire page table hierarchy (e.g., without changing the PGD and PUD).

[0056] For example, when a process is migrated from one CPU node to another, even if the underlying physical memory that the process accesses (e.g., the CXL memory expander 708) remains the same, the physical address used to reach that memory may differ depending on which CXL host adapter (CHA) is local to that node. Instead of rebuilding or rewriting the entire page table (PGD, PUD, PMD, PTE), the system can redirect the transformation by modifying the PMD entry to point to a different page table (another PTE). This separate page table contains valid mappings for the new node's local CHA address space. As a result, the system can access the same physical memory area across different CPU nodes with fewer page table rewrites. This redirection mechanism avoids address contention by ensuring that each CPU node accesses shared memory via a PTE page that reflects its own node's local physical address space.

[0057] CHA706 and CHA707 maintain a unique physical address space for the CXL memory expander 708. As a result, each CXL host adapter is local to its respective node (for example, CHA706 is local to node 701, and CHA707 is local to node 702), thus exposing overlapping memory areas within the CXL memory to each node in a different way. For example, a memory region exposed on CHA706 can be accessed through a certain physical address, while the same memory region exposed on CHA707 can be accessed through a different physical address. This allows each node to access memory via its local CXL host adapter, minimizing latency and optimizing performance.

[0058] This memory configuration ensures that each process uses the appropriate physical address corresponding to its own node's local CXL host adapter. For example, process A running on node 701 translates its virtual address to a physical address exposed through CHA706, and process B running on node 702 translates its virtual address to a physical address exposed through CHA707. This approach avoids contention and ensures efficient memory access across nodes. Furthermore, Figure 7 shows that in each node, the CR3 register, a system control register containing the physical address of the page directory, points to the base of the paging hierarchy, allowing the CPU to efficiently translate the virtual addresses of each process.

[0059] Therefore, Figure 7 illustrates a scenario in which two different processes (A and B) run independently on separate CPU nodes (701 and 702), with each process accessing the shared CXL memory expander 708. Process A and Process B use different PTEs, each pointing to different physical address ranges corresponding to the same CXL memory expander 708. Process A and Process B access the same physical memory, but because they are exposed via separate host adapters (706, 707), they use different addresses to access it. This configuration allows processes running on separate CPU nodes to independently manage and access memory via localized paths.

[0060] Figure 8 illustrates the use of a PTE to manage memory allocation to support process transitions according to an embodiment of the present invention. In contrast to Figure 7, which shows two independent processes accessing shared CXL memory from separate CPU nodes (701 and 702), Figure 8 shows a scenario in which a single process C migrates from CPU node 801 to another CPU node 802. To ensure that process C can continue to access the same memory region via the local host adapter after the migration, the system uses dual PTEs corresponding to the same memory pages within the CXL memory expander 808. These dual PTEs are located at different physical addresses exposed by CHA805 and CHA806, respectively.

[0061] Referring to Figure 8, a hierarchical paging structure consisting of PGD, PUD, PMD, and multiple PTEs is shown. These tables are managed by the MMU, which is configured to translate virtual address 803 to physical address 804. The conversion proceeds in stages. PGD ​​divides the virtual address space at a higher level, and each entry refers to PUD, which further subdivides the address range. The PUD is then configured to point to the PMD table. The PMD is then configured to point to a PTE table that translates virtual addresses to specific physical addresses.

[0062] In the example shown in Figure 8, the system is configured to maintain dual PTEs. The base address of each PTE is identified along the path (803a, 803b). Dual PTE represents two separate address ranges corresponding to two different CXL host adapters (CHA805 and CHA806). These host adapters are configured to provide access to the shared memory area within the CXL memory expander 808. The PTE pointed to by path 803a and the PTE pointed to by path 803b consist of different physical address ranges, but both are ultimately mapped to the same physical memory page 807. CHA805 is locally connected to CPU node 801, and CHA806 is locally connected to CPU node 802.

[0063] To support CXL memory allocation, a set of memory pages (e.g., 8KB total) may be reserved for the final level of paging. This final level structure may include a PTE for standard-sized 4KB pages, a PMD for large 2 megabyte (MB) pages, and a PUD for extra-large 1 gigabyte (GB) pages. When a program allocates memory within the CXL memory expander 808, the associated PTEs are initialized and configured so that one entry (e.g., a PTE from path 803a) corresponds to the base address used by CHA805, and a second entry (e.g., a PTE from path 803b) corresponds to the base address used by CHA806. The second entry can be calculated by applying a known offset between the address ranges of the two adapters.

[0064] During execution, if process C migrates from CPU node 801 to CPU node 802, the system is configured to update the corresponding PMD entry to refer to the PTE mapped to the local CXL host adapter (e.g., CHA805). This update is triggered by detecting a change in the running CPU node and is performed by adjusting the PMD entry to point to the new base address. This adjustment is made by adding or subtracting a fixed offset (e.g., ±4KB) to cause the PMD entry to point to path 803a instead of path 803b. This redirection ensures that subsequent memory accesses issued by process C always occur via the local adapter by controlling the exposure of physical addresses at the PTE level.

[0065] Therefore, Figure 8 shows a migration-aware memory translation mechanism that maintains an efficient access path to shared memory using dual PTEs and dynamic PMD updates. By aligning memory access to the node-local host adapter (CHA806 in the post-migration case), the system reduces interconnect traffic, avoids remote memory access, and maintains coherence between CPU nodes using CXL's cache-coherent protocol.

[0066] Figure 9 is a flowchart illustrating a memory management method in a computing system according to an embodiment of the present invention. The method shown in Figure 9 is implemented by a processor, memory controller, system-on-a-chip (SoC), or other processing unit capable of managing virtual memory. In one embodiment, the method is executed by system software (e.g., an operating system (OS)) running on a general-purpose CPU.

[0067] Referring to Figure 9, in step S901, a virtual node is created by combining two or more physical nodes coupled to the CXL switch. For example, a processor implementing this method identifies overlapping memory areas exposed on both nodes and allocates them to the virtual nodes. In step S902, the physical address of the data stored in memory is identified based on the offset between the node's address ranges. This is implemented by maintaining dual PTE pages. In this case, the base address of the mapping for the second node is derived by adding a fixed offset (e.g., ±4KB) to the base address of the first node. The PPMD ​​entry is updated during execution to select a PTE page based on the CPU node in which it is running.

[0068] Figure 10 is a block diagram showing a schematic configuration of a storage system according to an embodiment of the present invention. Referring to Figure 10, the storage system 1000 includes a host 1001 and a storage device 1002. Although the diagram shows one host and one storage device, the storage system 1000 may include multiple hosts and / or multiple storage devices.

[0069] The storage device 1002 may be an SSD (Solid State Drive), UFS (Universal Flash Storage), HDD (Hard Disk Drive), eMMC (Embedded Multimedia Card), CF (CompactFlash®) card, SD (Secure Digital) card, etc. The storage device 1002 includes a controller (processor) 1003 and a storage medium 1004 connected to the controller 1003. The host 1001 and / or storage device 1002 include a CXL switch. The storage medium 1004 may include volatile memory, non-volatile memory, or both, and may further include one or more flash memory chips (or other storage mediums).

[0070] The controller 1003 may include one or more processors, one or more error correction circuits, one or more field-programmable gate arrays (FPGAs), one or more host interfaces, one or more flash bus interfaces, or a combination thereof. The controller 1003 is configured to facilitate the transfer of data / commands between the host 1001 and the storage medium 1004. The host 1001 receives data / commands from the controller 1003 and processes them in conjunction with the storage medium 1004, then transmits them to the storage device 1002. As described herein, the methods, processes, and algorithms may be implemented on a storage device controller such as controller 1003.

[0071] The embodiments and operations of the subject matter described herein can be implemented in digital electronic circuits, or in computer software, firmware, or hardware (including the structures disclosed herein and their structural equivalents), or in one or more combinations thereof. Embodiments of the subject matter described herein can be implemented as one or more computer programs, that is, as one or more modules of computer program instructions encoded on a computer storage medium for execution by a data processing device or for controlling the operation of a data processing device. Furthermore, or alternatively, program instructions can be encoded into artificially generated propagating signals, such as mechanically generated electrical, optical, or electromagnetic signals, which are generated to encode information for transmission to a suitable receiving device for execution by a data processing device.

[0072] Computer storage media may be computer-readable storage devices, computer-readable storage boards, random-access or serial-access memory arrays or devices, or combinations thereof, or may be included in such devices. Furthermore, while computer storage media are not propagating signals themselves, they can be sources or destinations for computer program instructions encoded with artificially generated propagating signals. Computer storage media can also be one or more separate physical components or media (e.g., multiple compact discs (CDs), disks, or other storage devices), or may be contained within them. Furthermore, the operations described herein can be performed as operations executed by a data processing device on data stored in one or more computer-readable storage devices, or on data received from other sources.

[0073] Thus, specific embodiments of the subject matter have been described herein. Other embodiments are within the scope of the following claims. In some cases, preferred results can still be obtained by performing the operations described in the claims in a different order. Furthermore, the steps depicted in the attached diagram do not necessarily require a specific or sequential order to obtain favorable results. In certain implementations, multitasking and parallel processing may be advantageous. As will be apparent to those skilled in the art, the innovative concepts described herein are modifiable and adaptable to a wide range of applications. Therefore, the scope of the claimed subject matter should not be limited to the specific exemplary teachings described above, but rather defined by the following claims. [Explanation of Symbols]

[0074] 200 Systems 201, 202 CPU sockets 203, 204 PCIe switches 205, 206 CXL Host Adapters 207 Shared CXL Switch 208a, 208b, 208n CXL Memory Expander 300b Memory Allocator 301b, 302b logical nodes 305b Virtual Node 306b CXL memory area Nodes 311b~314b

Claims

1. A memory management method in a computing system, The steps include: creating a virtual node by combining two or more physical nodes coupled to a Compute Express Link (CXL) switch; A memory management method characterized by comprising the step of identifying the physical address of data stored in memory based on the offset between the address ranges of the two or more physical nodes.

2. The memory management method according to claim 1, characterized in that the two or more physical nodes each expose different address ranges corresponding to the shared memory area within the CXL memory expander.

3. The memory management method according to claim 1, further comprising the step of maintaining a memory allocation table that associates one or more pages in the shared memory area with the virtual node.

4. The memory management method according to claim 1, characterized in that the offset is determined based on the difference between the base addresses assigned to the two or more physical nodes.

5. The memory management method according to claim 1, characterized in that each of the two or more physical nodes is connected to the CXL switch via two or more CXL host adapters.

6. The memory management method according to claim 1, further comprising the step of obtaining a memory page using the aforementioned physical address.

7. The memory management method according to claim 1, further comprising the step of updating the page middle directory (PMD) to point to page table entries (PTEs) when a process is migrated between central processing unit (CPU) nodes.

8. The memory management method according to claim 1, further comprising the step of updating a page middle directory (PMD) entry to point to a page table entry (PTE) associated with a base address corresponding to a local CXL host adapter when a process is migrated between central processing unit (CPU) nodes.

9. The memory management method according to claim 1, further comprising the step of selecting a page table entry (PTE) associated with a base address corresponding to a local CXL host adapter in order to manage access to a shared memory area.

10. The memory management method according to claim 1, characterized in that the data stored in the memory is accessed with lower latency compared to accessing the memory without using the virtual node.

11. A device for memory management in a computing system, A compute express link (CXL) switch configured to connect two or more physical nodes, Equipped with a processor, The aforementioned processor, A virtual node is generated by combining the two or more physical nodes mentioned above. A device for memory management, characterized in that it is configured to identify the physical address of data stored in the memory based on the offset between the address ranges of the two or more physical nodes.

12. The memory management apparatus according to claim 11, characterized in that the two or more physical nodes each expose different address ranges corresponding to the shared memory area in the CXL memory expander.

13. The device for memory management according to claim 11, further characterized in that the processor is configured to maintain a memory allocation table that associates one or more pages in the shared memory area with the virtual node.

14. The device for memory management according to claim 11, characterized in that the offset is determined based on the difference between the base addresses assigned to the two or more physical nodes.

15. It also has two or more CXL host adapters, The device for memory management according to claim 11, characterized in that each of the two or more physical nodes is coupled to the CXL switch via the two or more CXL host adapters.

16. The device for memory management according to claim 11, further characterized in that the processor is configured to acquire a memory page using the physical address.

17. The device for memory management according to claim 11, further characterized in that the processor is configured to update the page middle directory (PMD) to point to page table entries (PTEs) when a process is migrated between central processing unit (CPU) nodes.

18. The device for memory management according to claim 11, further characterized in that the processor is configured to update a page middle directory (PMD) entry to point to a page table entry (PTE) associated with a base address corresponding to a local CXL host adapter when a process is migrated between central processing unit (CPU) nodes.

19. The device for memory management according to claim 11, further characterized in that the processor is configured to select a page table entry (PTE) associated with a base address corresponding to a local CXL host adapter in order to manage access to a shared memory area.

20. The device for memory management according to claim 11, characterized in that the data stored in the memory is accessed with lower latency compared to accessing the memory without using the virtual node.