Multi-machine system synchronization method based on CXL shared cache consistent memory
By defining spin locks and mutex locks in CXL shared cache coherent memory, combined with hardware atomic operations and PCIe MSI interrupt mechanism, the synchronization problem of CXL shared memory across host processes in multi-machine systems is solved, achieving safe concurrent access and avoiding data contention and system crashes.
Patent Information
- Application Number
- CN202510824051.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-03
AI Technical Summary
The existing Linux synchronization mechanism cannot effectively solve the concurrent access of cross-host processes to CXL shared memory in a multi-machine system, resulting in data competition and system crash risks, and cannot guarantee the security and consistency of cross-host access.
Spin lock and mutex variables are defined in CXL shared cache coherent memory. Hardware-supported atomic operation primitives and PCIe MSI interrupt mechanisms are used to achieve synchronized access across host systems. Spin lock and mutex mechanisms are used to manage lock states and wait queues to ensure secure concurrent access of processes.
It enables secure and efficient synchronous access to CXL shared memory across host processes in a multi-machine system, avoiding data contention and system crashes, and improving the security and consistency of data access.
Smart Images

Figure CN120743577A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of high-performance computing and memory interconnection, and in particular to a multi-machine system synchronization method based on CXL shared cache coherent memory. Background Art
[0002] Emerging applications such as high-performance computing, deep learning, and graphics computing are placing increasing demands on the capacity and performance of computer memory resources. On the one hand, the growth rate of traditional DRAM density is gradually slowing, and the cost and complexity of expanding to larger capacities are increasing. Memory bandwidth increases are lagging behind computing performance, and the pin efficiency of parallel DDR interfaces is low. Adding DDR channels for expansion leads to high pin counts and complex wiring. On the other hand, compute and storage resources in current data center servers are statically configured according to a fixed ratio, while different applications exhibit different compute-to-memory ratios. For example, deep learning typically involves multiple stages: feature extraction, model training, and inference, each with unique resource requirements. This varying compute-to-memory ratio leads to unused memory, and the cost of unused memory has become a major component of hyperscale data center infrastructure maintenance expenses. This creates a resource mismatch between the static resource allocation of servers and the dynamic demands of emerging applications, necessitating dynamic resource allocation.
[0003] The emerging Compute Express Link (CXL) interconnect technology provides a new solution for memory expansion, dynamic resource management, and data sharing between heterogeneous devices. CXL is an open, industry-standard interconnect technology. CXL 1.x defines a series of subprotocols, including CXL.io, CXL.cache, and CXL.mem. These protocols, combined, support high-bandwidth, low-latency connections between host processors (CPUs) and heterogeneous devices such as accelerators (such as GPUs and FPGAs), intelligent network interface cards (NICs), and memory expansion devices. CXL 1.x enables processors and devices to share data and operate on shared data within the same memory region, improving data access performance, reducing data movement, and reducing software stack complexity. Building on the widespread success of CXL 1.x, CXL 2.0 supports additional features, including memory pooling and single-stage switching, while maintaining backward compatibility. Memory pooling allows multiple hosts to dynamically allocate memory from a pool of memory devices. Single-stage switching allows multiple devices to connect to a single host through a switch while maintaining low latency and high bandwidth. Building on CXL 2.0, CXL 3.x introduces multi-level hierarchical switching capabilities, allowing complex networks to be constructed through cascading and fan-out. Furthermore, CXL 3.x introduces shared cache-coherent memory (CXL Shared Coherent Memory), which allows the same memory region in a remote memory pool to be mapped to the physical address space of multiple hosts. Multiple host systems can simultaneously access and update a shared cache-coherent memory region within the same coherence domain, ensuring that each host can read the latest data in a timely manner. Cache coherence in shared memory regions is managed by hardware through a back invalidation (BI) mechanism, enabling devices to implement snoop filters to track cache lines cached in peer caches.
[0004] The new shared cache-coherent memory feature provides a low-overhead data transfer and synchronization mechanism, allowing multiple systems to efficiently share data, perform synchronization operations, and pass messages. By eliminating the need to copy communication data, CXL can reduce latency by an order of magnitude compared to Ethernet and RDMA-based data center networks. However, shared access to data in CXL shared memory by multiple systems requires a synchronization mechanism that supports concurrent access to ensure the correct reading and writing of shared data.
[0005] Currently, a variety of synchronization mechanisms have been provided in the Linux system, such as atomic operations, spin locks, semaphores, etc. These synchronization mechanisms are designed for a single-machine system multi-core environment to support the correct access of multiple processes executed in parallel to shared data, such as list modification operations. However, the scope of the existing synchronization mechanism is limited to a single system, and is used to implement access to shared data by multiple processes in a system. For the situation where multiple processes belonging to multiple independent host systems concurrently access data in CXL shared memory, this cross-host parallel access requirement exceeds the application scope of the existing Linux synchronization mechanism. Direct use of these synchronization mechanisms may lead to data contention, inconsistent access, and even system crashes, and the security of cross-host access on CXL shared memory cannot be guaranteed. To this end, the present invention proposes a multi-machine system synchronization method based on CXL shared cache consistent memory, aiming to build a safe and efficient synchronization mechanism for multi-system parallel access to CXL shared memory data, and realize synchronized access to data in CXL shared memory by multiple processes across hosts.
[0006] After searching, the application publication number CN101794271 B is a method and device for implementing multi-core memory consistency. Among them, a method for implementing multi-core memory consistency includes: the second-level cache of the first processor group receives a control signal from the first processor group to read the first data; if the first data is currently maintained by the second processor group, the first data is read from the first-level cache of the second processor group through the fast consistency interface of the first-level cache of the second processor group, wherein the second-level cache of the first processor group is connected to the fast consistency interface of the first-level cache of the second processor group; the read first data is provided to the first processor group for processing through the second-level cache of the first processor group. The technical solution of this embodiment of the invention solves the problem of memory consistency between clusters of the ARM Cortex-A9 architecture.
[0007] The differences between that invention and the present invention are as follows: 1) The searched invention focuses on cache coherence between multi-core processors within the ARM Cortex-A9, which pertains to inter-core collaboration within a single-machine system. It relies on a specific hardware architecture and interconnected ARM L2 caches, cannot be extended to multi-machine systems, and does not involve cross-host memory sharing scenarios. The present invention addresses the synchronization issue when sharing memory across independent hosts using the CXL protocol, addresses data synchronization issues in shared memory across multi-host systems, and is not tied to a specific processor architecture. Summary of the Invention
[0008] The present invention aims to solve the above problems in the prior art. It proposes a multi-machine system synchronization method based on CXL shared cache coherent memory. The technical solution of the present invention is as follows:
[0009] A multi-machine system synchronization method based on CXL shared cache coherent memory, comprising:
[0010] Step 1, defining and managing a spin lock variable in the CXL shared cache coherent memory;
[0011] Step 2: Use the atomic operation primitives supported by the hardware to acquire and release the spin lock, and the atomic operation primitives ensure the integrity of the acquisition and release operations;
[0012] Step 3: Before accessing the shared resource, check the spin lock status. If the spin lock is occupied, busy wait until the spin lock is successfully acquired.
[0013] Step 4: define and manage mutex variables and wait queues in the CXL shared cache coherent memory;
[0014] Step 5: Use the atomic operation primitives supported by the hardware to acquire and release the mutex lock, and the atomic operation primitives ensure the integrity of the acquisition and release operations;
[0015] Step 6: Notify the host of the change in the state of the mutex lock through the PCIe MSI interrupt mechanism, and trigger the interrupt mechanism when the state of the mutex lock changes from occupied to unoccupied.
[0016] Step 7: monitor the state of the mutex lock, and when the state of the mutex lock changes, wake up the process in the waiting queue so that the process can access the shared resource.
[0017] Furthermore, in step 1, a spin lock variable is defined and managed in the CXL shared cache coherent memory, and the spin lock variable includes a lock status bit, a lock owner identifier, and a reference counter, wherein the lock status bit is used to indicate whether the lock is occupied, the lock owner identifier is used to record the host system ID that occupies the lock, and the reference counter is used to record how many systems are currently referencing the spin lock variable.
[0018] Furthermore, in step 3, before accessing the shared resource, busy waiting is performed until the spin lock is successfully acquired. The busy waiting strategy includes a cyclic comparison and exchange operation and a spin test lock status bit. The cyclic comparison and exchange operation uses a hardware atomic operation primitive to ensure the atomicity of lock acquisition. The spin test lock status bit ensures that the lock is acquired immediately after the lock is released, thereby preventing the idle state of the lock from being discovered by multiple competitors at the same time.
[0019] Furthermore, in step 4, a mutex lock variable and a waiting queue are defined and managed in the CXL shared cache coherent memory. The mutex lock variable includes a lock status bit, a lock owner identifier, a reference counter, and a waiting queue. The waiting queue is used to manage host system threads that cannot immediately access shared resources because the lock is occupied, ensuring that they can be awakened according to a certain strategy after the lock is released.
[0020] Furthermore, in step 6, the mutex state change is notified across hosts via a PCIe MSI interrupt mechanism, wherein the PCIe MSI interrupt mechanism includes a PCIe MSI-X interrupt mechanism, wherein the PCIe MSI-X interrupt mechanism implements an efficient notification and process wake-up mechanism across host systems by sending an interrupt signal to a PCIe switch or a downstream device.
[0021] Furthermore, in step 7, the process in the waiting queue is awakened, and the process awakening mechanism includes but is not limited to: a priority awakening strategy based on lock state changes, dynamically adjusting the awakening order of processes in the waiting queue; a fair awakening strategy based on lock state changes, ensuring that all waiting processes have the opportunity to access shared resources; a first-in-first-out (FIFO) awakening strategy based on lock state changes, waking up the processes in the order in which they join the waiting queue.
[0022] Furthermore, the CXL shared cache coherent memory supports a device DAX mode, wherein the device DAX mode allows user space applications to directly access the CXL shared memory without going through a traditional file system interface, thereby improving data access efficiency and flexibility.
[0023] Furthermore, it also includes the management of the shared memory, specifically including: shared memory allocation and recovery interface, used to support shared memory application scenarios of different granularities, coarse-grained shared memory management and fine-grained shared memory management, where coarse-grained shared memory management supports shared memory allocation with a granularity greater than 4KB, and fine-grained shared memory management supports space allocation of integer multiples of cache size 64B granularity.
[0024] The advantages and beneficial effects of the present invention are as follows:
[0025] This paper proposes a solution for shared memory synchronization in multi-host systems by extending native atomic operations and the CXL consistency protocol. Its advantages lie in the fact that traditional Linux synchronization mechanisms (such as mutexes and semaphores) are only applicable to single-host multi-core environments. By leveraging CXL hardware-supported multi-host cache coherence and software co-design, this paper designs a low-latency, highly reliable synchronization method for multiple hosts. This method enables secure concurrent access to shared memory across multiple independent hosts, avoiding the risk of data races and system crashes. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 This is a schematic diagram of the architecture of a multi-machine system synchronization method based on CXL shared cache coherent memory according to a preferred embodiment of the present invention;
[0027] Figure 2 Schematic diagram of the shared spin lock mechanism based on CXL shared cache coherent memory;
[0028] Figure 3 Flowchart of the shared spin lock mechanism based on CXL shared cache consistent memory;
[0029] Figure 4 Schematic diagram of a shared mutex mechanism based on CXL shared cache coherent memory;
[0030] Figure 5 Flowchart of the shared mutex mechanism based on CXL shared cache coherent memory. DETAILED DESCRIPTION
[0031] The following will describe the technical solutions in the embodiments of the present invention in detail with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention.
[0032] The technical solution of the present invention to solve the above technical problems is:
[0033] A multi-machine system synchronization method based on CXL shared cache coherent memory includes two mechanisms: a shared spin lock mechanism based on CXL shared cache coherent memory and a shared mutex lock mechanism based on CXL shared cache coherent memory.
[0034] 1. The shared spin lock mechanism based on CXL shared cache coherent memory uses a busy-wait approach to synchronize spins in multi-processor systems for competing access to shared resources in CXL shared memory. The mechanism defines spin locks in CXL shared memory and uses hardware atomic primitives to synchronize short-term access to CXL shared resources across multiple systems. This mechanism is suitable for scenarios with short critical sections and predictable wait times, such as concurrency control in interrupt contexts and kernel threads.
[0035] The shared spinlock mechanism based on CXL shared cache-coherent memory allocates storage space for the spinlock variable from CXL shared cache-coherent memory and defines the synchronization spinlock variable in CXL shared memory. Leveraging cache-coherent shared memory and hardware-supported atomic operation primitives across multiple systems, the processors in the multi-machine system can promptly access the value of the shared spinlock variable. Based on the value of the shared spinlock variable, the system determines whether the resource currently being accessed in CXL shared memory is available. If unavailable, the shared spinlock variable value is repetitively tested until the shared resource becomes available.
[0036] 2. The shared mutex mechanism based on CXL shared cache-coherent memory uses a blocking sleep mechanism for multi-machine, multi-processor systems competing for access to shared resources in CXL shared memory. It includes a CXL shared memory mutex module and a shared memory-based multi-machine system notification module. Compared to spin locks, this significantly reduces CPU resource waste and is suitable for scenarios with intense lock contention and long critical sections, such as list operations.
[0037] The CXL shared memory mutex lock module allocates storage space for the mutex lock variable from CXL shared cache-coherent memory and defines the synchronization mutex lock variable in CXL shared memory. Leveraging cache-coherent shared memory and hardware-supported atomic operation primitives across multiple systems, the multi-machine system processors can promptly access the value of the shared mutex lock variable. The value of the shared mutex lock variable is then determined based on whether it is 1. If it is 1, indicating that the shared data controlled by the lock is available, the current host process can acquire the shared mutex lock and access the shared resource. Otherwise, indicating that the shared mutex lock is already occupied, the current process is blocked and suspended, added to the shared mutex lock waiting queue, and released after the process occupying the shared mutex lock completes access.
[0038] The multi-machine system notification module based on shared memory is used to realize efficient event notification and process awakening between multi-machine systems, ensuring that after a host processor releases a shared mutex, the waiting processes suspended in the shared mutex waiting queue, local processes or processes in other host systems, can be awakened to access shared memory resources to continue execution. Specifically, the module monitors the state changes of the mutex variables. When a process in a host system releases the lock, it notifies other host systems through the PCIe MSI interrupt mechanism based on the interrupt triggered by writing to shared memory. The host system that receives the notification determines whether the process at the head of the mutex waiting queue belongs to the local system. If so, it wakes up the corresponding waiting process, occupies the shared mutex to access the data in the shared memory, and realizes synchronous access to CXL shared memory data across multiple host system processes.
[0039] The present invention proposes a multi-system synchronization method based on CXL shared cache consistent memory. Through a shared spin lock mechanism based on CXL shared cache consistent memory and a shared mutex mechanism based on CXL shared cache consistent memory, the two synchronization mechanisms are applied in different scenarios to achieve synchronous access to data in CXL shared memory by multiple processes in multiple machine systems.
[0040] Figure 1 This is a schematic diagram of the architecture of a multi-machine system synchronization method based on CXL shared cache coherent memory. Figure 1As shown, multiple hosts simultaneously map a shared memory region into the system through CXL, enabling memory sharing between hosts. Cache coherence is ensured by CXL 3.0 hardware, providing a consistent view of data in different hosts' caches and shared memory. Hosts can perform atomic operations by obtaining exclusive control of cache lines within their caches. After data is updated, it is shared globally through cache coherence. From the host perspective, there are two ways to manage CXL shared memory. The first is device DAX mode, in which CXL shared memory is exposed as a DAX device. Userspace applications can open the DAX device file using the open() system call and map the shared memory into the user process's address space using the mmap() system call. This mode provides user programs with flexibility and varying granularity in managing CXL memory. However, its disadvantage is that it requires modifications to existing programs and can only be used by applications that have opened DAX device-mapped shared memory. The second mode is system RAM mode, in which CXL shared memory is integrated into the system's memory hierarchy and managed by the operating system as a CPU-less NUMA memory. In this mode, CXL shared memory behaves like traditional DRAM; all applications can access it through a standard memory allocation interface, achieving wider compatibility without requiring application changes. In this invention, each host manages CXL shared memory using the device DAX mode. The host system obtains the starting virtual address, starting physical address, and size of the CXL shared memory using the device direct_access() function in the kernel.
[0041] Figure 1 The lower middle section illustrates the management of CXL shared memory space. First, global shared memory information is stored in the 4KB space at the beginning of the shared memory. This information includes the shared memory space size, shared memory block size, shared spinlock starting offset, number of available shared spinlocks, shared mutex starting offset, number of available shared spinlocks, free shared memory block offsets, total number of free shared memory blocks, and the current number of free memory blocks. Next, a fixed-size shared spinlock and shared mutex array space is provided. The starting offset and shared spinlock number of this space allow for quick location of shared spinlock and mutex variables. Finally, the shared memory data storage space is managed at a 4KB page granularity. Free shared memory is managed using a linked list. To support shared memory application scenarios with varying granularity, shared memory management is divided into coarse-grained shared memory management and fine-grained shared memory management. Coarse-grained shared memory management supports shared memory allocations with granularity greater than 4KB, while fine-grained shared memory management supports allocations with granularity that is an integer multiple of the 64B cache size. It provides shared spin locks, shared memory mutex locks, and shared memory allocation and recovery interfaces respectively.
[0042] Figure 2 This diagram illustrates the shared spinlock mechanism based on CXL shared cache-coherent memory. Each shared mutex contains fields including a name string, a spinlock variable, and a reference counter. The name allows multiple systems to reference the same shared mutex variable. The spinlock variable controls access to shared resources. The reference counter records the number of systems currently referencing the shared spinlock variable. Different host systems request shared spinlocks through the shared spinlock allocation interface. After passing the shared spinlock name, the host system allocates a shared spinlock variable from the shared spinlock array and associates the name with the spinlock number. Other host systems are then associated with the same shared spinlock using the same shared spinlock name. Thanks to the hardware-level cache coherence features and atomic operation primitives supported by CXL 3.x, cache data remains consistent across multiple host systems. Modifications to the shared spinlock variable by one host are visible to other host systems. Based on the value of the shared spinlock variable, the host system determines whether the resource currently being accessed in CXL shared memory is available. If unavailable, the host system repeatedly tests the value of the shared spinlock variable until the shared resource becomes available.
[0043] Figure 3 This is a flowchart of the shared spin lock mechanism based on CXL shared cache coherent memory. When multiple host systems share CXL memory, locks are required when multiple kernel threads within the same host or different kernel threads on multiple hosts perform read and modify operations on data in shared memory. Based on whether access to the data is interrupted and the time it takes to access the data, a shared spin lock is selected to protect access to the shared data. The process for accessing shared memory data using a shared spin lock is as follows:
[0044] In step S301, before accessing shared memory data, a shared spin lock is requested through a shared spin lock allocation function. The input parameter is the shared spin lock name. With the same name, multiple systems can point to the same shared spin lock.
[0045] In step S302, in the shared spin lock allocation function, the shared spin lock area is searched by name to locate the shared spin lock number. If an existing shared spin lock variable is found, step S304 is executed. Otherwise, step S303 is executed.
[0046] In step S303, no shared spin lock variable with the same name is found, and a new shared spin lock variable is allocated from the idle shared spin lock array.
[0047] In step S304, a shared spin lock variable with the same name is found, a reference is associated with the shared spin lock with the same name in other systems, and the shared spin lock reference counter is incremented by 1.
[0048] In step S305 , all system kernel paths accessing shared memory data determine whether the value of the shared spin lock variable is 0. If it is 0, step S306 is executed. Otherwise, step S305 is executed again.
[0049] In step S306 , a system kernel path occupies the shared spin lock, sets the value of the shared spin lock from 1 to 0, and accesses the shared memory data.
[0050] In step S307 , after accessing the shared memory data, the system kernel path occupying the shared spin lock releases the shared spin lock and sets the value of the shared spin lock variable from 0 to 1.
[0051] In step S308 , the shared spin lock reference counter is decremented by 1, and the value of the reference counter is determined to be 0. If so, step S308 is executed, otherwise the process ends.
[0052] In step S309, the shared spin lock name field is set to null, and the shared spin lock variable is reclaimed.
[0053] Figure 4 This diagram illustrates a shared mutex mechanism based on CXL shared cache coherent memory. This shared mutex, based on a shared spinlock, adds a wait queue and queue access spinlock properties, allowing multiple host system kernel paths to compete for access to shared memory data through a busy-sleep mechanism. Different host systems apply for a shared mutex through the shared mutex allocation interface. Multiple host systems reference the same shared mutex by name. When accessing shared memory data, if the shared memory data is available, the shared memory data is locked for access. Otherwise, the kernel thread is added to the shared mutex wait queue, where it sleeps and waits for a scheduled wakeup to access shared memory resources. Because the awakened kernel thread may have data from a different host system, the present invention uses a PCIe MSI interrupt mechanism based on a write-to-shared memory interrupt to notify other host systems and wake up the suspended kernel control path threads belonging to that system. Portions of shared memory are mapped to a PCIe device. When new data is written to the shared memory space, a PCIe MSI interrupt is triggered. The interrupt handler is executed, adding the kernel threads belonging to the system in the shared mutex queue to the execution queue. After execution, the kernel threads occupy the shared mutex and access the shared memory data.
[0054] Figure 5 This is a flowchart of the shared mutex mechanism based on CXL shared cache coherent memory. Based on whether the access to the data is interrupted and the time it takes to access the data, a shared mutex is selected to protect access to the shared data. The process is as follows:
[0055] In step 5301, before accessing shared memory data, a shared mutex is requested through a shared mutex allocation function. The input parameter is the shared mutex name. Multiple systems point to the same shared mutex with the same name.
[0056] In step S502, in the shared mutex allocation function, the shared mutex area is searched by name to locate the shared mutex number. If an existing shared mutex variable is found, step S304 is executed. Otherwise, step S303 is executed.
[0057] In step S503, if no shared mutex lock variable with the same name is found, a new shared mutex lock variable is allocated from the idle shared spin lock array.
[0058] In step S504, a shared mutex variable with the same name is found, a reference is associated with the shared mutex with the same name in other systems, and the shared mutex reference counter is incremented by 1.
[0059] In step S505 , all system kernel paths accessing shared memory data determine whether the value of the shared mutex variable is 0, and if so, execute step S506 , otherwise execute step S507 .
[0060] In step S506, a system kernel path occupies the shared mutex lock, sets the value of the shared spin lock from 1 to 0, and accesses the shared memory data. Then, step S508 is executed.
[0061] In step S507, the shared mutex is already occupied, and the system kernel path thread is added to the waiting list of the shared mutex, and the kernel thread is suspended to wait for other system kernel threads to complete access to the shared memory data.
[0062] In step S508 , the shared mutex is released, and the value of the shared mutex is set from 0 to 1.
[0063] In step S509, the shared mutex lock reference counter is decremented by 1. If the reference counter value is 0, step S510 is executed. Otherwise, step S511 is executed.
[0064] In step S510 , the shared mutex variable name field is set to null, and the shared mutex variable is recycled.
[0065] In step S511, it is determined whether the thread at the head of the shared mutex variable waiting queue belongs to the local system, if so, executing step S506. Otherwise, executing step S512.
[0066] In step S512, the first thread in the waiting queue is awakened across the host system through a notification mechanism based on triggering a PCIe MSI interrupt when writing to the shared memory, and then step S506 is executed.
[0067] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions.
[0068] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0069] The above embodiments should be understood as merely illustrating the present invention and not as limiting the scope of protection of the present invention. After reading the contents of the present invention, technicians may make various changes or modifications to the present invention, and these equivalent changes and modifications also fall within the scope defined by the claims of the present invention.
Claims
1. A multi-machine system synchronization method based on CXL shared cache consistent memory, characterized in that: The following steps are involved: Step 1, defining and managing a spin lock variable in the CXL shared cache coherent memory; Step 2: Use the atomic operation primitives supported by the hardware to acquire and release the spin lock, and the atomic operation primitives ensure the integrity of the acquisition and release operations; Step 3: Before accessing the shared resource, check the spin lock status. If the spin lock is occupied, busy wait until the spin lock is successfully acquired. Step 4: define and manage mutex variables and wait queues in the CXL shared cache coherent memory; Step 5: Use the atomic operation primitives supported by the hardware to acquire and release the mutex lock, and the atomic operation primitives ensure the integrity of the acquisition and release operations; Step 6: Notify the host of the change in the state of the mutex lock through the PCIe MSI interrupt mechanism, and trigger the interrupt mechanism when the state of the mutex lock changes from occupied to unoccupied. Step 7: monitor the state of the mutex lock, and when the state of the mutex lock changes, wake up the process in the waiting queue so that the process can access the shared resource.
2. A multi-machine system synchronization method based on CXL shared cache coherent memory according to claim 1, characterized in that: In step 1, a spin lock variable is defined and managed in the CXL shared cache coherent memory. The spin lock variable includes a lock status bit, a lock owner identifier, and a reference counter. The lock status bit is used to indicate whether the lock is occupied, the lock owner identifier is used to record the host system ID that occupies the lock, and the reference counter is used to record how many systems are currently referencing the spin lock variable.
3. The multi-machine system synchronization method based on CXL shared cache coherent memory according to claim 1, characterized in that: In step 3, before accessing the shared resource, busy waiting is performed until the spin lock is successfully acquired. The busy waiting strategy includes a cyclic comparison and exchange operation and a spin test lock status bit. The cyclic comparison and exchange operation uses a hardware atomic operation primitive to ensure the atomicity of lock acquisition. The spin test lock status bit ensures that the lock is acquired immediately after the lock is released, thereby preventing the idle state of the lock from being discovered by multiple competitors at the same time.
4. The multi-machine system synchronization method based on CXL shared cache coherent memory according to claim 1, characterized in that: In step 4, a mutex lock variable and a waiting queue are defined and managed in the CXL shared cache coherent memory. The mutex lock variable includes a lock status bit, a lock owner identifier, a reference counter, and a waiting queue. The waiting queue is used to manage host system threads that cannot immediately access shared resources due to the lock being occupied, ensuring that they can be awakened according to a certain strategy after the lock is released.
5. The multi-machine system synchronization method based on CXL shared cache coherent memory according to claim 1, characterized in that: In step 6, the mutex state change is notified across hosts through a PCIe MSI interrupt mechanism, wherein the PCIe MSI interrupt mechanism includes a PCIe MSI-X interrupt mechanism, wherein the PCIe MSI-X interrupt mechanism implements an efficient notification and process wake-up mechanism across host systems by sending an interrupt signal to a PCIe switch or a downstream device.
6. The multi-machine system synchronization method based on CXL shared cache coherent memory according to claim 1, characterized in that: In step 7, the processes in the waiting queue are woken up, and the process wake-up mechanism includes but is not limited to: a priority wake-up strategy based on lock state changes to dynamically adjust the wake-up order of processes in the waiting queue; a fair wake-up strategy based on lock state changes to ensure that all waiting processes have the opportunity to access shared resources; a first-in-first-out (FIFO) wake-up strategy based on lock state changes to wake up the processes in the order in which they join the waiting queue.
7. The multi-machine system synchronization method based on CXL shared cache coherent memory according to claim 1, characterized in that: The access mode supported by the CXL shared cache coherent memory is the device DAX mode, wherein the device DAX mode allows user space applications to directly access the CXL shared memory without going through the traditional file system interface, thereby improving data access efficiency and flexibility.
8. The multi-machine system synchronization method based on CXL shared cache coherent memory according to claim 1, characterized in that: It also includes steps for managing the shared memory, specifically including: shared memory allocation and recovery interfaces for supporting shared memory application scenarios of different granularities, coarse-grained shared memory management and fine-grained shared memory management, wherein coarse-grained shared memory management supports shared memory allocation with a granularity greater than 4KB, and fine-grained shared memory management supports space allocation of integer multiples of the cache size 64B granularity.
Citation Information
Patent Citations
Implementation method and device of consistency of multi-core internal memory
CN101794271B
Cited By
A cross-node hierarchical blocking lock and fault recovery method based on CXL shared memory
CN122653905A