Smart network interface controller unloading-based remote memory system
By using intelligent network cards to uninstall memory node management unit and unilateral RDMA communication in remote memory systems, the CPU overhead problem of user-mode runtime library method is solved, and the invisible access and low-power management of memory nodes is realized.
Patent Information
- Application Number
- PCT/CN2024/104879
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-02
- Filing Date
- 2024-07-11
- Publication Date
- 2025-07-10
AI Technical Summary
The existing remote memory system with user-mode runtime library method increases CPU computing management overhead and network communication overhead on memory nodes, and cannot adapt to the separate architecture of memory.
The function of uninstalling the remote memory node management unit of the intelligent network card is adopted, and the memory management is implemented using the intelligent network card's SoC, and the memory of the memory node is directly read through single-sided RDMA to reduce the CPU participation of the memory node.
It effectively reduces the CPU computing management overhead of memory nodes, realizes invisible access to memory nodes, and reduces the power bill and CPU resource usage during runtime.
Smart Images

Figure CN2024104879_10072025_PF_FP_ABST
Abstract
Description
A remote memory system based on smart network card offloading Technical Field
[0001] The present invention relates to the field of communication technology, and in particular to a remote memory system based on smart network card offloading. Background Art
[0002] Memory is the most constrained and fiercely contested resource in today's data centers. With the increasing popularity of in-memory workloads such as machine learning applications and key-value storage, user demand for memory in data center computing clusters is exploding. Insufficient server memory resources have become a major bottleneck affecting application performance.
[0003] One approach to resolving this memory bottleneck is to transform memory on other servers into remote memory for local applications, creating a cross-server memory pooling architecture. This allows compute nodes to access memory on remote memory nodes, meaning servers are no longer limited to local memory and can instead utilize free memory elsewhere in the cluster. Thanks to the rapid development of network performance, the emerging RDMA (Remote Direct Memory Access) technology, short for remote direct memory access, was developed to address server-side data processing delays during network transmission. As a high-throughput, low-latency network communication technology, it is widely used in remote memory systems.
[0004] Traditional remote memory systems are based on the Linux kernel's virtual memory abstraction. When a required memory page is not in local memory, a page fault is triggered. The page fault handler detects whether remote access is required, then retrieves the required memory page from the remote memory node, transfers it over the network to the local memory cache, and ultimately makes the data available to the application. However, accessing remote memory through kernel page fault processing has the following bottlenecks: ① The kernel page fault handler must enter kernel state, increasing overhead due to frequent context switching; ② The kernel lacks application semantics, limiting the granularity of page retrieval to the kernel page size (4KB). Therefore, accessing small-granularity objects suffers from severe read-write amplification (LA) issues, as the network must transfer at least a page size, wasting network bandwidth.
[0005] To address the limitations imposed by the kernel on traditional remote memory systems, one popular solution is to access remote memory through a user-mode runtime library. This approach offers the advantage of allowing applications to program using exposed key-value or data structure interfaces, effectively leveraging application semantics and enabling fine-grained object-level access to remote data. This avoids the high overhead of kernel page fault handling and addresses the problem of read-write amplification. The AIFM designed by Zhenyuan Ruan et al. is used as an example to illustrate a mainstream user-mode runtime library-based remote memory architecture. As shown in Figure 1, this approach utilizes a fast, low-overhead pointer abstraction that supports remote access. It exposes data structure programming interfaces that convey semantic information to the runtime library for compute node applications, enabling applications written based on this runtime library to access remote memory at an object-level, independent of kernel page fault handling. A memory node management unit is placed remotely to process remote memory access requests. Based on the object ID, it locates the address of the corresponding data in the data address management module, thereby facilitating object-level access to the remote memory node. This design's network communication is based on the user-mode DPDK's TCP network protocol stack, bypassing the kernel and reducing kernel context switches. However, to ensure performance, DPDK uses thread polling to receive packets, which consumes a large amount of CPU resources. Another approach, FaRM, uses unilateral RDMA semantics for communication between compute nodes and memory nodes, but the memory nodes still require the CPU to poll the request buffer to receive data.
[0006] Previous research has shown that memory nodes in a split architecture have very limited CPU resources. However, the current user-mode runtime library approach imposes two major overheads on the computing resources of memory nodes:
[0007] 1. Memory node management unit overhead: The management unit on the remote memory node needs to create a service process dedicated to remote memory request processing, which consumes the CPU resources of the memory node and results in additional computing overhead in addition to the memory provided by the memory node.
[0008] 2. Network communication overhead: User-mode network protocol stacks based on DPDK or RDMA bypass the kernel, but they typically use CPU polling mode, which causes the CPU providing network services to be 100% occupied at all times, resulting in huge overhead for the limited CPU resources of memory nodes.
[0009] As can be seen, current remote memory systems based on user-mode runtimes are not adaptable to architectures with fully separated storage and computing. Emerging Smart NICs are already widely used in data centers. For example, previous research has utilized the SoC (System on a Chip) on these cards to accelerate distributed transactions.
[0010] Therefore, technicians in this field are committed to developing a remote memory system based on smart network card offloading, using the SoC on the smart network card to offload the functions of the remote memory node management unit, and directly read the memory of the memory node through the SoC unilateral RDMA, realizing remote memory network management and seamless access to memory nodes.
[0011] Summary of the Invention
[0012] In view of the above-mentioned defects of the prior art, the technical problem to be solved by the present invention is that the memory node management unit and the memory node network protocol stack increase the computing management overhead of the CPU.
[0013] To achieve the above objectives, the present invention provides a remote memory system based on smart network card offloading, characterized in that it includes a computing node, a smart network card, and a memory node, wherein the computing node sends a remote memory access request to the smart network card; the smart network card is used to parse the remote memory access request and access the memory node; the memory node is used to initialize and save the remote memory; the smart network card includes a remote memory request processing module, a data address management module, and an RDMA communication front end, and the memory node includes an RDMA communication back end; the use phases of the remote memory system include an initialization phase, an application registration phase, and an operation phase.
[0014] Furthermore, the RDMA communication backend is used to specify the size of the memory area during initialization, and send the virtual base address and remote key of the memory area to the smart network card.
[0015] Furthermore, the remote memory request processing module includes RDMA bilateral communication function, request data parsing function and memory node interaction function; the data address management module is used to manage the memory area offset mapping table; the RDMA communication front end is used to store the virtual base address, remote key and the sending queue and completion queue for RDMA communication with the memory area.
[0016] Furthermore, the initialization phase includes:
[0017] The RDMA communication front end initiates an RDMA initialization connection request to the RDMA communication back end, completes the RDMA queue pair connection, creates a completion queue corresponding to the queue pair, and creates a buffer for receiving data transmitted from the memory node through the queue pair;
[0018] The memory node registers a memory area according to the size of the RDMA initialization connection request;
[0019] The memory node sends the virtual base address and remote key of the registered memory area to the buffer on the smart network card through a queue pair;
[0020] The smart network card stores the virtual base address and the remote key in the RDMA communication front end.
[0021] Furthermore, the application registration phase includes:
[0022] The application of the computing node starts, initiates an RDMA communication connection to the smart network card, and then initiates a registration request to inform the smart network card of the required remote memory size;
[0023] The smart network card receives the registration request, allocates and records the starting position of the remote memory area available to the application, and creates an object ID to memory area offset mapping table for the application in the data address management module. All subsequent data using the remote memory space will be recorded in the memory area offset mapping table;
[0024] After registration is complete, a notification is returned to inform the computing node that the application can continue to run.
[0025] Furthermore, the operation phase includes:
[0026] Step 1: The thread running the application in the computing node initiates a bilateral RDMA remote memory read / write request;
[0027] Step 2: After receiving the request, the smart network card performs address translation to obtain the virtual address on the memory node;
[0028] Step 3: The SmartNIC initiates a unilateral RDMA read / write operation to the memory node, during which the memory node is unaware.
[0029] Step 4: The smart network card polls the completion queue to obtain the completion result, and returns the read and write result to the computing node through a bilateral RDMA request.
[0030] Furthermore, the step 1 further includes:
[0031] The computing node sends a request and simultaneously sends a buffer for receiving the result; the request data includes the operation type, object ID and data size. If the operation type is a write request, it also includes a data field to be written.
[0032] Furthermore, the step 2 further includes:
[0033] Step 2.1: The request sent by the computing node is placed in the request buffer of the remote memory request processing module; the remote memory request processing module polls the request notification queue to receive the computing node request, and then notifies the thread in the thread pool to process the request in the buffer;
[0034] Step 2.2: The remote memory request processing module parses the request data and searches for the memory area offset corresponding to the object ID field through the data address management module;
[0035] Step 2.3: The remote memory request processing module delivers the memory area offset and the requested data to the RDMA communication front end;
[0036] Step 2.4: The RDMA communication front end obtains the virtual address of the object according to the virtual base address and the memory area offset.
[0037] Furthermore, the step 3 further includes:
[0038] Step 3.1: The RDMA communication front end initiates a unilateral RDMA request through a virtual address and fills it into a sending queue;
[0039] Step 3.2: The RDMA hardware device reads and writes the locations corresponding to the virtual addresses on the memory nodes in sequence according to the order in the sending queue; after completion, the completion results are filled into the completion queue in the sending order.
[0040] Furthermore, the step 4 further includes:
[0041] The RDMA communication front end polls the completion queue, and the result of each request is returned to the computing node through RDMA bilateral send semantics. For write requests, the Smart NIC needs to add or modify the memory area offset mapping table of the data address management module after completion. For read requests, the Smart NIC needs to return the content of the read data to the computing node.
[0042] Compared with the prior art, the present invention has at least the following beneficial technical effects:
[0043] 1. This invention uses a SmartNIC to offload the memory management unit (MMU) of remote memory based on the user-mode runtime library, thus enabling in-network management of remote memory. Compute nodes communicate only with the SmartNIC, and the SmartNIC's SoC parses remote requests sent from the compute node and manages the address mapping of object IDs to actual data locations. This effectively reduces the computational management overhead of the CPU on the memory node when the remote memory system is enabled. The low-power SmartNIC SoC takes over the remote memory management function, effectively reducing electricity costs and other expenses during operation.
[0044] 2. This invention achieves seamless access to memory nodes. By using a smart network card to read and write data from the memory node via unilateral RDMA, the overhead of communicating with the compute node is transferred to the smart network card. This process is completely passive for the memory node, and the memory node CPU does not participate in the reading and writing process. Therefore, the CPU resource overhead is close to zero during operation.
[0045] The concept, specific structure and technical effects of the present invention will be further described below in conjunction with the accompanying drawings to fully understand the purpose, characteristics and effects of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] FIG1 is a schematic diagram of a system architecture of the prior art;
[0047] FIG2 is a schematic diagram of a system architecture of a preferred embodiment of the present invention;
[0048] FIG3 is a schematic diagram of the system use phases of a preferred embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following describes several preferred embodiments of the present invention with reference to the accompanying drawings to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the scope of protection of the present invention is not limited to the embodiments mentioned herein.
[0050] The smart network card realizes the integration of network and computing resources on the same card, and can achieve higher-performance network functions at a lower cost than ordinary CPUs. The present invention adds a smart network card to the memory node, and offloads the management unit overhead originally on the memory node to the SoC of the smart network card for centralized management of remote memory-related functions. The smart network card is used to receive and process remote memory requests, and the memory node is only responsible for running a resident process to initialize and save all remote memories. When the remote memory is running, the smart network card SoC sends a unilateral RDMA read and write request to the memory node. The memory node acts as a passive receiver, and the smart network card completes the processing of the request. The CPU does not need to do any processing on these requests.
[0051] This embodiment provides a remote memory system based on smart network card offloading, as shown in Figure 2, including computing nodes, smart network cards, and memory nodes. The computing node design follows the computing node design of AIFM. Applications need to be written based on the user-mode runtime library provided by AIFM on the computing node. When using remote memory, the computing node will send a remote memory access request to the smart network card. The RDMA communication backend is deployed on the memory node, which is responsible for specifying the size of the memory area during initialization. The base address of the memory area is BaseVA. The remote memory request processing module, data address management module, and RDMA communication front end are deployed on the smart network card.
[0052] The remote memory request processing module's main components include RDMA bilateral communication, request data parsing, and memory node interaction. The data address management module manages the mapping table between object IDs and memory region offsets. The RDMA communication frontend stores the BaseVA of the memory region on the memory node and the remote key R_Key required for access, as well as the send queue and completion queue for RDMA communication with the memory region.
[0053] The process after a remote memory access request is sent to the SmartNIC is as follows:
[0054] Step a: The compute node sends a request to the SmartNIC using the RDMA send semantics ibv_post_send and simultaneously sends the buffer containing the received result using the recv semantics.
[0055] Step b: Through the RDMA bilateral communication function, the request buffer of the remote memory request processing module has been issued in advance through RDMA's recv semantics, and the request will be placed in the request buffer; the thread on the SoC receives the compute node request through the ibv_poll_cq polling request notification queue, and then notifies the thread in the thread pool to process the request in the buffer.
[0056] Step c. The request is divided into multiple fields, including the operation type (read / write), object ID, and data size. If it is a write request, it also includes the data field to be written. By parsing the request data function, the remote memory request processing module parses the data request and searches for the memory area offset corresponding to the object ID through the data address management module for the object ID field. This process is the key translation process from the local object to the remote memory address.
[0057] Step d: Through the memory node interaction function, the remote memory request processing module delivers the offset and requested data to the RDMA communication front end.
[0058] Step e: The RDMA communication front end obtains the virtual address Obj-VA of the object according to BaseVA+offset, initiates a unilateral RDMA request through the virtual address ibv_post_send and fills it into the send queue.
[0059] Step f: The RDMA hardware device reads and writes the locations corresponding to the Obj-VA on the memory node in the order in which they are sent. This process does not require the participation of the memory node. Upon completion, the completion results are added to the completion queue in the order in which they are sent.
[0060] Step g: The RDMA communication front-end polls the completion queue. The result of each request is returned to the computing node through the RDMA bilateral send semantics. For write requests, the successful processing information is returned to the computing node, and the mapping table information in the data address management module is added or modified. For read requests, the content of the read data needs to be returned to the computing node.
[0061] This embodiment implements RDMA transparent access to memory on a memory node using the universal RDMA Verbs API. This is done according to the process shown in Figure 3. The process is divided into three phases: initialization, application registration, and operation. The detailed steps are as follows:
[0062] Step 1: Initialization phase
[0063] Step 1.1: The RDMA communication frontend on the SmartNIC initiates an RDMA initialization connection request to the RDMA communication backend on the memory node, completes the RDMA queue pair (QP) connection, and creates the completion queue (CQ) corresponding to the QP and the buffer used to receive data transmitted from the memory node through the QP.
[0064] Step 1.2: The memory node registers the RDMA memory region (MR) using the ibv_reg_mr() API according to the requested size.
[0065] In step 1.3, the memory node sends the virtual base address BaseVA of the registered memory area and the remote key R_Key required for remote access to the memory area to the buffer on the SmartNIC through the QP. This process is a bilateral RDMA operation.
[0066] Step 1.4: The SmartNIC extracts the BaseVA and R_Key from the buffer and saves them to the RDMA communication front end. With these two key pieces of information, subsequent unilateral RDMA access can be performed.
[0067] 2. Application registration stage
[0068] Step 2.1: The compute node application starts to initiate an RDMA communication connection to the Smart NIC, and then initiates a registration request to inform the Smart NIC of the required remote memory size.
[0069] Step 2.2: The SmartNIC receives the registration request, allocates and records the starting location of the remote memory area available to the application, and creates a mapping table from object ID to memory area offset for the application in the data address management module. All subsequent data using the remote memory space will be recorded in this table.
[0070] Step 2.3: After registration is complete, a notification is returned to inform the application that it can continue running.
[0071] 3. Operation phase
[0072] Step 3.1: The application thread initiates a bilateral RDMA remote memory read / write request.
[0073] Step 3.2: After receiving the request, the SmartNIC performs the address translation in step c to obtain the virtual address on the memory node.
[0074] Step 3.3: The SmartNIC initiates a unilateral RDMA read / write operation to the memory node. The memory node is unaware of this process.
[0075] Step 3.4: The SmartNIC polls the completion queue CQ to obtain the completion result.
[0076] Step 3.5: After the write operation is completed, the smart network card needs to update the mapping table of the data address management module.
[0077] Step 3.6: The SmartNIC returns the read and write results to the compute node via a bilateral RDMA request.
[0078] Step 3.7: The application thread continues to run seamlessly, and the entire running process is completed by the user-mode runtime library of AIFM on the computing node, and the application is unaware of this.
[0079] The above describes in detail the preferred embodiments of the present invention. It should be understood that those skilled in the art can make numerous modifications and variations based on the concepts of the present invention without inventive effort. Therefore, any technical solutions that can be derived by those skilled in the art through logical analysis, reasoning, or limited experimentation based on the concepts of the present invention and the prior art should be within the scope of protection defined by the claims.
Claims
1. A remote memory system based on intelligent network card offloading, characterized in that It includes computing nodes, smart network cards, and memory nodes. Among them, the computing nodes send remote memory access requests to the smart network cards; the smart network cards are used to parse the remote memory access requests and access the memory nodes; the memory nodes are used to initialize and save remote memory; the smart network cards include a remote memory request processing module, a data address management module, and an RDMA communication front end, and the memory nodes include an RDMA communication back end; the usage phases of the remote memory system include an initialization phase, an application registration phase, and a running phase.
2. The remote memory system based on intelligent network card offloading according to claim 1, wherein The RDMA communication back end is used to specify the size of the memory area during initialization and send the virtual base address and remote key of the memory area to the smart network card.
3. The remote memory system based on intelligent network card offloading according to claim 2, wherein The remote memory request processing module includes an RDMA bilateral communication function, a function of parsing request data, and a function of interacting with the memory node; the data address management module is used to manage the memory area offset mapping table; the RDMA communication front end is used to save the virtual base address, remote key, and the send queue and completion queue for RDMA communication with the memory area.
4. The remote memory system based on intelligent network card offloading according to claim 3, characterized in that The initialization phase includes: The RDMA communication front end initiates an RDMA initialization connection request to the RDMA communication back end, completes the connection of the queue pair of RDMA, creates a completion queue corresponding to the queue pair, and creates a Buffer for receiving data transmitted by the memory node through the queue pair; The memory node registers a memory area according to the size of the RDMA initialization connection request; The memory node sends the virtual base address and remote key of the memory area after registration to the Buffer on the smart network card through the queue pair; The smart network card saves the virtual base address and remote key in the RDMA communication front end.
5. The remote memory system based on intelligent network card offloading according to claim 4, characterized in that, The application registration phase includes: The application on the computing node starts, initiates an RDMA communication connection to the smart network card, and then initiates a registration request to inform the smart network card of the required size of the remote memory; The smart network card receives the registration request, allocates and records the starting position of the remote memory area available for the application, and creates a mapping table of object ID to memory area offset for the application in the data address management module, and then any data using the remote memory space will be recorded in the memory area offset mapping table; After the registration is completed, a notice is returned to inform the application on the computing node that it can continue to run.
6. The remote memory system based on intelligent network card offloading according to claim 5, characterized in that The running phase includes: Step 1, the thread running the application in the computing node initiates a bilateral RDMA remote memory read / write request; Step 2, after receiving the request, the smart network card performs address translation to obtain the virtual address on the memory node; Step 3, the smart network card initiates a unilateral RDMA read / write operation to the memory node; Step 4, the smart network card polls the completion queue to obtain the completion result and returns the read / write result to the computing node through the bilateral RDMA request.
7. The remote memory system based on intelligent network card offloading according to claim 6, wherein The Step 1 further includes: The computing node sends a request and at the same time issues a buffer for receiving the result; the request data includes the operation type, object ID, and data size. If the operation type is a write request, it also includes the data field to be written.
8. The remote memory system based on intelligent network card offloading according to claim 7, characterized in that Step 2 further includes: Step 2.1: Place the requests sent by the computing node into the request buffer of the remote memory request processing module; the remote memory request processing module polls the request notification queue to receive requests from the computing node, and then notifies the threads in the thread pool to process the requests in the buffer; Step 2.2: The remote memory request processing module parses the request data, and looks up the memory area offset corresponding to the object ID through the data address management module for the object ID field; Step 2.3: The remote memory request processing module delivers the memory area offset and the request data to the RDMA communication front end; Step 2.4: The RDMA communication front end obtains the virtual address where the object is located according to the virtual base address and the memory area offset.
9. The remote memory system based on intelligent network card offloading according to claim 8, wherein Step 3 further includes: Step 3.1: The RDMA communication front end initiates a unilateral RDMA request through the virtual address and fills it into the send queue; Step 3.2: The RDMA hardware device sequentially reads and writes the positions corresponding to the virtual addresses on the memory node according to the order in the send queue; after completion, fills the completion results into the completion queue according to the send order. Step 3.2: The RDMA hardware device sequentially reads and writes the positions corresponding to the virtual addresses on the memory node according to the order in the send queue; after completion, fills the completion results into the completion queue according to the send order.
10. The remote memory system based on intelligent network card offloading as claimed in claim 9, wherein, Step 4 further includes: The RDMA communication front end polls the completion queue, and the result after each request is completed is returned to the computing node through the RDMA bilateral send semantics; for write requests, the smart network card needs to add or change the memory area offset mapping table of the data address management module after completion; for read requests, the smart network card needs to return the content of the read data to the computing node together.
Citation Information
Patent Citations
RDMA (Remote Direct Memory Access)-based distributed memory file system
CN108268208A
Intelligent network card, network storage method of intelligent network card and medium
CN114285676A
Data access system, method, equipment and network card
CN115270033A
Data storage system, intelligent network card and computing node
CN116166179A
Memory access control system for RDMA network card
CN116932430A