A Separate Memory System and Management Method
By introducing DPU composed of DRAM and CPU in a separate memory system, data access and memory management are optimized, and the problems of high latency and insufficient available memory capacity in a separate memory system are solved, achieving more efficient data access and storage expansion.
Patent Information
- Application Number
- CN202510542725.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-28
AI Technical Summary
In a separate memory system, the access delay between the computing server and the memory server is high, which is not conducive to the expansion of the available memory capacity of the computing server.
A cascading dynamic random access memory DRAM and a data processor DPU composed of a central processor CPU is introduced between the computing server and the memory server. DRAM is the local cache of remote memory. It optimizes data access through the cache management module, the remote memory mapping management module and the remote memory service module, and uses the RDMA network for memory aggregation and management.
It significantly reduces read/write latency, expands the available storage capacity of the computing server, reduces average latency by about 30%, and reduces RDMA network transmission through local cache, improving access efficiency.
Smart Images

Figure CN120066994B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of distributed storage of computer systems, and relates to a separated memory system and a management method thereof. Background Art
[0002] Separate memory is a memory resource pooling management technology that enables a host to access not only local memory but also the idle memory (or memory pool) of other hosts. A separate memory system typically includes computing servers and memory servers. In a separate memory system, an application running on a computing server requests additional memory from a memory server, and the computing server and the memory server generally communicate through a high-speed, low-latency Remote Direct Memory Access (RDMA) network, enabling the computing server to directly access the memory of the memory server.
[0003] The separate memory system provides the memory of the memory server to the computing server for use through an RDMA network. Compared with local memory, cross-server memory access in a separate memory system generally has a higher latency. Therefore, when designing and evaluating a separate memory system, it is necessary to comprehensively consider factors related to latency to ensure that the system can provide efficient services in different application scenarios.
[0004] Patent Application Publication No. CN119149210A, titled "Memory Scheduling Method, System, and Product Based on a Separate Memory System", discloses a separate memory system and a management method. The invention determines the target memory device accessed by the current task based on the demand parameters of the current task and the actual operating parameters of the separate memory system. Before actually deploying the memory device, based on the execution demand parameters of the current task, the demand parameters corresponding to different tasks and the actual operating parameters of the separate memory system are used to preliminarily determine the target memory device to be accessed by the current task. To reduce the access latency corresponding to the current task, the access cost of the current target memory device is estimated based on the historical call count and access latency of the target computing accelerator corresponding to the current task, and the scheduling memory device for the current task is determined according to the access cost, so that the access cost of the scheduling memory device accessed by the target computing accelerator corresponding to each task is relatively small, and the access execution efficiency of the target computing accelerator of the current task is improved. However, it uses the Remote Direct Memory Access (RDMA) network as the underlying communication mechanism, resulting in a still relatively high access latency between the computing server and the memory server, and it is not conducive to the expansion of the available memory capacity of the computing server. Summary of the Invention
[0005] The object of the present invention is to overcome the defects of the above-mentioned prior art, and propose a disaggregated memory system and a management method, which are used to solve the problem of high access latency in the existing disaggregated memory system and expand the available memory capacity of the computing server.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A disaggregated memory system includes a host side composed of multiple computing servers and a remote memory side composed of multiple memory servers interconnected through a Remote Direct Memory Access (RDMA) network; a Data Processing Unit (DPU) is loaded between the host side and the remote memory side; the DPU includes a cascaded Dynamic Random Access Memory (DRAM) and a Central Processing Unit (CPU), where the DRAM is used to cache the data stored by the computing server in the remote memory side; the CPU is used to provide a running environment for aggregating the available memory of the remote memory side.
[0008] As an optimization, the CPU includes a cascaded cache management module, a remote memory mapping management module, a remote memory service module, and a remote memory recycling module connected to the input end of the remote memory mapping management module, where:
[0009] The cache management module is used to search for the block storage page address of the read request or write request sent by the computing server, and divert the read request or write request to the DRAM or the remote memory mapping management module according to the search result;
[0010] The remote memory mapping management module is used to maintain the mapping relationship between the block storage page address and the memory page address of the memory server, and send the read / write request to the remote memory service module according to the mapping result;
[0011] The remote memory service module is used to perform read / write operations on the memory space of the memory server;
[0012] The remote memory recycling module is used to periodically release the invalid memory in the memory server according to the mapping relationship of the remote memory mapping management module.
[0013] A management method for a disaggregated memory system includes the following steps:
[0014] (1) The cache management module initializes the DRAM memory space:
[0015] The cache management module initializes the DRAM memory space, including multiple memory pages, as well as the Least Recently Used (LRU) queue and the First In First Out (FIFO) queue;
[0016] (2)The computing server issues read / write requests;
[0017] (3)The cache management module manages the read / write requests:
[0018] The cache management module searches for the block storage addresses of the read requests or write requests sent by the computing server through the LRU queue and the FIFO queue, and distributes the read requests or write requests to the DRAM or the remote memory mapping management module according to the search results;
[0019] (4)The remote memory mapping management module maintains the mapping table and forwards the read / write requests:
[0020] The remote memory mapping management module searches for the mapped memory server addresses in the mapping table through the addresses of the read requests and write requests, writes the memory server address corresponding to the read request into the read request, and at the same time invalidates the memory server address corresponding to the write request, and then sends the write request to the remote memory service module;
[0021] (5)The remote memory service module processes the read / write requests:
[0022] The remote memory service module reads data from the memory server according to the memory server address of the read request and returns it to the remote memory mapping management module; collects the memory occupancy rate of the memory server and writes the data of the write request into the memory server with the lowest memory utilization rate;
[0023] (6)The remote memory recycling module releases the invalid memory pages in the memory server:
[0024] The remote memory recycling module periodically releases the invalid memory in the memory server according to the mapping table of the remote memory mapping management module.
[0025] Compared with the prior art, the present invention has the following advantages:
[0026] A data processor DPU composed of cascaded dynamic random access memories DRAM and central processing units CPU is loaded between the host end and the remote memory end of the present invention. The DRAM therein serves as a local cache for the remote memory, enabling persistent caching of data frequently accessed by the computing server, so that high-frequency access requests can be directly responded to through the DRAM, avoiding excessive RDMA network transmissions, thereby significantly reducing read / write latency. At the same time, the CPU aggregates the available memory of the remote memory end using the RDMA network interface, and a single computing server can be connected to multiple memory servers simultaneously, expanding the available storage capacity of the operating system of the computing server. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 is a schematic diagram of the overall structure of the split memory system of the present invention.
[0028] Figure 2 is a flowchart of the management method of the split memory system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] Refer to Figure 1 , the split memory system of the present invention includes a host end composed of multiple computing servers and a remote memory end composed of multiple memory servers interconnected through a remote direct memory access RDMA network, and a single computing server can be connected to multiple memory servers simultaneously to expand its available storage capacity; a data processor DPU is loaded between the host end and the remote memory end; the DPU includes cascaded dynamic random access memories DRAM and central processing units CPU, wherein the DRAM is used to cache data stored by the computing server at the remote memory end; the CPU is used to provide an operating environment for aggregating the available memory of the remote memory end.
[0031] The CPU includes four parts: a cascaded cache management module, a remote memory mapping management module, a remote memory service module, and a remote memory recycling module connected to the input end of the remote memory mapping management module, wherein:
[0032] The cache management module searches for the block storage address of the read request or write request sent by the computing server in the LRU and FIFO queues of the DRAM, and distributes the read request or write request to the DRAM or the remote memory mapping management module according to the search result.
[0033] The remote memory mapping management module manages the remote memory mapping table to record the mapping relationship between the block storage page address and the page address in the memory server at the remote memory end. The mapping table not only contains the specific information of the allocated pages in the memory server, but also reflects the distribution and occupancy rate of the current memory resources. When the cache management module fails to find the required data in the LRU and FIFO queues of the DRAM, it will quickly locate the memory server containing the target page through the mapping table, and send read requests and write requests to the remote memory service module according to the mapping result, and jointly update the memory status in real time with the remote memory service module.
[0034] The remote memory service module uses a B+ tree to manage the memory occupancy of the memory server, and uses the RDMA network to send data to the memory server and read data from the memory server, so as to realize the read / write operation of the memory server. When selecting a memory server to write data, the present invention uses a pre-placement strategy to keep the memory resource usage status of the memory server balanced.
[0035] The remote memory recycling module is used to periodically release the invalid memory in the memory server according to the mapping relationship of the remote memory mapping management module. Since the DPU adopts a remote update mechanism, over time and with the accumulation of multiple write operations, some data may become obsolete or invalid, and the storage space occupied by these invalid data needs to be recycled. The remote memory recycling module will periodically compare the quotient of the number of valid memory pages of each memory server and the total number of memory pages with a preset recycling threshold according to the mapping table of the remote memory mapping management module. If the quotient is less than the recycling threshold, it will search for the invalid memory pages in the mapping table of the remote memory mapping management module, delete the information of the found invalid memory pages from the mapping table, and release these invalid memory pages in the memory server at the same time.
[0036] Refer to Figure 2 , a management method for a disaggregated memory system, includes the following steps:
[0037] Step 1) The cache management module initializes the DRAM memory space:
[0038] The cache management module divides the DRAM memory space into multiple memory pages with a fixed capacity, and LRU queues and FIFO queues with a configurable capacity ratio, and assigns a unique identification number to each memory page.
[0039] Step 2) The computing server issues read / write requests:
[0040] When the computing server runs out of memory, the operating system will write the pages that are not frequently used in the computing server to the DPU from the memory, or read the data in the DPU to the memory. Therefore, when accessing the remote memory server, read / write requests will be issued.
[0041] Step 3) The cache management module manages read / write requests:
[0042] (3a) The cache management module checks whether the FIFO queue is full. If the FIFO queue is full, it converts the address at the tail of the FIFO queue and the data pointed to by the memory page number into a write request and sends it to the remote memory mapping management module; otherwise, it does nothing.
[0043] (3b) Checks whether the LRU queue is full. If the LRU queue is full, it evicts the address and memory page number at the tail of the LRU queue to the head of the FIFO queue; otherwise, it does nothing.
[0044] (3c) Writes the data of the write request into the memory page in the DRAM memory space, and at the same time inserts the memory page number and the address of the write request into the head of the LRU queue.
[0045] (3d) The cache management module checks whether the read request hits in the DRAM cache. If so, it returns the data to the computing server and inserts the memory page where the data is located and the address of the read request into the head of the LRU queue; otherwise, it reads from the target memory server according to the mapping table managed by the remote memory mapping management module, returns the read data to the computing server and writes it into the memory page in the DRAM memory space, and then inserts the memory page number and the address of the read request into the head of the LRU queue.
[0046] Step 4) The remote memory mapping management module maintains the mapping table and transfers read / write requests:
[0047] (4a) The remote memory mapping management module checks from the mapping table whether the read request sent by the cache management module has been allocated according to the address of the read request. If so, it fills the address into the read request and forwards it to the remote memory service module; otherwise, it returns a read error message to the cache management module.
[0048] (4b) The remote memory mapping management module requests to allocate a memory page from the remote memory service module according to the write request sent by the cache management module.
[0049] Step 5) The remote memory service module processes read / write requests:
[0050] The remote memory service module records the page address information in the remote memory server in a B+ tree structure. When it receives a read / write request sent from the remote memory mapping management module, the processing methods are as follows:
[0051] (5a) After receiving the read request sent from the remote memory mapping management module, the remote memory service module finds the memory server mapped by the address of the read request and reads the data from the target memory server and returns it to the remote memory mapping management module.
[0052] After the remote memory service module receives the write request sent from the remote memory mapping management module in (5b), it queries the memory occupancy rate of each memory server at the remote memory end, and then calculates the standard deviation between the memory occupancy rate of each memory server after placing the data of the write request on it and the average memory occupancy rate of all memory servers in sequence. Then, it selects the memory server with the smallest standard deviation as the target memory server to write the data of the write request, where :
[0053] ;
[0054] Among them, is the number of memory servers, is the memory occupancy rate of the th memory server, , is the average memory occupancy rate of all memory servers;
[0055] In (5c), the remote memory service module uses the RDMA network to send data to the memory server and reads data from the memory server. The specific operation is to initiate a mapping request using the RDMA bilateral send SEND operation. The memory server receives it through the RECEIVE operation and feeds back the execution result through the bilateral operation as well. The read and write operations of the data are implemented through the RDMA unilateral operations READ / WRITE.
[0056] Step 6) The remote memory recovery module releases the invalid memory pages in the memory server:
[0057] In (6a), the remote memory recovery module periodically determines whether the quotient of the number of valid memory pages of each memory server and the total number of memory pages is less than a preset recovery threshold. If so, it searches for the invalid memory pages in the mapping table of the remote memory mapping management module and records the information of these invalid memory pages.
[0058] In (6b), the remote memory recovery module deletes the information of the found invalid memory pages from the mapping table and releases the invalid memory pages in the memory server.
[0059] The cache management module in the CPU of the present invention uses the DRAM as the local cache of the remote memory, so that the data frequently accessed by the computing server is persistently cached in the DRAM of the DPU. The high-frequency access requests can be directly responded through the local cache, avoiding excessive RDMA network transmissions, reducing the access latency of cache hit requests, and thus significantly reducing the overall read / write latency. Experimental data shows that the average latency of the present invention is reduced by about 30% compared with the prior art; at the same time, the available block storage of the computing server also increases as the number of memory servers increases.
[0060] As described above, it is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be covered within the protection scope of the present invention.
Claims
1. A split memory system, comprising a host side composed of multiple computing servers and a remote memory side composed of multiple memory servers interconnected through a Remote Direct Memory Access (RDMA) network; characterized in that, A data processor DPU is loaded between the host side and the remote memory side; the DPU includes a cascaded dynamic random access memory DRAM and a central processing unit CPU, where the DRAM is used to cache the data stored by the computing server in the remote memory side; the CPU is used to provide a running environment for aggregating the available memory of the remote memory side; The CPU includes a cascaded cache management module, a remote memory mapping management module, and a remote memory service module, and a remote memory recycling module connected to the input end of the remote memory mapping management module, where: The cache management module is used to look up the block storage page address of the read request or write request sent by the computing server, and divert the read request or write request to the DRAM or the remote memory mapping management module according to the lookup result; The remote memory mapping management module is used to maintain the mapping relationship between the block storage page address and the memory page address of the memory server, and send the read / write request to the remote memory service module according to the mapping result; The remote memory service module is used to read / write the memory space of the memory server; The remote memory recycling module is used to periodically release the invalid memory in the memory server according to the mapping relationship of the remote memory mapping management module.
2. A management method for a split memory system, characterized in that, It includes the following steps: (1) The cache management module initializes the DRAM memory space: The cache management module initializing the DRAM memory space includes multiple memory pages, as well as an LRU queue and a FIFO queue; (2) The computing server issues a read / write request; (3) The cache management module manages the read / write request: Look up the block storage address of the read request or write request sent by the computing server through the LRU queue and the FIFO queue, and divert the read request or write request to the DRAM or the remote memory mapping management module according to the lookup result; (4) The remote memory mapping management module maintains the mapping table and transfers the read / write request: The remote memory mapping management module respectively looks up the mapped memory server address in the mapping table through the addresses of the read request and the write request, writes the memory server address corresponding to the read request into the read request, and at the same time sets the memory server address corresponding to the write request to invalid, and then sends the write request to the remote memory service module; (5) The remote memory service module processes the read / write request: The remote memory service module reads data from the memory server according to the memory server address of the read request and returns it to the remote memory mapping management module; Collect the memory occupancy rate of the memory server and write the data of the write request into the memory server with the lowest memory utilization rate; (6) The remote memory recycling module releases the invalid memory pages in the memory server: The remote memory recycling module periodically releases the invalid memory in the memory server according to the mapping table of the remote memory mapping management module.
3. The method according to claim 2, characterized in that, In step (1), the cache management module initializes the DRAM memory space, and the implementation steps are: The cache management module divides the DRAM memory space into multiple memory pages with a fixed capacity, as well as an LRU queue and a FIFO queue with a configurable capacity ratio, and assigns a unique identification number to each memory page.
4. The method according to claim 2, wherein The cache management module described in step (3) manages read / write requests, and the implementation steps are as follows: (3a) The cache management module checks the FIFO queue. When the FIFO queue is full, it converts the data pointed to by the address and memory page number at the tail of the queue into a write request and sends it to the remote memory mapping management module. Then it checks the LRU queue. When the LRU queue is full, it evicts the address and memory page number at its tail to the head of the FIFO queue; (3b) The cache management module writes the data of the write request into the memory page in the DRAM memory space, and at the same time inserts the memory page number and the address of the write request into the head of the LRU queue; (3c) The cache management module checks whether the read request hits in the DRAM cache. If so, it returns the data to the computing server and inserts the memory page where the data is located and the address of the read request into the head of the LRU queue; Otherwise, it initiates a read request to the remote memory mapping management module, returns the read data to the computing server and writes it into the memory page in the DRAM memory space, and inserts the memory page number and the address of the read request into the head of the LRU queue.
5. The method according to claim 2, characterized in that The mapping table described in step (4) includes multiple triples. Each triple contains block storage memory page number information, memory server memory page address information, and valid flag bit information. The block storage memory page number and the memory server memory page address are composed of 64 bits, and the valid flag bit is composed of 1 bit.
6. The method according to claim 2, characterized in that The method for invalidating the memory server address corresponding to the write request described in step (4) is as follows: The remote memory mapping management module sets the valid flag bit of the triple corresponding to the block storage memory page number carried in the write request in the mapping table to 0, and at the same time writes the memory server page address information corresponding to the mapping table into the write request.
7. The method according to claim 2, characterized in that, The steps for the remote memory service module to process read / write requests described in step (5) are as follows: (5a) After receiving the read request sent from the remote memory mapping management module, the remote memory service module finds the memory server where the address of the read request is located, reads the data from the target memory server and returns it to the remote memory mapping management module; After the remote memory service module receives the write request sent from the remote memory mapping management module, it queries the memory occupancy rates of each memory server on the remote memory side, and calculates the standard deviation between the memory occupancy rate of each memory server after writing the data of the write request and the average memory occupancy rate of all memory servers in sequence. , and selects the memory server with the smallest standard deviation as the target memory server to write the data of the write request, where: ; Among them, is the number of memory servers, is the memory occupancy rate of the th memory server, and is the average memory occupancy rate of all memory servers; (5c) The remote memory service module interacts with the memory server using the RDMA network to realize the reading out of the data of the read request and the writing of the data of the write request.
8. The method according to claim 2, wherein The steps for the remote memory recycling module to release the invalid memory pages in the memory server described in step (6) are as follows: (6a) The remote memory recycling module periodically judges whether the quotient of the number of valid memory pages of each memory server and the total number of memory pages is less than a preset recycling threshold. If so, it finds the invalid memory pages in the mapping table of the remote memory mapping management module and records the information of these invalid memory pages; (6b) The remote memory recycling module deletes the information of the found invalid memory pages from the mapping table and releases the invalid memory pages in the memory server.
Citation Information
Patent Citations
Memory scheduling method, system and product based on separated memory system
CN119149210A
Remote memory system based on intelligent network card unloading
CN117785789A
Data access method, device and equipment
CN119690323A
Data processing unit for compute nodes and storage nodes
US20190012278A1