Separated memory system, memory management method and memory management device

By designing a separate structure between computing pool and memory pool in a separate memory system, using CXL protocol and RDMA network for data transmission, the problems of limited scalability and large remote access overhead in large-scale data centers are solved, and high-performance memory management is achieved.

CN120123089APending Publication Date: 2025-06-10GUANGDONG JINGTIE STORAGE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510209731.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing separate memory systems are limited in scalability in large-scale data centers, and have high remote access overhead, which affects system performance.

Method used

A separate memory system is designed. Through the separation structure between the computing pool and the memory pool, the CXL protocol and the RDMA network are used for data transmission, and the global memory space is set so that the computing nodes can access the memory nodes in the memory pool in a local memory manner, reducing remote access overhead.

Benefits of technology

It improves the scalability of the detached memory system, reduces remote access overhead, improves system performance, and is suitable for large-scale data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123089A_ABST
    Figure CN120123089A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of memory management, and aims to provide a separated memory system, a memory management method and a memory management device. The separated memory system comprises a computing pool and a memory pool, the computing pool comprises a computing node, an address translator and a computing manager, the computing node is in communication connection with the address translator and the computing manager, and data transmission is carried out between the computing node and the computing manager through a CXL protocol; the memory pool comprises a memory node and a memory manager in communication connection with the memory node, and data transmission is carried out between the memory manager and the computing manager through an RDMA network; the number of the computing nodes and the number of the memory nodes are both set to be multiple, and the computing manager is used for mapping all the computing nodes in the computing pool and all the memory nodes in the memory pool to a global memory space. The method is high in expandability, and meanwhile, the remote access overhead can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of memory management, and particularly relates to a disaggregated memory system, a memory management method, and a memory management device. Background Art

[0002] A disaggregated memory system separates computing resources and memory resources from a traditional monolithic server and reorganizes them into independent resource pools including a computing pool and a memory pool. Existing disaggregated memory systems, as Figure 1 shown, the computing pool consists of multiple computing nodes, each of which is equipped with a large number of CPUs (central processing units) and a small amount of local DRAM (Dynamic Random Access Memory); the memory pool consists of multiple memory nodes, which are dedicated to providing large-scale memory support and have almost no computing power. In existing disaggregated memory systems, data transmission between the computing pool and the memory pool is usually achieved through an RDMA (Remote Direct Memory Access) network or a CXL (Compute Express Link, a high-speed interconnection protocol) protocol to ensure efficient resource interaction and collaboration.

[0003] However, in the process of using the prior art, the inventors found that the existing disaggregated memory systems have at least the following problems: a. Limited scalability and not suitable for large-scale data centers: As the scale of the data center continues to expand, the number of nodes increases. The RDMA network depends on the physical network. Although it supports microsecond-level latency, with the increase in the number of nodes, the network bandwidth and latency pressure increase significantly, especially in high-concurrency scenarios; while CXL protocol transmission is limited by the physical length of the PCIe (Peripheral Component Interconnect Express, a high-speed serial computer expansion bus standard) bus and signal transmission attenuation problems, making the current CXL solutions mostly used in the single-machine range (such as inside a single server) and unable to achieve remote large-scale memory pool deployment across servers, thus affecting the application of the disaggregated memory system in large-scale data centers.

[0004] b. High overhead of remote access: Although the RDMA network transmission has a microsecond-level latency, it is still much higher than the nanosecond-level latency of local memory access (for example, the read and write latency of DRAM is usually between dozens and hundreds of nanoseconds). This results in the impact on system performance. For data-intensive tasks or applications with strict latency requirements, the overhead of remote access of computing nodes is particularly obvious, leading to low resource utilization efficiency and system performance degradation; although the CXL protocol transmission supports memory semantics with lower latency, its remote memory access is still limited by the PCIe link bandwidth and latency. Summary of the Invention

[0005] The present invention aims to solve the above technical problems at least to a certain extent. The present invention provides a disaggregated memory system, a memory management method, and a memory management device.

[0006] To achieve the above object, the present invention adopts the following technical solutions: In a first aspect, the present invention provides a disaggregated memory system, including a computing pool and a memory pool; the computing pool includes computing nodes, an address converter, and a computing manager. The computing nodes are respectively communicatively connected to the address converter and the computing manager, and data transmission between the computing nodes and the computing manager is performed through the CXL protocol; the memory pool includes memory nodes and a memory manager communicatively connected to the memory nodes, and data transmission between the memory manager and the computing manager is performed through the RDMA network; the number of both the computing nodes and the memory nodes is set to be multiple, and the computing manager is configured to map each computing node in the computing pool and each memory node in the memory pool to a global memory space.

[0007] In a possible design, the computing pool further includes a first message manager. Correspondingly, the memory pool further includes a second message manager. The computing manager is communicatively connected to the computing manager through the first message manager and the second message manager in sequence, and data transmission between the first message manager and the second message manager is performed through the RDMA network, and data transmission between the computing manager and the first message manager is performed through the CXL protocol.

[0008] In a possible design, any computing node includes a CPU, local memory, and a cache. The CPU is communicatively connected to the local memory, the CPU is communicatively connected to the address converter, and the CPU is further communicatively connected to the cache and the first message manager respectively through the computing manager. Data transmission between the CPU and the computing manager, and between the computing manager and the cache is performed through the CXL protocol.

[0009] In a possible design, the computing manager adjusts the maximum storage capacity of the cache of each computing node by using an elastic capacity allocation strategy.

[0010] In a possible design, the computing manager adjusts the maximum storage capacity of the cache of each computing node by using an elastic capacity allocation strategy, including: The computing manager periodically collects the task priority, cache hit rate, and the number of read / write accesses to pages in the cache of each computing node; The computing manager calculates the cache demand weights of each computing node respectively according to the task priority, cache hit rate, and the number of read / write accesses to pages in the cache of each computing node; The computing manager obtains the cache allocation ratio of each computing node respectively according to the cache demand weights of each computing node, and obtains the cache capacity to be allocated for each computing node according to the cache allocation ratio of each computing node; The computing manager adjusts the maximum storage capacity of the cache of each computing node according to the cache capacity to be allocated for each computing node.

[0011] In a possible design, the cache demand weight of the i-th computing node in the distributed memory system is: W_i = α * P_i + β * H_i + γ * R_i; In the formula, P_i is the task priority of the i-th computing node; H_i is the cache hit rate of the i-th computing node; R_i is the number of read / write accesses to pages in the cache of the i-th computing node; α, β, and γ are preset weight coefficients; The cache allocation ratio of the i-th computing node in the distributed memory system is: A_i = W_i / (ΣW_j); In the formula, ΣW_j is the sum of the cache demand weights of all computing nodes in the distributed memory system; The cache capacity to be allocated for the i-th computing node in the distributed memory system is: S_i = A_i * S_total; In the formula, S_total is the total cache capacity of the distributed memory system.

[0012] In a second aspect, the present invention provides a memory management method, which is executed by any one of the computing nodes in the distributed memory system as described in any one of the above; the method includes: Issue a memory access request, and based on the global memory space through the converter, convert the virtual address in the memory access request into a physical address, and then use the physical address as the target address of the memory access request; wherein, the local memories of each computing node in the computing pool and the virtual-to-physical address mapping relationships of each memory node in the memory pool are mapped in the global memory space; Determine whether the target address belongs to the local memory of any of the computing nodes. If so, access the local memory according to the memory access request. If not, proceed to the next step; Send the memory access request to the computing manager, so that the computing manager accesses the memory node corresponding to the target address according to the memory access request.

[0013] In a possible design, each computing node includes a cache, and the cache is communicatively connected to the computing manager. The virtual-to-physical address mapping relationships of the caches of each computing node in the computing pool are also mapped in the global memory space; correspondingly, the computing manager accessing the memory node corresponding to the target address according to the memory access request includes: The computing manager determines whether the target address of the memory access request belongs to the cache of any of the computing nodes. If so, access the cache according to the memory access request. If not, the computing manager accesses the memory node corresponding to the target address according to the memory access request.

[0014] In a possible design, when the memory access request is a write request, the target address of the memory access request belongs to the cache of any of the computing nodes. At this time, the computing manager writes data to the cache according to the memory access request, and then when the data storage capacity of the cache is greater than a preset capacity threshold, batches the data in the cache and writes it to the memory node in the memory pool.

[0015] In a third aspect, the present invention provides a memory management device for implementing the above-mentioned operating system exception diagnosis method; the memory management device includes: A request issuing module, configured to issue a memory access request, and based on the global memory space through the converter, convert the virtual address in the memory access request into a physical address, and then use the physical address as the target address of the memory access request; wherein, the local memories of each computing node in the computing pool and the virtual-to-physical address mapping relationships of each memory node in the memory pool are mapped in the global memory space; A memory access module, communicatively connected to the request issuing module, is configured to determine whether the target address belongs to the local memory of any of the computing nodes. If so, it accesses the local memory according to the memory access request. If not, it sends the memory access request to the computing manager so that the computing manager accesses the memory node corresponding to the target address according to the memory access request.

[0016] The beneficial effects of the present invention are as follows: The disaggregated memory system of the present invention has strong scalability and can reduce the remote access overhead at the same time. Specifically, in the implementation process of the present invention, when any computing node issues a memory access request, the virtual address in the memory access request can be converted into a physical address based on the global memory space through the address converter, and then the physical address is used as the target address of the memory access request. If the target address belongs to the local memory of any computing node, the local memory is directly accessed. Otherwise, the memory node in the memory pool can be accessed through the computing manager. The memory manager is responsible for dynamically allocating memory resources so that the memory resources of the memory pool can be used by each computing node as needed and support concurrent access to the memory pool by multiple computing nodes. In the present invention, data transmission is carried out through the CXL protocol in the computing pool, and a global memory space is set, so that the computing nodes in the computing pool can call the memory nodes in the memory pool in the same way as accessing the local memory, thereby reducing the overhead of the computing nodes in the computing pool accessing the remote memory in the memory pool. At the same time, data transmission between the memory manager and the computing manager is carried out through the RDMA network, so that the present invention can maintain the scalability inside and between systems under the requirement of latency closer to the local memory, and is applicable to large-scale data centers.

[0017] Other beneficial effects of the present invention will be further described in the specific implementation manner. Description of the Drawings

[0018] Figure 1 is a block diagram of a disaggregated memory system in the prior art; Figure 2 is a block diagram of the disaggregated memory system in Embodiment 1; Figure 3 is a flowchart of the memory management method in Embodiment 2; Figure 4 is a block diagram of the memory management device in Embodiment 3. Specific Embodiment

[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the present invention in combination with the accompanying drawings and the descriptions of the embodiments or the prior art. Obviously, the following descriptions of the structures of the accompanying drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts. It should be noted here that the descriptions of these embodiments are used to help understand the present invention, but do not constitute a limitation to the present invention.

[0020] Embodiment 1: This embodiment discloses a split memory system, as Figure 2 shown, including a computing pool and a memory pool; the computing pool includes computing nodes, an address converter, and a computing manager. The computing nodes are respectively communicatively connected to the address converter and the computing manager, and data is transmitted between the computing nodes and the computing manager through the CXL protocol; the memory pool includes memory nodes and a memory manager communicatively connected to the memory nodes, and data is transmitted between the memory manager and the computing manager through the RDMA network; the number of both the computing nodes and the memory nodes is set to be multiple, and the computing manager is used to map each computing node in the computing pool and each memory node in the memory pool to the global memory space.

[0021] This embodiment has strong scalability and can reduce the remote access overhead at the same time. Specifically, during the implementation of this embodiment, when any computing node issues a memory access request, the address converter can convert the virtual address in the memory access request into a physical address based on the global memory space, and then use the physical address as the target address of the memory access request. If the target address belongs to the local memory of the any computing node, the local memory is directly accessed; otherwise, the memory node in the memory pool can be accessed through the computing manager. The memory manager is responsible for dynamically allocating memory resources so that the memory resources of the memory pool can be used by each computing node on demand and support concurrent access to the memory pool by multiple computing nodes. In this embodiment, data is transmitted through the CXL protocol within the computing pool, and by setting the global memory space, the computing nodes in the computing pool can call the memory nodes in the memory pool in the same way as accessing local memory, thereby reducing the overhead of the computing nodes in the computing pool accessing the remote memory (including each memory node in the memory pool). At the same time, data is transmitted between the memory manager and the computing manager through the RDMA network, so that this embodiment can maintain the scalability inside and between systems under the requirement of latency closer to local memory and is applicable to large-scale data centers.

[0022] In this embodiment, the computing pool further includes a first message manager. Correspondingly, the memory pool further includes a second message manager. The computing manager is communicatively connected to the computing manager through the first message manager and the second message manager in sequence. Data is transmitted between the first message manager and the second message manager through an RDMA network, and data is transmitted between the computing manager and the first message manager through a CXL protocol.

[0023] It should be noted that in this embodiment, the first message manager and the second message manager are used to coordinate the communication between the computing nodes in the computing pool and the memory nodes in the memory pool to ensure efficient and reliable data transmission between the computing pool and the memory pool. They can also detect and handle errors in network transmission, thereby improving the fault tolerance of the disaggregated memory system.

[0024] Specifically, during implementation, when any computing node needs to perform a remote access, that is, needs to access a memory node in the memory pool, it can be accessed through the computing manager. At this time, the first message manager can process the memory access requests sent by the computing manager and coming from each computing node, perform target address resolution, and can also call the memory manager in the memory pool through the second message manager, so that the memory manager can perform operations such as data reading or writing on the memory node matching the target address.

[0025] In this embodiment, any computing node includes a CPU, local memory, and a cache. The CPU is communicatively connected to the local memory, the CPU is communicatively connected to the address converter, and the CPU is also communicatively connected to the cache and the first message manager through the computing manager respectively. Data is transmitted between the CPU and the computing manager, and between the computing manager and the cache through a CXL protocol.

[0026] It should be noted that in this embodiment, the computing manager is responsible for coordinating the task scheduling and data communication of multiple computing nodes, and is also used to map the virtual-to-physical address mapping relationships of the local caches and memories of each computing node in the computing pool, as well as the remote memories in the memory pool, to the global memory space. The address converter can convert the virtual address in the memory access request sent by the CPU in each computing node to a physical address based on the global memory space, so that each computing node can use the current physical address as the target address of the memory access request. If the target address belongs to its local memory, it is directly accessed. If the target address belongs to the cache or remote memory, subsequent access is performed through the computing manager.

[0027] As the tasks executed by the computing pool are different, the computing and storage requirements of each computing node in the computing pool will change, and the memory requirements fluctuate greatly. However, the existing memory management mechanism cannot release or reallocate memory in time, resulting in memory fragmentation and unused memory space. If a cache with a fixed storage capacity is used, it is easy to cause problems such as insufficient or excessive cache capacity of the computing node, resulting in low utilization rate of the cache resources in the computing pool. Therefore, in this embodiment, the following further improvements are made: The computing manager adjusts the maximum storage capacity of the cache of each computing node by using an elastic capacity allocation strategy.

[0028] Specifically, in this embodiment, the computing manager adjusts the maximum storage capacity of the cache of each computing node by using an elastic capacity allocation strategy, including: A1. The computing manager periodically collects the task priority (P), cache hit rate (H), and the number of read and write accesses to pages in the cache (R) of each computing node; wherein, the task priority (P) is used to represent the importance level of the task of the computing node, and the value range is from 1 to K (K is the highest priority); the cache hit rate (H) is the ratio of the number of cache hits to the total number of accesses, and the value range is [0, 1]; the number of read and write accesses to pages in the cache (R) is the number of read and write operations of the cache pages per unit time.

[0029] A2. The computing manager calculates the cache demand weight (W) of each computing node according to the task priority (P), cache hit rate (H), and the number of read and write accesses to pages in the cache (R) of each computing node respectively; it should be noted that the computing manager allocates cache capacity to the computing node based on the cache demand weight (W). Among them, the cache demand weight of the i-th computing node in the split memory system is: W_i = α * P_i + β * H_i + γ * R_i; Wherein, \(P_i\) is the task priority of the \(i\)th computing node, which is pre-normalized to \([0, 1]\), and \(P_i=(P'_i - P_{min}) / (P_{max}-P_{min})\), \(P_{min}\) and \(P_{max}\) are respectively the minimum and maximum values of the initial task priorities of all computing nodes, and \(P'_i\) is the initial task priority of the \(i\)th computing node; \(H_i\) is the cache hit rate of the \(i\)th computing node; \(R_i\) is the number of read / write accesses to pages in the cache of the \(i\)th computing node, which is pre-normalized to \([0, 1]\), and \(R_i=(R'_i - R_{min}) / (R_{max}-R_{min})\), \(R_{min}\) and \(R_{max}\) are respectively the minimum and maximum values of the initial number of read / write accesses to pages in the cache of all computing nodes, and \(R'_i\) is the initial number of read / write accesses to pages in the cache of the \(i\)th computing node; \(\alpha\), \(\beta\) and \(\gamma\) are preset weight coefficients, and \(\alpha+\beta+\gamma = 1\), which are used to adjust the contribution degrees of various indicators.

[0030] A3. The computing manager respectively obtains the cache allocation ratios of each computing node according to the cache demand weights of each computing node, and obtains the cache capacity to be allocated for each computing node according to the cache allocation ratios of each computing node; wherein, the cache allocation ratio of the \(i\)th computing node in the split memory system is: \(A_i = W_i / (\sum W_j)\); Wherein, \(\sum W_j\) is the sum of the cache demand weights of all computing nodes in the split memory system; The cache capacity to be allocated for the \(i\)th computing node in the split memory system is: \(S_i = A_i*S_{total}\); Wherein, \(S_{total}\) is the total cache capacity of the split memory system.

[0031] A4. The computing manager adjusts the maximum storage capacity of the cache of each computing node according to the cache capacity to be allocated for each computing node.

[0032] Specifically, in this embodiment, the computing manager adopts an elastic capacity allocation strategy, and periodically dynamically adjusts the maximum storage capacity of the cache of each computing node according to the cache capacity to be allocated for each computing node, that is, every time interval \(T\) (such as 1 second), recalculate the cache demand weights (\(W\)) of each computing node and perform cache capacity allocation, and can also trigger reallocation immediately when the cache hit rate (\(H_i\)) of a certain computing node drops significantly.

[0033] As an example, assume that there are 3 computing nodes in the split memory system, the total cache capacity is 12 MB, the weight coefficients are respectively set as \(\alpha = 0.4\), \(\beta = 0.3\), \(\gamma = 0.3\), and the data collected by the computing manager in a certain period is as follows: Computing Node 1: P_1 = 0.8, H_1 = 0.9, R_1 = 0.7; Computing Node 2: P_2 = 0.5, H_2 = 0.7, R_2 = 0.5; Computing Node 3: P_3 = 0.3, H_3 = 0.6, R_3 = 0.4; Subsequently, the computing manager calculates the cache requirement weights of the 3 computing nodes as follows: W_1 = 0.4 * 0.8 + 0.3 * 0.9 + 0.3 * 0.7 = 0.80; W_2 = 0.4 * 0.5 + 0.3 * 0.7 + 0.3 * 0.5 = 0.56; W_3 = 0.4 * 0.3 + 0.3 * 0.6 + 0.3 * 0.4 = 0.42; Next, the computing manager calculates the cache allocation ratios of the 3 computing nodes as follows: A_1 = 0.80 / (0.80 + 0.56 + 0.42) = 0.45; A_2 = 0.56 / (0.80 + 0.56 + 0.42) = 0.31; A_3 = 0.42 / (0.80 + 0.56 + 0.42) = 0.24; Finally, the computing manager calculates the cache capacities to be allocated for the 3 computing nodes as follows, in order to adjust the maximum storage capacities of the caches of each computing node in the current cycle: S_1 = 0.45 * 12 = 5.40MB; S_2 = 0.31 * 12 = 3.72MB; S_3 = 0.24 * 12 = 2.88MB.

[0034] Based on this, in this embodiment, the computing manager can dynamically analyze the computing and storage requirements of each computing node in the computing pool, so as to automatically adjust and allocate caches with appropriate capacities for each computing node. Thus, without adding hardware, the load processing capacity of the split memory system can be improved, remote access can be reduced, the average latency can be decreased, and at the same time, the resource utilization rate in the computing pool can be increased, and the problems of memory fragmentation and resource waste in the computing pool can be avoided.

[0035] Embodiment 2: This embodiment discloses a memory management method, which is executed by any one of the computing nodes in the split memory system described in Embodiment 1; as Figure 3 shown, the method may but is not limited to include the following steps: S1. Send a memory access request, and based on the global memory space through the converter, convert the virtual address in the memory access request into a physical address, and then use the physical address as the target address of the memory access request; wherein, the local memories of each computing node in the computing pool and the virtual-to-physical address mapping relationships of each memory node in the memory pool are mapped in the global memory space.

[0036] It should be noted that in this embodiment, each computing node in the computing pool can directly access remote memories through the global memory space, similar to accessing its local memory, thereby reducing the overhead of the computing node for remote memory access and improving the overall performance of the disaggregated memory system.

[0037] Specifically, during the implementation process, the computing manager uses the CXL.io protocol (Compute ExpressLink.io, a high-speed interconnection sub-protocol, which is an improved PCIe 5.0 protocol for initialization, link, device discovery and enumeration, and register access, providing a non-uniform load / store interface for I / O devices) to map the local memories, caches of each computing node in the computing pool and the virtual-to-physical address mapping relationships of each memory node in the memory pool to the global memory space. Since the CXL protocol supports memory semantics, in the global memory space, each memory address can be recognized and accessed by each computing node, enabling each computing node to call remote memories in the same way as calling local memories, thereby reducing the overhead of the computing node for remote memory access.

[0038] In this embodiment, any computing node includes a cache, and the cache is communicatively connected to the computing manager. The virtual-to-physical address mapping relationships of the caches of each computing node in the computing pool are also mapped in the global memory space.

[0039] S2. Determine whether the target address belongs to the local memory of any computing node. If so, access the local memory according to the memory access request; if not, proceed to step S3.

[0040] S3. Send the memory access request to the computing manager so that the computing manager accesses the memory node corresponding to the target address according to the memory access request. Specifically, in this embodiment, when the computing manager accesses the memory node corresponding to the target address, it accesses through the first message manager, the second message manager, and the memory manager in sequence.

[0041] In step S3, the computing manager accesses the memory node corresponding to the target address according to the memory access request, including: The computing manager determines whether the target address of the memory access request belongs to the cache of any of the computing nodes. If so, the cache is accessed according to the memory access request. If not, the computing manager accesses the memory node corresponding to the target address according to the memory access request.

[0042] In this embodiment, when the memory access request is a write request, the target address of the memory access request belongs to the cache of any of the computing nodes. At this time, the computing manager writes data to the cache according to the memory access request, and then when the data storage capacity of the cache is greater than a preset capacity threshold, the data in the cache is batch-written to the memory node in the memory pool.

[0043] Specifically, in this embodiment, based on the caches of each computing node, a data caching mechanism is introduced. That is, during the implementation process, when the memory access request is a write request, the computing manager writes data to the cache according to the memory access request to temporarily store the data in the cache, and then when the data storage capacity of the cache is greater than a preset capacity threshold (such as set to 80% of the maximum storage capacity of the cache, which is not limited here), the data in the cache is batch-written to the memory node in the memory pool, thereby reducing the number of single remote transmissions, effectively avoiding congestion problems in remote transmissions, and alleviating the problem of performance fluctuations of the disaggregated memory system caused by frequently performing write operations to the memory node.

[0044] In addition, in this embodiment, when the memory access request is a read request and the target address of the memory access request belongs to any memory node, the computing manager reads data from the any memory node through the memory manager according to the memory access request, and then stores the read data in the cache to quickly respond to future similar requests of any of the computing nodes. In this embodiment, the caches of each computing node serve as a buffer between its local memory and the memory nodes in the memory pool, thereby enabling fine-grained memory management, while avoiding congestion problems during the remote transmission between each computing node and the memory nodes in the memory pool, and effectively reducing the network transmission overhead brought by each small-scale write operation.

[0045] Embodiment 3: This embodiment discloses a memory management device for implementing the memory management method in Embodiment 2; as Figure 4 shown, the memory management device includes: A request sending module, configured to send a memory access request, and convert a virtual address in the memory access request into a physical address based on the global memory space through the converter, and then use the physical address as the target address of the memory access request; wherein, the local memories of each computing node in the computing pool and the virtual-to-physical address mapping relationships of each memory node in the memory pool are mapped in the global memory space; A memory access module, communicatively connected to the request sending module, configured to determine whether the target address belongs to the local memory of any of the computing nodes. If so, access the local memory according to the memory access request; if not, send the memory access request to the computing manager, so that the computing manager accesses the memory node corresponding to the target address according to the memory access request.

[0046] It should be noted that for the working process, working details and technical effects of the memory management device provided in Embodiment 3, reference can be made to Embodiments 1 and 2, which will not be elaborated here.

[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A separate memory system, comprising a computing pool and a memory pool; characterized in that: The computing pool includes computing nodes, an address converter and a computing manager. The computing nodes are respectively connected to the address converter and the computing manager for communication. The computing nodes and the computing manager perform data transmission via the CXL protocol. The memory pool includes memory nodes and a memory manager connected to the memory nodes for communication. The memory manager and the computing manager perform data transmission via the RDMA network. The number of the computing nodes and the number of the memory nodes are both set to be multiple, and the computing manager is used to map each computing node in the computing pool and each memory node in the memory pool to a global memory space.

2. A separate memory system according to claim 1, characterized in that: The computing pool also includes a first message manager, and correspondingly, the memory pool also includes a second message manager. The computing manager is communicated with the computing manager through the first message manager and the second message manager in sequence, and data is transmitted between the first message manager and the second message manager through an RDMA network, and data is transmitted between the computing manager and the first message manager through a CXL protocol.

3. A separate memory system according to claim 2, characterized in that: Any computing node includes a CPU, a local memory and a cache. The CPU is communicatively connected to the local memory, the CPU is communicatively connected to the address translator, and the CPU is also communicatively connected to the cache and the first message manager respectively through the computing manager. Data is transmitted between the CPU and the computing manager, and between the computing manager and the cache through the CXL protocol.

4. The split memory system according to claim 3, characterized in that: The computing manager uses an elastic capacity allocation strategy to adjust the maximum storage capacity of the cache of each computing node.

5. A separate memory system according to claim 4, characterized in that: The computing manager uses an elastic capacity allocation strategy to adjust the maximum storage capacity of the cache of each computing node, including: The computing manager periodically collects the task priority, cache hit rate and cache page read and write access times of each computing node; The computing manager calculates the cache demand weight of each computing node according to the task priority, cache hit rate and the number of page read and write accesses in the cache of each computing node; The computing manager obtains the cache allocation ratio of each computing node according to the cache demand weight of each computing node, and obtains the cache capacity to be allocated of each computing node according to the cache allocation ratio of each computing node; The computing manager adjusts the maximum storage capacity of the cache of each computing node according to the cache capacity to be allocated of each computing node.

6. The split memory system according to claim 5, characterized in that: The cache demand weight of the i-th computing node in the split memory system is: W_i=α*P_i+β*H_i+γ*R_i; Where, P_i is the task priority of the i-th computing node; H_i is the cache hit rate of the i-th computing node; R_i is the number of page read and write accesses in the cache of the i-th computing node; α, β and γ are preset weight coefficients; The cache allocation ratio of the i-th computing node in the split memory system is: A_i=W_i / (ΣW_j); Wherein, ΣW_j is the sum of the cache demand weights of all computing nodes in the split memory system; The cache capacity to be allocated for the i-th computing node in the split memory system is: S_i=A_i*S_total; Wherein, S_total is the total cache capacity of the split memory system.

7. A memory management method, characterized in that: The method is performed by any computing node in the split memory system according to any one of claims 1 to 6; the method comprises: A memory access request is issued, and the virtual address in the memory access request is converted into a physical address through the converter based on the global memory space, and then the physical address is used as the target address of the memory access request; wherein the local memory of each computing node in the computing pool and the virtual-to-real address mapping relationship of each memory node in the memory pool are mapped in the global memory space; Determine whether the target address belongs to the local memory of any computing node, if yes, access the local memory according to the memory access request, if no, proceed to the next step; The memory access request is sent to the computing manager, so that the computing manager accesses a memory node corresponding to the target address according to the memory access request.

8. A memory management method according to claim 7, characterized in that: Any computing node includes a cache, and the cache is in communication with the computing manager, and the global memory space also maps the virtual-real address mapping relationship of the cache of each computing node in the computing pool; correspondingly, the computing manager accesses the memory node corresponding to the target address according to the memory access request, including: The computing manager determines whether the target address of the memory access request belongs to the cache of any of the computing nodes. If so, the cache is accessed according to the memory access request. If not, the computing manager accesses the memory node corresponding to the target address according to the memory access request.

9. A memory management method according to claim 8, characterized in that: When the memory access request is a write request, the target address of the memory access request belongs to the cache of any computing node. At this time, the computing manager writes data to the cache according to the memory access request, and then when the data storage capacity of the cache is greater than the preset capacity threshold, the data in the cache is written in batches to the memory nodes in the memory pool.

10. A memory management device, characterized in that: Used to implement the operating system abnormality diagnosis method as claimed in claim 7; The memory management device comprises: A request issuing module, used for issuing a memory access request, and converting the virtual address in the memory access request into a physical address based on the global memory space through the converter, and then using the physical address as the target address of the memory access request; wherein the local memory of each computing node in the computing pool and the virtual-to-real address mapping relationship of each memory node in the memory pool are mapped in the global memory space; A memory access module is communicatively connected to the request issuing module, and is used to determine whether the target address belongs to the local memory of any of the computing nodes. If so, the local memory is accessed according to the memory access request; if not, the memory access request is sent to the computing manager so that the memory node corresponding to the target address can be accessed through the computing manager according to the memory access request.