Memory resource management system, method and apparatus, and device and storage medium

Through the combined system of computing node modules, high-speed interconnection switching modules, memory resource modules and network switching modules, the problem of low efficiency in memory resource management and maintenance in the existing technology is solved, flexible allocation of memory resources and rapid location of faulty hardware are achieved, the utilization rate of server hardware resources is improved and operation and maintenance costs are reduced.

WO2025200547A1PCT designated stage Publication Date: 2025-10-02INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/136495
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-29
Filing Date
2024-12-03
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing technologies are unable to timely summarize each hardware failure in the computing resource pool and memory pool, and are unable to quickly locate faulty hardware, resulting in inefficient memory resource management and maintenance.

Method used

A system combining computing node modules, high-speed interconnection switching modules, memory resource modules, and network switching modules is used. The baseboard management controller and memory resource management processor are used to aggregate and locate fault information. The memory expansion controller is used to monitor memory status. The network switching module is used to achieve network interconnection between modules, enabling flexible allocation of memory resources and rapid location of faulty hardware.

Benefits of technology

It achieves effective management and maintenance of memory resources, improves the utilization of server hardware resources, reduces operation and maintenance costs, and does not affect the normal operation of computing tasks during fault repair.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024136495_02102025_PF_FP_ABST
    Figure CN2024136495_02102025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the technical field of storage, and in particular to a memory resource management system, method and apparatus, and a device and a storage medium, which aim to effectively manage and maintain memory resources. The system comprises computing node modules, a compute express link switch module, a memory resource module and a network switch module, wherein each computing node module comprises a first baseboard management controller and a central processing unit; the compute express link switch module comprises a compute express link switch component, a second baseboard management controller and a memory resource management central processing unit, and is used for managing memory resources, receiving fault information and performing fault localization; the memory resource module comprises a third baseboard management controller, memory expander controllers and memories, the third baseboard management controller being used for controlling the memory expander controllers, and the memory expander controllers being used for monitoring and managing the memories; and the network switch module is used for realizing network interconnection among the modules.
Need to check novelty before this filing date? Find Prior Art

Description

Memory resource management system, method, device, equipment and storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on March 29, 2024, with application number 202410372719.1, and entitled “A Memory Resource Management System, Method, Apparatus, Equipment and Storage Medium,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] Embodiments of the present application relate to the field of storage technology, and more specifically, to a memory resource management system, method, apparatus, device, and storage medium. Background Art

[0004] With the continuous development of computer technology, the demand for memory resources is also increasing, giving rise to memory resource pooling technology. Memory resource pooling technology mainly includes computing resource pools and memory pools. It can achieve flexible allocation of large amounts of memory within the memory pool, greatly improving the utilization of server hardware resources. In a memory pooling environment, how to effectively manage and maintain the memory resources in the memory pool to ensure the normal operation of the business under the memory pooling architecture is a key research issue in memory pooling technology.

[0005] In the related art, a sensor is installed in the memory pool to monitor the memory in the memory pool, and maintenance personnel are used to maintain the memory pool.

[0006] The related technology cannot timely summarize the failures of each hardware in the computing resource pool and the memory pool, nor can it quickly locate the failed hardware, and cannot effectively maintain and manage memory resources. Summary of the Invention

[0007] The embodiments of the present application provide a memory resource management system, method, apparatus, device, and storage medium, which are intended to effectively manage and maintain memory resources.

[0008] In a first aspect, the present application provides a memory resource management system, the system comprising:

[0009] Computing node module, high-speed interconnection switching module, memory resource module, network switching module;

[0010] The computing node module includes a first baseboard management controller and a central processing unit, wherein the first baseboard management controller is used to control the central processing unit;

[0011] The high-speed interconnection switching module includes a high-speed interconnection switching component, a second baseboard management controller, and a memory resource management processor. The high-speed interconnection switching component is used to manage memory resources. The memory resource management processor is used to allocate corresponding memory to the computing node module through the high-speed interconnection switching component. The second baseboard management controller is used to receive fault information sent by the memory resource management processor, and receive fault information sent by the computing node module and the memory resource module through the network switching component, and determine the corresponding faulty hardware according to the fault information.

[0012] The memory resource module includes a third baseboard management controller, a memory expansion controller, and a memory. The third baseboard management controller is used to control the memory expansion controller, and the memory expansion controller is used to monitor and manage the memory.

[0013] The network switching module is used to realize network interconnection between computing node modules, high-speed interconnection switching modules, and memory resource modules.

[0014] In some embodiments, the high-speed interconnect switch component is connected to the central processing unit and the memory resource management processor, and the memory resource management processor is connected to the second baseboard management controller;

[0015] The third baseboard management controller is connected to the memory expansion controller, and the memory expansion controller is connected to the high-speed interconnect switch component and the memory;

[0016] The network switching component is connected to the computing node module, the high-speed interconnection switching module, and the memory resource module.

[0017] In some embodiments, the third baseboard management controller obtains memory status information through the memory expansion controller;

[0018] The third baseboard management controller generates first fault information when any value in the status information in the memory exceeds a preset threshold;

[0019] The third baseboard management controller sends the first fault information to the network switching module;

[0020] The network switching module sends the first fault information to the second baseboard management controller.

[0021] In some embodiments, the third baseboard management controller polls and obtains alarm information generated during the operation of the memory;

[0022] When the alarm information is any one of the preset alarm information, generating second fault information;

[0023] The third baseboard management controller sends the second fault information to the network switching module;

[0024] The network switching module sends the second fault information to the second baseboard management controller.

[0025] In some embodiments, the first baseboard management controller generates third fault information when identifying fault alarm information of the memory expansion controller;

[0026] The first baseboard management controller sends the third fault information to the network switching module;

[0027] The network switching module sends the third fault information to the second baseboard management controller.

[0028] In some embodiments, the memory resource management processor identifies a fault in a high-speed interconnect switch component;

[0029] The memory resource management processor sends the identified fourth fault information to the second baseboard management controller;

[0030] The second baseboard management controller summarizes all received fault information.

[0031] In some embodiments, upon receiving the fault information, the second baseboard management controller determines the computing node and memory corresponding to the fault information based on the memory topology interconnection relationship read by the memory resource management processor from the high-speed interconnection switch component;

[0032] The second baseboard management controller records the fault information into the fault log;

[0033] The second baseboard management controller triggers the first baseboard management controller to check the operating status of the central processing unit;

[0034] The second baseboard management controller controls the computing node to shut down when the first baseboard management controller detects that the central processing unit cannot perform computing tasks;

[0035] The second baseboard management controller adds an abnormal memory mark to the memory;

[0036] The second baseboard management controller sends the memory information of the memory to the memory management processor;

[0037] The memory management processor stops the memory configuration task when receiving the memory information;

[0038] The second baseboard management controller sends a resource allocation command to the memory management processor;

[0039] The memory management processor performs memory resource configuration upon receiving a resource allocation command;

[0040] The second baseboard management controller controls the computing node to restart.

[0041] In some embodiments, the second baseboard management controller clears the abnormal memory flag when detecting that the memory repair is successful.

[0042] A second aspect of an embodiment of the present application provides a memory resource management method, the method comprising:

[0043] During the memory operation in the memory pool, the memory status information of each memory in the memory pool is polled and monitored;

[0044] When any value in the memory status information exceeds a preset threshold, generating first fault information;

[0045] The first fault information is sent to the second baseboard management controller.

[0046] In some embodiments, the method further comprises:

[0047] During the operation of the memory in the memory pool, the alarm information issued by each memory in the memory pool is polled and monitored;

[0048] When the alarm information is any one of the preset alarm information, generating second fault information;

[0049] The second fault information is sent to the second baseboard management controller.

[0050] In some embodiments, the method further comprises:

[0051] During the startup of the computing node, determining whether there is first fault alarm information in the memory expansion controller;

[0052] When first fault alarm information exists in the memory expansion controller, third fault information is generated;

[0053] The third fault information is sent to the second baseboard management controller.

[0054] In some embodiments, the method further comprises:

[0055] During startup of the memory resource management processor, identifying second fault alarm information in the high-speed interconnection switch control component;

[0056] generating fourth fault information when the second fault warning information is identified;

[0057] The fourth fault information is sent to the second baseboard management controller.

[0058] In some embodiments, the method further comprises:

[0059] When fault information is received, the computing node and memory corresponding to the fault information are determined based on the memory topology interconnection relationship;

[0060] Check whether the central processing unit corresponding to the computing node is operating normally;

[0061] Shut down the computing node if it detects that the central processor cannot operate normally;

[0062] Add an exception flag to the memory;

[0063] Sending memory information of the memory to the memory resource management processor;

[0064] Allocate new memory to the compute node;

[0065] When the memory allocation of the compute node is complete, restart the compute node.

[0066] In some embodiments, allocating new memory to a compute node includes:

[0067] Filter out multiple memories that are not marked as abnormal from the memory pool;

[0068] Determine any memory in an idle state among the multiple memories to which no abnormal operation mark is added;

[0069] When the memory is in normal operation, the memory is allocated to the corresponding memory of the computing node.

[0070] In some embodiments, the method further comprises:

[0071] When it is detected that the memory with the abnormal operation mark is repaired, the abnormal operation mark corresponding to the memory is deleted.

[0072] A third aspect of an embodiment of the present application provides a memory resource management device, the device comprising:

[0073] A memory status information determination module is used to poll and monitor the memory status information of each memory in the memory pool during the memory operation period in the memory pool;

[0074] A first fault information generating module, configured to generate first fault information when any value in the memory status information exceeds a preset threshold;

[0075] The first fault information sending module is used to send the first fault information to the second baseboard management controller.

[0076] In some embodiments, the apparatus further comprises:

[0077] The memory alarm monitoring module is used to poll and monitor the alarm information issued by each memory in the memory pool during the operation of the memory in the memory pool;

[0078] A second fault information generating module is configured to generate second fault information when the alarm information is any one of the preset alarm information;

[0079] The second fault information sending module is used to send the second fault information to the second baseboard management controller.

[0080] In some embodiments, the apparatus further comprises:

[0081] A first fault alarm information detection module is used to determine whether there is first fault alarm information in the memory expansion controller during the startup process of the computing node;

[0082] A third fault information generating module, configured to generate third fault information when the first fault warning information exists in the memory expansion controller;

[0083] The third fault information sending module is used to send the third fault information to the second baseboard management controller.

[0084] In some embodiments, the apparatus further comprises:

[0085] A second fault alarm information detection module is used to identify second fault alarm information in the high-speed interconnection switching control component during the startup process of the memory resource management processor;

[0086] a fourth fault information generating module, configured to generate fourth fault information when the second fault warning information is identified;

[0087] The fourth fault information sending module is used to send the fourth fault information to the second baseboard management controller.

[0088] In some embodiments, the method further comprises:

[0089] A hardware determination module is used to determine the computing node and memory corresponding to the fault information based on the memory topology interconnection relationship when the fault information is received;

[0090] The operation status detection module is used to detect whether the central processing unit corresponding to the computing node is operating normally;

[0091] A computing node shutdown module is used to shut down the computing node when it is detected that the central processing unit cannot operate normally;

[0092] An operation abnormality mark adding module is used to add an operation abnormality mark to the memory;

[0093] A memory information sending module, used for sending memory information of the memory to the memory resource management processor;

[0094] Memory allocation module, used to allocate new memory to computing nodes;

[0095] When the memory allocation of the compute node is complete, restart the compute node.

[0096] In some embodiments, the memory allocation module includes:

[0097] The memory filtering submodule is used to filter out multiple memories that have not been marked as abnormal from the memory pool;

[0098] A memory determination submodule, configured to determine any memory in an idle state among a plurality of memories not marked with abnormal operation;

[0099] The memory allocation submodule is used to allocate memory to the memory corresponding to the computing node when the memory is in normal operation.

[0100] In some embodiments, the apparatus further comprises:

[0101] The operation abnormality mark deletion module is used to delete the operation abnormality mark corresponding to the memory when it is detected that the memory with the operation abnormality mark is repaired.

[0102] A fourth aspect of an embodiment of the present application provides a non-volatile readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps in the method of the first aspect of the present application are implemented.

[0103] The fifth aspect of the embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method of the first aspect of the present application are implemented.

[0104] The memory resource management system provided by the present application is adopted, and the system includes: a computing node module, a high-speed interconnection switching module, a memory resource module, and a network switching module; the computing node module includes a first baseboard management controller and a central processing unit, and the first baseboard management controller is used to control the central processing unit; the high-speed interconnection switching module includes a high-speed interconnection switching component, a second baseboard management controller, and a memory resource management processor, and the high-speed interconnection switching component is used to manage memory resources, and the memory resource management processor is used to allocate corresponding memory to the computing node module through the high-speed interconnection switching component, and the second baseboard management controller is used to receive fault information sent by the memory resource management processor, and receive fault information sent by the computing node module and the memory resource module through the network switching component, and determine the corresponding faulty hardware according to the fault information; the memory resource module includes a third baseboard management controller, a memory extension controller, and memory, and the third baseboard management controller is used to control the memory extension controller, and the memory extension controller is used to monitor and manage memory; the network switching module is used to realize network interconnection between the computing node module, the high-speed interconnection switching module, and the memory resource module.

[0105] In this system, a complete memory resource management system is composed of a computing node module, a high-speed interconnection switching module, a memory resource module and a network switching module. The central processing unit on the computing node module uses the memory on the memory resource module when executing computing tasks, and monitors and manages the memory connected to the memory resource module through the memory extension controller on the memory resource module. The memory resource management processor can allocate corresponding memory to each computing node through the high-speed interconnection switching component, realizing flexible allocation of memory resources. The second baseboard management controller receives fault information sent by the remaining modules through the network, summarizes the faults, and can locate the faulty hardware, which facilitates the maintenance of the memory resource pool, thereby realizing effective management and effective maintenance of the memory resource system. BRIEF DESCRIPTION OF THE DRAWINGS

[0106] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0107] FIG1 is a schematic diagram of the structure of a memory resource management system proposed in one embodiment of the present application;

[0108] FIG2 is a schematic diagram of a network interconnection of a memory resource management system proposed in one embodiment of the present application;

[0109] FIG3 is a flowchart of a memory resource management method proposed in one embodiment of the present application;

[0110] FIG4 is a schematic diagram of a memory resource management device proposed in an embodiment of the present application;

[0111] FIG5 is a schematic diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0112] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0113] Refer to Figure 1, which is a schematic diagram of the memory resource management system proposed in one embodiment of the present application. As shown in Figure 1, the system includes:

[0114] Computing node module, high-speed interconnection switching module, memory resource module, and network switching module.

[0115] In this embodiment, the above modules are modules composed of components and circuits integrated on a circuit board.

[0116] In this embodiment, the compute node module is the component that performs computing tasks in the memory resource management system. The high-speed interconnect (CXL, Compute Express Link) switch module is used to manage and configure memory resources. The memory resource module monitors and manages memory and transfers the high-speed interconnect bus of the memory to the high-speed interconnect switch core for computing node connection. The network switch module connects to each board and enables interconnection between the boards through the network.

[0117] The computing node module includes a first baseboard management controller and a central processing unit. The first baseboard management controller is used to control the central processing unit, and the central processing unit is used to execute computing tasks.

[0118] In this embodiment, as shown in Figure 1, the computing node module (CPU Board) includes a first baseboard management controller (CPU Board BMC) and a central processing unit (CPU). The CPU is connected to the first baseboard management controller via an LPC (Low Pin Count Bus). The CPU is connected to the outside world via a PCIE (High-Speed ​​Serial Bus) and is connected to a high-speed interconnect switch control component. When executing computing tasks, the CPU uses memory resources in the memory pool.

[0119] The high-speed interconnection switching module includes a high-speed interconnection switching component, a second baseboard management controller, and a memory resource management processor. The high-speed interconnection switching component is used to manage memory resources. The memory resource management processor is used to allocate corresponding memory to the computing node module through the high-speed interconnection switching component. The second baseboard management controller is used to receive fault information sent by the memory resource management processor, and receive fault information sent by the computing node module and the memory resource module through the network switching component, and determine the corresponding faulty hardware according to the fault information.

[0120] In this embodiment, as shown in FIG1 , the high-speed interconnection switch module includes a high-speed interconnection switch component (CXL SW Board), a second baseboard management controller (CXL SW BMC), and a memory resource management processor (mCPU, Management CPU).

[0121] The high-speed interconnection switching component is a component based on the high-speed interconnection protocol, which is used to manage the memory in the memory pool. One end of the high-speed interconnection switching component is connected to the central processing unit through the PCIE interface, and the other end is connected to the MXC (Memory Expander Controller) through the corresponding interface. The high-speed interconnection switching component and the memory resource management processor are connected through the PCIE bus and also through the UART (Universal Asynchronous Receiver Transmitter) / I2C (Inter-Integrated Circuit) interface.

[0122] The second baseboard management controller receives the second fault information sent by the first baseboard management controller on the compute node module through the network switch module, receives the first fault information sent by the third baseboard management controller, and receives the third fault information sent by the memory resource management processor. The controller aggregates the fault information and locates the fault based on the pre-stored memory topology interconnection relationship to locate the faulty hardware. As shown in Figure 1, the second baseboard management controller is connected to the memory resource management processor via the LPC interface and the SGMII (Serial GMII) interface.

[0123] The memory resource management processor controls the allocation of memory resources through high-speed interconnection switching components, and allocates corresponding memory resources to multiple computing node modules. The memory in the memory resource module can be arbitrarily allocated to the corresponding computing node module, realizing flexible allocation of memory resources.

[0124] The memory resource module includes a third baseboard management controller, a memory expansion controller, and memory. The third baseboard management controller is used to control the memory expansion controller, and the memory expansion controller is used to monitor and manage the memory. The third baseboard management controller is connected to the memory expansion controller, and the memory expansion controller is connected to the high-speed interconnection switching component and the memory.

[0125] In this embodiment, as shown in Figure 1, the third baseboard management controller (DIMM BMC Board) is connected to the memory expansion controller. The third baseboard management controller is used to monitor and manage the memory. When an abnormality in the memory operating status is detected or an alarm information is issued during the operation of the memory, the third baseboard management controller generates corresponding fault information and sends it to the second baseboard management controller.

[0126] The memory expansion controller is used to control and manage the memory. The memory management controller can obtain hardware sensor information such as temperature information, voltage information, power consumption information, etc. during memory operation, and then transmit this information to the third baseboard management controller through the SMBus (microcontroller communication link management) protocol. The memory resource module includes multiple memory expansion controllers, each memory expansion controller is connected to multiple memories, and the memory expansion controller and the memory interact through the SMBus (microcontroller communication link management) protocol.

[0127] The network switching module is used to realize the network interconnection between the computing node module, the high-speed interconnection switching module, and the memory resource module. The network switching component is connected to the computing node module, the high-speed interconnection switching module, and the memory resource module.

[0128] In this embodiment, as shown in Figure 1, the network switching module is connected to the high-speed interconnection switching module, the computing node module, and the memory resource module. The network switching module can realize the interconnection between these three modules.

[0129] Refer to Figure 2, which is a schematic diagram of a network interconnection of a memory resource management system proposed in an embodiment of the present application. As shown in Figure 2, the first baseboard management controller, the second baseboard management controller, the third baseboard management controller and the network switch module

[0130] In this embodiment, the hardware topology connection between the high-speed interconnection switching component and the upstream and downstream is fixed. The upstream is connected to the port of the computing node module, and the downstream is connected to the port of the memory resource module. The memory resource management processor configures the upstream and downstream interconnection relationship of the high-speed interconnection switching component and the memory resource slice management. The second baseboard management controller interacts with the memory resource management processor to obtain which downstream ports correspond to the memory used by the upstream computing node of the current high-speed interconnection switching component, and form a connection topology between the computing resources and the memory resources.

[0131] The third baseboard management controller obtains memory status information through the memory expansion controller.

[0132] In this embodiment, the memory status information is the values ​​of various indicators of the memory operation obtained by pre-deployed sensors during the memory operation process.

[0133] In this embodiment, the third baseboard management controller interacts with the memory expansion controller via the SMBus (Microcontroller Management Bus) protocol, and obtains memory information of the memory during operation from the memory expansion controller.

[0134] The third baseboard management controller generates first fault information when any value in the status information in the memory exceeds a preset threshold.

[0135] In this embodiment, the state information of the memory includes multiple values, each value has a preset threshold, and when any value in the state information of the memory exceeds the preset threshold, the first fault information is generated.

[0136] For example, the temperature information in the memory status information shows that the current memory temperature is 90 degrees Celsius, and the preset temperature threshold is 80 degrees Celsius, then the first fault information is generated at this time.

[0137] The third baseboard management controller sends the first fault information to the network switching module.

[0138] In this embodiment, after generating the first fault information, the third baseboard management controller sends the first fault information to the network switching module.

[0139] The network switching module sends the first fault information to the second baseboard management controller.

[0140] In this embodiment, after receiving the first fault information, the network switching module sends the first fault information to the second baseboard management controller.

[0141] The third baseboard management controller polls and obtains alarm information generated by the memory during operation.

[0142] In this embodiment, the third baseboard management controller interacts with the memory expansion controller through MCTP over SMBus (a computer management transport protocol) protocol, and polls to obtain alarm information (Mailbox Event Record) during the operation of the memory.

[0143] When the alarm information is any one of the preset alarm information, the second fault information is generated.

[0144] In this embodiment, the preset alarm information is relatively important alarm information that is preset.

[0145] In this embodiment, when the alarm information is any one of the preset alarm information, the second fault information is generated.

[0146] For example, the preset alarm information can be General Media Event Record (general media event record), DRAM Event Record (dynamic random memory event record), Memory Module Event Record (memory module event record), Physical Switch Event Record (physical switch event record), Virtual Switch Event Record (network switch event record), MLD Port Event Record (network protocol event record) and Dynamic Capacity Event Record (memory capacity event record), etc.

[0147] The third baseboard management controller sends the second fault information to the network switching module.

[0148] In this embodiment, the third baseboard management controller sends the second fault information to the network switching module.

[0149] The network switching module sends the second fault information to the second baseboard management controller.

[0150] In this embodiment, after receiving the second fault information, the network switching module sends the second fault information to the second baseboard management controller.

[0151] The first baseboard management controller generates third fault information when identifying that the memory expansion controller has fault alarm information.

[0152] In this embodiment, in the computing node module, the BIOS (Basic Input Output System) sends an IPMI (Intelligent Platform Management Interface) command to the first baseboard management controller through the LPC bus for interaction.

[0153] In this embodiment, during the process of starting up the computing node in the computing node module, the BIOS identifies whether the fault alarm information (PCIE alarm information) of the memory expansion controller exists. When the fault alarm information exists in the memory expansion controller, the computer node cannot use the corresponding memory resources normally, and thus cannot execute the computing task. At this time, the BIOS sends an IPMI command to the first baseboard management controller. After identifying the fault alarm information of the corresponding memory expansion controller, the first baseboard management controller generates third fault information.

[0154] The first baseboard management controller sends the third fault information to the network switching module.

[0155] In this embodiment, the first baseboard management controller sends the third fault information to the network switching module.

[0156] The network switching module sends the third fault information to the second baseboard management controller.

[0157] In this embodiment, after receiving the third fault information, the network switching component sends the third fault information to the second baseboard management controller.

[0158] The memory resource management processor identifies faults in high-speed interconnect switching components.

[0159] In this embodiment, in the high-speed interconnection switch module, the BIOS sends an IPMI command to the second baseboard management controller through the LCP bus for interaction.

[0160] In this embodiment, during the boot process of the memory resource management processor, the BIOS of the memory resource management processor identifies the PCIE alarm information sent by the high-speed interconnection switch component.

[0161] The memory resource management processor sends the identified fourth fault information to the second baseboard management controller.

[0162] In this embodiment, the memory resource management processor sends the identified fourth fault information to the second baseboard management controller in the form of an IPMI command.

[0163] The second baseboard management controller summarizes all received fault information.

[0164] In this embodiment, the second baseboard management controller summarizes all received fault information, including the first fault information, the second fault information, the third fault information, and the fourth fault information.

[0165] In this embodiment, the second baseboard management controller serves as the entire machine management CMC (Chassis Management Controller) to perform unified fault aggregation. The first baseboard management controller and the third baseboard management controller establish a connection with the second baseboard management controller through a network hardware link and a command interface to report fault information of hardware such as memory, memory expansion controller, and interface.

[0166] When the second baseboard management controller receives the fault information, it determines the computing node and memory corresponding to the fault information according to the memory topology interconnection relationship read by the memory resource management processor from the high-speed interconnection switch component.

[0167] In this embodiment, upon receiving the fault information, the second baseboard management controller determines the computing node and memory corresponding to the fault information according to the memory topology interconnection relationship read by the memory resource management processor from the high-speed interconnection switch component.

[0168] In this embodiment, the memory resource management processor will read the computing node or memory expansion controller connected to each interface from the high-speed interconnection component, and then obtain the topological interconnection relationship of the entire system. The second baseboard management controller can obtain the topological interconnection relationship of the entire system from the memory resource management processor, and then determine the computing node and memory corresponding to the fault information based on the topological interconnection relationship. As long as the number of one hardware is known, the number of the corresponding other hardware can be known, and the memory module group configured with memory slices can also be located. A memory module consists of multiple memories, and a computing node can also correspond to a memory module.

[0169] For example, the fault information is sent by DIMM0 (memory 0). From the topological interconnection structure, it can be seen that the computing node corresponding to DIMM0 is CPU0. Therefore, it is determined that the memory corresponding to the fault information is DIMM0 and the computing node is CPU0.

[0170] The second baseboard management controller records the fault information in the fault log.

[0171] In this embodiment, after determining the computing node and memory corresponding to the fault information, the second baseboard management controller records the fault information in the fault log.

[0172] The second baseboard management controller triggers the first baseboard management controller to check the operating status of the central processing unit.

[0173] In this embodiment, the second baseboard management controller sends a command to the first baseboard management controller through the network switching module, triggering the first baseboard management controller to check the operating status of the central processing unit.

[0174] When the first baseboard management controller detects that the central processing unit is unable to perform computing tasks, the second baseboard management controller controls the computing node to shut down.

[0175] In this embodiment, when the first baseboard management controller detects that the central processing unit is unable to execute a computing task, the second baseboard management controller controls the computing node where the central processing unit is located to shut down through the network.

[0176] The second baseboard management controller adds an abnormal memory mark to the memory.

[0177] In this embodiment, the second baseboard management controller adds an abnormal memory mark to the faulty memory, marking the memory or memory module that triggered the abnormality and was not repaired in all memory resources.

[0178] The second baseboard management controller sends the memory information of the memory to the memory management processor.

[0179] In this embodiment, the second baseboard management controller sends memory information corresponding to the memory marked as abnormal to the memory management processor, and the memory management processor will no longer use this part of memory resources in subsequent memory allocation or slice configuration tasks.

[0180] When the memory management processor receives the memory information, it stops the memory configuration task.

[0181] In this embodiment, when the memory management processor receives the memory information of the faulty memory, it will no longer use the memory resource in subsequent memory allocation or memory switching configuration tasks and stop the corresponding memory configuration task.

[0182] The second baseboard management controller sends a resource allocation command to the memory management processor.

[0183] In this embodiment, the second baseboard management controller sends a resource allocation command, which is used to instruct the memory management processor to reallocate memory to the corresponding computing node.

[0184] The memory management processor performs memory resource configuration upon receiving the resource allocation command.

[0185] In this embodiment, upon receiving the resource allocation command, the memory management processor performs memory resource configuration and reallocates memory that can be normally used for the computing node.

[0186] The second baseboard management controller controls the computing node to restart.

[0187] In this embodiment, the second baseboard management controller restarts the computing node after the computing node is allocated new memory that can be used normally.

[0188] The second baseboard management controller clears the abnormal memory mark when detecting that the memory repair is successful.

[0189] In this embodiment, when the second baseboard management controller detects that the memory repair is successful, it means that the memory has been repaired and can be used normally. At this time, the abnormal memory mark added to the memory is removed, and the memory resource management processor can use this part of the memory resources during subsequent configuration.

[0190] In this embodiment, the memory resource management system achieves the purpose of flexibly allocating memory in the memory resource pool through a high-speed interconnection switching module, and collects and summarizes the fault information of each hardware through the second baseboard management controller, flexibly locates the faulty hardware, promptly notifies maintenance personnel to repair it, and allocates the memory resources of the all-in-one machine to different computing resource service nodes, which greatly improves the utilization rate of server hardware resources and reduces operation and maintenance costs. The memory can be suspended during the repair period and resumed after the repair is completed, which will not affect the operation of the computing task. Allocating the memory resources of the all-in-one machine to different computing resource service nodes greatly improves the utilization rate of server hardware resources and reduces operation and maintenance costs.

[0191] Referring to FIG3 , FIG3 is a flow chart of a memory resource management method proposed in an embodiment of the present application. The method is applied to a memory resource management system, and the specific steps are as follows:

[0192] S11: During the operation of the memory in the memory pool, the memory status information of each memory in the memory pool is polled and monitored.

[0193] In this embodiment, the memory pool is a memory cluster composed of multiple memories in the memory resource module.

[0194] In this embodiment, the memory operation device in the memory pool polls and monitors the memory status information of each memory in the memory pool.

[0195] For example, there are 10 memories in total, namely memory 0 to memory 9, and the third baseboard management controller obtains the memory status information of each memory from memory 0 to memory 9 through the corresponding memory expansion controller.

[0196] S12: When any value in the memory status information exceeds a preset threshold, first fault information is generated.

[0197] In this embodiment, when any value in the memory status information exceeds a preset threshold, first fault information is generated.

[0198] S13: Send the first fault information to the second baseboard management controller.

[0199] In this embodiment, the third baseboard management controller sends the first fault information to the second baseboard management controller.

[0200] In this embodiment, the method further includes:

[0201] S14: During the operation of the memory in the memory pool, the alarm information issued by each memory in the memory pool is polled and monitored.

[0202] In this embodiment, during the operation of the memories in the memory pool, the memories may issue various alarm messages, and the third baseboard management controller polls and monitors the alarm messages issued by each memory in the memory pool.

[0203] S15: When the alarm information is any one of the preset alarm information, generate second fault information.

[0204] In this embodiment, the preset alarm information is a relatively serious alarm information that is preset and affects the memory operation.

[0205] In this embodiment, when the alarm information issued by the memory is any one of the preset alarm information, it indicates that a problem affecting normal operation of the memory occurs, and the second fault information is generated.

[0206] S16: Send the second fault information to the second baseboard management controller.

[0207] In this embodiment, after the second fault information is generated, the second fault information is sent to the second baseboard management controller.

[0208] In this embodiment, the method further includes:

[0209] S17: During the startup of the computing node, determine whether there is first fault alarm information in the memory expansion controller.

[0210] In this embodiment, the first fault alarm information is PCIE alarm information sent by the memory expansion controller corresponding to the computing node.

[0211] In this embodiment, during the startup of the computing node, the first baseboard management controller determines whether the memory expansion controller has the first fault alarm information.

[0212] S18: When the first fault alarm information exists in the memory expansion controller, generate third fault information.

[0213] In this embodiment, when the first fault alarm information exists in the memory expansion controller, the third fault information is generated.

[0214] S19: Send the third fault information to the second baseboard management controller.

[0215] In this embodiment, the first baseboard management controller sends the third fault information to the second baseboard management controller through the network interconnection module.

[0216] In this embodiment, the method further includes:

[0217] S110: During the startup process of the memory resource management processor, second fault alarm information in the high-speed interconnection switch control component is identified.

[0218] In this embodiment, the second fault alarm information is PCIE alarm information issued by the high-speed interconnection switch control component.

[0219] In this embodiment, during the startup process of the memory resource management processor, second fault alarm information in the high-speed interconnection switch control component is identified.

[0220] S111: When the second fault warning information is identified, fourth fault information is generated.

[0221] In this embodiment, when the memory resource management processor identifies the second fault warning information, it generates fourth fault information.

[0222] S112: Send the fourth fault information to the second baseboard management controller.

[0223] In this embodiment, the memory resource management processor sends the fourth fault information to the second baseboard management controller.

[0224] In this embodiment, the method further includes:

[0225] When receiving the fault information, S21 determines the computing node and memory corresponding to the fault information according to the memory topology interconnection relationship.

[0226] In this embodiment, upon receiving the fault information, the second baseboard management controller obtains the memory interconnection topology from the memory resource management processor, and determines the computing node and memory corresponding to the fault information based on the memory interconnection topology.

[0227] In this embodiment, the fault information includes the serial number of the hardware that sends the fault information, and the memory interconnection topology records each computing node corresponding to each memory. Therefore, the computing node and memory corresponding to the fault information can be determined based on the memory interconnection topology.

[0228] S22: Check whether the central processing unit corresponding to the computing node is operating normally.

[0229] In this embodiment, the second baseboard management controller commands the first baseboard management controller through the network to detect whether the central processing unit on the computing node is operating normally.

[0230] S23: When it is detected that the central processing unit cannot operate normally, the computing node is shut down.

[0231] In this embodiment, when it is detected that the central processing unit cannot operate normally, the computing node is shut down.

[0232] S24: Add an abnormal operation mark to the memory.

[0233] In this embodiment, the fault exception flag is a flag used to indicate that an abnormal event has occurred in the memory and the memory cannot operate normally.

[0234] S25: Send the memory information of the memory to the memory resource management processor.

[0235] In this embodiment, after adding an abnormal operation mark to the faulty memory, the memory information of the memory is sent to the memory resource management processor.

[0236] S26: Allocate new memory for the compute node.

[0237] In this embodiment, the memory resource management processor allocates new memory to the computing node so that the computing node can normally execute the computing task.

[0238] S27: When the memory allocation of the computing node is completed, restart the computing node.

[0239] In this embodiment, when new memory is allocated to the computing node, the computing node is restarted.

[0240] In this embodiment, allocating new memory to a computing node includes:

[0241] S27-1: Filter out multiple memories that are not marked as abnormal from the memory pool.

[0242] In this embodiment, when allocating memory to a computing node, first, a plurality of memories that are not marked with abnormal operation are screened out from the memory pool.

[0243] S27-2: Determine any memory in an idle state among the plurality of memories not marked with abnormal operation.

[0244] In this embodiment, after a plurality of memories are determined, any one memory in an idle state is determined from the plurality of memories.

[0245] S27-3: When the memory is in a normal operating state, the memory is allocated as memory corresponding to the computing node.

[0246] In this embodiment, when the memory is in a normal operating state, the memory is allocated as the memory corresponding to the computer node.

[0247] In this embodiment, the method further includes:

[0248] S31: When it is detected that the memory with the abnormal operation flag added is repaired, the abnormal operation flag corresponding to the memory is deleted.

[0249] In this embodiment, the memory resource management controller regularly checks memory marked with an abnormality. When it detects that a memory marked with an abnormality has been repaired, it removes the abnormality flag corresponding to the memory. This memory is then restored to normal operation in the memory pool and can be used by the memory resource management processor when allocating memory.

[0250] In this embodiment, the above method is used in the memory resource management system to flexibly allocate memory resources in the memory pool. In the event of a memory failure, normal memory can be allocated to the computing node, ensuring the normal operation of the computing task, and quickly locating the faulty hardware, thereby improving the operation and maintenance efficiency of the entire system.

[0251] Based on the same inventive concept, an embodiment of the present application provides a memory resource management device. Referring to FIG4 , FIG4 is a schematic diagram of a memory resource management device 400 proposed in an embodiment of the present application. As shown in FIG4 , the device includes:

[0252] The memory status information determining module 401 is used to poll and monitor the memory status information of each memory in the memory pool during the memory operation period.

[0253] A first fault information generating module 402 is configured to generate first fault information when any value in the memory status information exceeds a preset threshold;

[0254] The first fault information sending module 403 is configured to send the first fault information to the second baseboard management controller.

[0255] In some embodiments, the apparatus further comprises:

[0256] The memory alarm monitoring module is used to poll and monitor the alarm information issued by each memory in the memory pool during the operation of the memory in the memory pool;

[0257] A second fault information generating module is configured to generate second fault information when the alarm information is any one of the preset alarm information;

[0258] The second fault information sending module is used to send the second fault information to the second baseboard management controller.

[0259] In some embodiments, the apparatus further comprises:

[0260] A first fault alarm information detection module is used to determine whether there is first fault alarm information in the memory expansion controller during the startup process of the computing node;

[0261] A third fault information generating module, configured to generate third fault information when the first fault warning information exists in the memory expansion controller;

[0262] The third fault information sending module is used to send the third fault information to the second baseboard management controller.

[0263] In some embodiments, the apparatus further comprises:

[0264] A second fault alarm information detection module is used to identify second fault alarm information in the high-speed interconnection switching control component during the startup process of the memory resource management processor;

[0265] a fourth fault information generating module, configured to generate fourth fault information when the second fault warning information is identified;

[0266] The fourth fault information sending module is used to send the fourth fault information to the second baseboard management controller.

[0267] In some embodiments, the method further comprises:

[0268] A hardware determination module is used to determine the computing node and memory corresponding to the fault information based on the memory topology interconnection relationship when the fault information is received;

[0269] The operation status detection module is used to detect whether the central processing unit corresponding to the computing node is operating normally;

[0270] A computing node shutdown module is used to shut down the computing node when it is detected that the central processing unit cannot operate normally;

[0271] An operation abnormality mark adding module is used to add an operation abnormality mark to the memory;

[0272] A memory information sending module, used for sending memory information of the memory to the memory resource management processor;

[0273] Memory allocation module, used to allocate new memory to computing nodes;

[0274] When the memory allocation of the compute node is complete, restart the compute node.

[0275] In some embodiments, the memory allocation module includes:

[0276] The memory filtering submodule is used to filter out multiple memories that have not been marked as abnormal from the memory pool;

[0277] A memory determination submodule, configured to determine any memory in an idle state among a plurality of memories not marked with abnormal operation;

[0278] The memory allocation submodule is used to allocate memory to the memory corresponding to the computing node when the memory is in normal operation.

[0279] Based on the same inventive concept, another embodiment of the present application provides a non-volatile readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the memory resource management method of any of the above embodiments of the present application are implemented.

[0280] Based on the same inventive concept, another embodiment of the present application provides an electronic device. Referring to Figure 5, Figure 5 is a schematic diagram of an electronic device 500 proposed in an embodiment of the present application, including a memory 502, a processor 501, and a computer program stored in the memory and executable on the processor. When executed by the processor, the steps in the memory resource management method of any of the above embodiments of the present application are implemented.

[0281] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0282] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0283] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, devices, or computer program products. Therefore, the embodiments of the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0284] The present application embodiment is described with reference to the flow chart and / or block diagram of the method, terminal device (system), and computer program product according to the embodiment of the present application. It should be understood that each process and / or box in the flow chart and / or block diagram and the combination of the process and / or box in the flow chart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device produce a device for realizing the function specified in one process or multiple processes and / or one box or multiple boxes of the flow chart.

[0285] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0286] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce computer-implemented processing, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0287] Although some embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they are aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the embodiments of the present application and all changes and modifications that fall within the scope of the embodiments of the present application.

[0288] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.

[0289] The above is a detailed introduction to the memory resource management method, device, equipment and storage medium provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core ideas. At the same time, for general technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A memory resource management system, characterized in that: The system comprises: Computing node module, high-speed interconnection switching module, memory resource module, network switching module; The computing node module includes a first baseboard management controller and a central processing unit, wherein the first baseboard management controller is configured to control the central processing unit; The high-speed interconnection switching module includes a high-speed interconnection switching component, a second baseboard management controller, and a memory resource management processor. The high-speed interconnection switching component is configured to manage memory resources. The memory resource management processor is configured to allocate corresponding memory to the computing node module through the high-speed interconnection switching component. The second baseboard management controller is configured to receive fault information sent by the memory resource management processor and receive the fault information sent by the computing node module and the memory resource module through the network switching component, and determine the corresponding faulty hardware according to the fault information. The memory resource module includes a third baseboard management controller, a memory expansion controller, and a memory, wherein the third baseboard management controller is configured to control the memory expansion controller, and the memory expansion controller is configured to monitor and manage the memory; The network switching module is configured to realize network interconnection among the computing node module, the high-speed interconnection switching module, and the memory resource module.

2. The memory resource management system according to claim 1, wherein: The high-speed interconnection switching component is connected to the central processing unit and the memory resource management processor, and the memory resource management processor is connected to the second baseboard management controller; The third baseboard management controller is connected to the memory expansion controller, and the memory expansion controller is connected to the high-speed interconnection switch component and the memory; The network switching component is connected to the computing node module, the high-speed interconnection switching module, and the memory resource module.

3. The memory resource management system according to claim 1, wherein: The third baseboard management controller obtains the state information of the memory through the memory expansion controller; The third baseboard management controller generates first fault information when any value in the status information of the memory exceeds a preset threshold; The third baseboard management controller sends the first fault information to the network switching module; The network switching module sends the first fault information to the second baseboard management controller.

4. The memory resource management system according to claim 1, wherein: The third baseboard management controller polls and obtains alarm information generated by the memory during operation; When the warning information is any one of the preset warning information, generating second fault information; The third baseboard management controller sends the second fault information to the network switching module; The network switching module sends the second fault information to the second baseboard management controller.

5. The memory resource management system according to claim 4, characterized in that: The alarm information at least includes general media event records, dynamic random access memory event records, memory module event records, physical switch event records, network protocol event records and memory capacity event records.

6. The memory resource management system according to claim 1, wherein: The first baseboard management controller generates third fault information when identifying that the memory expansion controller has fault alarm information; The first baseboard management controller sends the third fault information to the network switching module; The network switching module sends the third fault information to the second baseboard management controller.

7. The memory resource management system according to claim 1, wherein: The memory resource management processor performs fault identification on the high-speed interconnection switching component; The memory resource management processor sends the identified fourth fault information to the second baseboard management controller; The second baseboard management controller summarizes all received fault information.

8. The memory resource management system according to claim 1, wherein: Upon receiving the fault information, the second baseboard management controller determines the computing node and the memory corresponding to the fault information according to the memory topology interconnection relationship read by the memory resource management processor from the high-speed interconnection switch component; The second baseboard management controller records the fault information into a fault log; The second baseboard management controller triggers the first baseboard management controller to check the operating status of the central processing unit; The second baseboard management controller controls the computing node to shut down when the first baseboard management controller detects that the central processing unit is unable to perform computing tasks; The second baseboard management controller adds an abnormal memory mark to the memory; The second baseboard management controller sends the memory information of the memory to the memory management processor; The memory management processor stops configuring tasks for the memory when receiving the memory information; The second baseboard management controller sends a resource allocation command to the memory management processor; The memory management processor performs memory resource configuration when receiving the resource allocation command; The second baseboard management controller controls the computing node to restart.

9. The memory resource management system according to claim 8, characterized in that: When the second baseboard management controller receives the fault information, it determines the computing node and the memory corresponding to the fault information according to the memory topology interconnection relationship read by the memory resource management processor from the high-speed interconnection switch component, including: The memory resource management processor reads the computing node or memory expansion controller connected to each interface from the high-speed interconnect component to obtain the memory topology interconnection relationship; The second baseboard management controller obtains the memory topology interconnection relationship from the memory resource management processor, and determines the computing node and the memory corresponding to the fault information according to the memory topology interconnection relationship.

10. The memory resource management system according to claim 8, wherein: The second baseboard management controller clears the abnormal memory flag when detecting that the memory repair is successful.

11. A memory resource management method, characterized in that: The method is applied to the memory resource management system according to any one of claims 1 to 10, comprising: During the operation of the memory in the memory pool, polling and monitoring the memory status information of each memory in the memory pool; When any value in the memory status information exceeds a preset threshold, generating first fault information; The first fault information is sent to the second baseboard management controller.

12. The memory resource management method according to claim 11, characterized in that: The method further comprises: During the operation of the memory in the memory pool, polling and monitoring alarm information issued by each memory in the memory pool; In a case where the warning information is any one of the preset warning information, generating second fault information; The second fault information is sent to the second baseboard management controller.

13. The memory resource management method according to claim 12, wherein: The method further comprises: During the startup of the computing node, determining whether there is first fault alarm information in the memory expansion controller; When the first fault alarm information exists in the memory expansion controller, generating third fault information; The third fault information is sent to the second baseboard management controller.

14. The memory resource management method according to claim 13, wherein: The method further comprises: During startup of the memory resource management processor, identifying second fault alarm information in the high-speed interconnection switch control component; generating fourth fault information when the second fault warning information is identified; The fourth fault information is sent to the second baseboard management controller.

15. The memory resource management method according to claim 14, wherein: The method further comprises: Upon receiving the fault information, determining the computing node and the memory corresponding to the fault information according to the memory topology interconnection relationship; Detecting whether the central processing unit corresponding to the computing node is operating normally; If it is detected that the central processing unit cannot operate normally, shut down the computing node; Adding an abnormal operation mark to the memory; Sending memory information of the memory to the memory resource management processor; Allocating new memory to the computing node; When the memory allocation of the computing node is completed, restart the computing node.

16. The memory resource management method according to claim 15, characterized in that: Allocating new memory to the computing node includes: Filtering out a plurality of memories to which the abnormal operation mark is not added from the memory pool; Determine any one of the memories that is in an idle state among the plurality of memories to which the abnormal operation mark is not added; When the memory is in a normal operating state, the memory is allocated as the memory corresponding to the computing node.

17. The memory resource management method according to claim 15, characterized in that: The method further comprises: When it is detected that the memory with the abnormal operation mark added thereto is repaired, the abnormal operation mark corresponding to the memory is deleted.

18. A memory resource management device, characterized in that: The device is applied to the memory resource management system according to any one of claims 1 to 10, and the device includes: A memory status information determination module is configured to poll and monitor the memory status information of each memory in the memory pool during the operation of the memory in the memory pool; a first fault information generating module, configured to generate first fault information when any value in the memory status information exceeds a preset threshold; The first fault information sending module is configured to send the first fault information to the second baseboard management controller.

19. A computer non-volatile readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 11 to 17 are implemented.

20. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 11 to 17 are implemented.

Citation Information

Patent Citations

  • Resource sharing device, resource management device, and resource management method

    CN115586964A

  • BMC (Baseboard Management Controller)-based memory resource processing equipment, method and device and medium

    CN115686872A

  • Memory resource management system, method, device and equipment and storage medium

    CN117992270A

  • Use of CXL expansion memory for metadata offload

    US20240020195A1