High-efficiency memory repair system and method based on pooling memory

By using a high-efficiency memory repair system based on pooled memory and employing repair technology based on memory pages, combined with out-of-band storage modules and hash calculation routing table entries, the system solves the problems of resource waste and insufficient repair granularity in memory pooling. This achieves efficient and flexible memory management and repair, improving the reliability and lifespan of the memory system.

CN120950311AActive Publication Date: 2025-11-14CORE TREND (ZHUHAI) TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511484508.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-11-14
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing memory pooling technologies suffer from severe resource waste, repair solutions rely on memory physical structure and storage materials, and lack sufficient granularity and flexibility in repair, making it difficult to adapt to the dynamic resource management needs of a pooled environment. Furthermore, existing memory repair solutions cannot meet the high reliability and high stability requirements of data centers.

Method used

Design a high-efficiency memory repair system based on pooled memory. The system uses a memory testing module to detect erroneous addresses at the cacheline level, a management module and an address routing module to repair memory pages, and an independent out-of-band storage module to repair damaged memory. A routing table based on hash calculation is established to achieve memory addressing and repair.

Benefits of technology

It improves memory utilization efficiency, enhances the reliability and availability of the memory system, reduces repair costs, achieves high compatibility with the operating system and flexible memory management, and meets the requirements of data centers for memory RAS systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120950311A_ABST
    Figure CN120950311A_ABST
Patent Text Reader

Abstract

The invention provides a high-efficiency memory repair system and method based on a pooling memory. The method comprises the following steps: a memory test module performs UCE error address detection on a memory bank by taking cache as granularity; the management module is used for converting the memory error address into a memory page address to which the memory error address belongs, and issuing a repair command for the memory page address; the address routing module replaces and adds a memory page address into a routing table by utilizing a routing table look-up algorithm according to the repair command so as to finish address repair, performs corresponding address conversion according to a repair state of an accessed host memory address, and routes the address to a port where the memory bank or the out-of-band storage module is located; and the out-of-band storage module is an out-of-band memory pool independent of the system memory pool and is used for repairing the damaged memory by taking the memory page as the granularity. According to the method, the effects of improving the utilization efficiency and compatibility of the memory and enhancing the RAS system of the memory are achieved through the high-efficiency memory repair system which is provided with the out-of-band memory module and takes the memory page capable of being dynamically configured as the granularity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and specifically to a high-efficiency memory repair system and method based on pooled memory. Background Technology

[0002] The three most critical elements in the core architecture of a data center are computing, storage, and networking. Among these, memory, as a volatile storage unit, directly impacts the overall operational efficiency and service capabilities of the data center through its performance and reliability. With the continuous development of technologies such as AI, the massive amounts of temporary data generated by computation require huge amounts of memory, frequently leading to memory shortages on computing nodes. Simultaneously, some peripheral nodes experience memory surplus and underutilization. From a timeline perspective, due to the diverse and unpredictable nature of the services handled by data centers, memory resource usage across different nodes is highly uneven, exhibiting both short-term bursts of resource shortages and long-term periods of idle resources. Therefore, to fully utilize the expensive memory resources of data centers, memory pooling technology has become a significant trend in data center development.

[0003] Memory pooling technology enables on-demand allocation and dynamic supply of memory at the system level, achieving flexible resource scheduling. However, on the one hand, data centers have high requirements for data security and service stability, and current RAS (Reliability, Availability, and Serviceability) assurance mechanisms, especially memory repair methods, in pooled memory environments are not yet perfect. On the other hand, to ensure high-quality and highly stable services, data centers need to periodically perform preventative replacements of highly sensitive components such as memory, typically every 3-5 years. However, the memory base of data centers is relatively large, leading to a significant waste of resources as a large amount of still repairable memory is discarded, greatly increasing the operating costs of cloud vendors.

[0004] Currently, the main memory repair solutions on the market include: Post Package Repair (PPR), Adaptive Double DRAM Device Correction (ADDDC), Memory Rank Sparing (MRS), Memory Page Retire (MPR), and Memory Patrol Scrub (MPS). Among these, PPR, ADDDC, and MRS technologies all replace uncorrectable memory errors (UCE) based on the physical structural units of the memory module (DIMM) (such as ROW, BANK, or RANK). However, the implementation of these solutions requires prior backup of the corresponding physical units. While this can ensure data security to a certain extent, it also results in significant resource waste. Furthermore, their feasibility heavily depends on the DIMM manufacturer's product manufacturing process and is affected by factors such as the number of UCE errors and repair time in practical applications, failing to meet the RAS (Recovery, Safety, and Availability) requirements of memory systems. MPR and MPS technologies are merely mechanisms for alarming and masking memory errors; they do not actually repair the memory.

[0005] In summary, existing reliability management solutions for pooled memory still suffer from the following shortcomings: First, significant resource waste exists. Both redundancy-based repair schemes and preventative device replacement strategies require substantial memory redundancy resources, drastically reducing the effective utilization rate of memory resources. Second, there is technological conservatism, with existing solutions limited by memory physical structure and storage materials. Third, there are architectural limitations; existing repair methods lack granularity and flexibility, making it difficult to adapt to the fine-grained and dynamic resource management needs of pooled environments. Therefore, there is an urgent need to design a memory repair system and method that is applicable to memory pooling, improves memory utilization efficiency and operating system compatibility, enhances the functionality of memory RAS systems, and thereby improves the reliability, availability, and lifespan of memory systems. Summary of the Invention

[0006] To address the common problems in existing technologies, the present invention aims to provide a high-efficiency memory repair system and method based on pooled memory. This invention, based on data center memory pooling, utilizes an out-of-band storage module and a high-efficiency memory repair system with dynamically configurable memory pages to improve memory utilization efficiency and compatibility with the operating system, enhance the functionality of the memory RAS system, and thus improve the reliability, availability, and lifespan of the memory system.

[0007] The present invention achieves the above objectives through the following technical solutions: A high-efficiency memory repair system based on pooled memory includes: a memory testing module, a management module, an address routing module, and an out-of-band storage module. The memory testing module performs error address detection on memory modules at the cacheline level and reports the detected UCE error addresses to the management module.

[0008] The management module is used to convert the UCE error address into its corresponding memory page address, query the repair status of the memory page address, generate a repair command for the memory page address based on its repair status, and send it to the address routing module.

[0009] The address routing module uses a routing lookup algorithm to replace and add the memory page address to the routing table according to the repair command to complete the address repair. It also performs corresponding address translation and routes the memory to the memory module or the port where the out-of-band storage module is located based on the repair status of the accessed host memory address.

[0010] The out-of-band storage module is an out-of-band memory pool independent of the system memory pool, used to repair damaged memory at the memory page level.

[0011] According to the high-efficiency memory repair system based on pooled memory provided by the present invention, the memory testing module is integrated into the DDR controller of the memory system, and its detection range is limited to the memory modules on one or more memory channels managed by the DDR controller. The UCE error address detected by the module is the address of the memory channel corresponding to the memory module with the UCE error.

[0012] According to the high-efficiency memory repair system based on pooled memory provided by the present invention, the memory testing module reports the UCE error information detected by the memory module to the management module. The UCE error information includes: the slot where the error occurred, the memory chip, the memory controller, the channel, the slot, the physical array, the sub-physical array, the logical array group, the logical array and the row address and column address, the device identifier, the memory chip identifier and the transfer number.

[0013] According to the high-efficiency memory repair system based on pooled memory provided by the present invention, the process of determining the repair status of the host memory address by the address routing module includes: Send the memory access address to the host and perform routing table entry calculation.

[0014] Using the calculation results from the routing table entries, access the routing table entries and determine the repair status of the host memory address.

[0015] If the host memory address is a normally accessible memory address, then its repair status is "not repaired".

[0016] If the host memory address is a repaired memory page address, then its repair status is repaired.

[0017] According to the high-efficiency memory repair system based on pooled memory provided by the present invention, the routing table entry is an entry for address translation and address routing merging. Its calculation method adopts hash calculation. The host memory address is calculated by hash function and a routing table index is generated. The corresponding entry in the routing table is accessed according to the index.

[0018] According to the high-efficiency memory repair system based on pooled memory provided by the present invention, the address routing module obtains the host memory address. If it is determined that the address is in an unrepaired state, it converts it into a readable and writable DDR control channel address and routes it to the port where the memory module is located. If it is determined that the address is in a repaired state, it converts it into the address of the out-of-band storage module and routes it to the port where the out-of-band storage module is located.

[0019] According to the high-efficiency memory repair system based on pooled memory provided by the present invention, the management module generates a memory repair log. The memory repair log is used to query the repair status of the memory page address and the RAS log information of the memory repair system. It includes: UCE error address, repaired memory page address, address of the repaired out-of-band storage module, repair result information, and abnormal cause information.

[0020] According to the high-efficiency memory repair system based on pooled memory provided by the present invention, the out-of-band storage module has an address encoding independent of the system memory pool. This address encoding is managed by the management module, and the address size is determined at the memory page granularity during system initialization.

[0021] The address routing module routes memory access commands for damaged and repaired addresses to the out-of-band storage module according to the routing table entries, and the out-of-band storage module performs memory read or write operations according to the memory access commands.

[0022] According to the high-efficiency memory repair system based on pooled memory provided by the present invention, the memory page of the out-of-band storage space of the out-of-band storage module is defined as the memory page. The page size of the memory page can be dynamically configured according to the memory damage situation, and its value is:

[0023] Where n is a non-negative integer.

[0024] The dynamic configuration range of its page size is:

[0025] in, This is the maximum capacity of the out-of-band storage space.

[0026] A high-efficiency memory repair method based on pooled memory, applied to the aforementioned high-efficiency memory repair system based on pooled memory, includes: S1: The memory testing module performs error address detection on the memory modules it manages at the cacheline level and reports the detected UCE error addresses.

[0027] S2: After receiving the UCE error address, the management module converts it into the corresponding memory page address and checks whether the memory page address has been repaired. If it has not been repaired, a repair command for the memory page address is issued.

[0028] S3: After receiving the repair command, the address routing module uses a routing lookup algorithm to replace and add the memory page address to the routing table, and reports the repair result to the management module; if the entry for the memory page address can be added, the repair result is successful, otherwise it is a failure.

[0029] S4: The management module reports the repair results in the form of a memory repair log, and the memory system queries the RAS log information through the memory repair log.

[0030] Therefore, compared with the prior art, the present invention has the following beneficial effects: 1. This invention designs an out-of-band storage module independent of the conventional system memory pool and performs memory repair at the memory page level. Compared with traditional repair solutions, it is no longer limited to the memory redundancy space of the memory module itself for memory repair. Instead, it provides a brand-new repair solution that uses out-of-band storage space as redundant memory. Memory repair is performed based on out-of-band storage space. The redundant storage space is no longer limited by the storage medium. The out-of-band storage medium can be DDR, flash, SSD, etc., or it can be the RAM space on the chip.

[0031] 2. This invention establishes routing and address translation table entries based on memory page addresses at the front end of the DDR controller, thereby enabling the feasibility of post-addressing repair of system memory. Compared with the physical structure of traditional basic DDR memory modules, which no longer distinguishes between basic memory structural units such as ROW, BANK, and RANK, this invention is a post-addressing memory repair solution with strong affinity to the upstream operating system and is no longer limited to the memory module structure itself, nor is it limited by the underlying technical structure of memory manufacturers.

[0032] 3. The memory page size unit in the out-of-band storage module of this invention is flexibly configurable. The smallest unit of memory page size is 64 bytes. Compared with traditional memory repair methods such as row (2K) and bank (GB level), it can save the resources required for memory repair and solve the problem of excessive redundant memory in traditional methods. Moreover, by dynamically adjusting the memory page size, the repairability of damaged addresses is enhanced and the applicable scenarios are wider.

[0033] 4. The memory testing module of this invention is designed to be integrated with the DDR controller and performs error detection and marking based on the cacheline. The error memory detection mechanism of this invention enables the memory system to autonomously complete the detection of memory module error addresses without relying on external expensive and complex automatic testing equipment, which significantly reduces equipment dependence and testing time, and lowers the overall cost.

[0034] 5. This invention enables memory fault diagnosis, predictive maintenance, and memory resource monitoring through memory repair logs, meeting the serviceability requirements of RAS systems, enhancing the memory pooled RAS system, and facilitating unified memory management in data centers.

[0035] 6. This invention realizes the feasibility of repairing memory storage units based on chip storage units by using routing table entries based on hash calculation, and ensures the bandwidth of memory access after repair.

[0036] 7. This invention establishes a complete memory repair RAS system based on memory pages as units, including memory UCE testing and reporting, memory address routing mechanism, memory repair method mechanism, and memory repair log maintenance mechanism. Compared with traditional memory repair schemes, this invention improves the RAS system for system memory management. The formulation of this overall repair strategy can perform system error memory repair more flexibly and efficiently, enhance the functionality of the memory RAS system, and improve the reliability, availability, and lifespan of the memory system.

[0037] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments. Attached Figure Description

[0038] Figure 1 This is a schematic diagram of a system embodiment of a high-efficiency memory repair system based on pooled memory according to the present invention.

[0039] Figure 2 This is a flowchart of a method embodiment of a high-efficiency memory repair method based on pooled memory according to the present invention. Detailed Implementation

[0040] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0041] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0042] See Figure 1 The present invention discloses a high-efficiency memory repair system based on pooled memory, comprising: a memory testing module 10, a management module 20, an address routing module 30, and an out-of-band storage module 40. The memory testing module 10 performs error address detection on memory modules at the cacheline level and reports the detected UCE error addresses to the management module 20.

[0043] The management module 20 is used to convert the UCE error address into its corresponding memory page address, query the repair status of the memory page address, generate a repair command for the memory page address based on its repair status, and send it to the address routing module 30.

[0044] The address routing module 30 uses a routing lookup algorithm to replace and add the memory page address to the routing table according to the repair command to complete the address repair. It also performs corresponding address translation and routes the memory address to the port where the memory module 40 is located based on the repair status of the accessed host memory address.

[0045] The out-of-band storage module 40 is an out-of-band memory pool independent of the system memory pool, used to repair damaged memory at the memory page level.

[0046] In this embodiment, the memory testing module 10 is integrated into the DDR controller of the memory system. Its detection range is limited to the memory modules on one or more memory channels managed by the DDR controller. The UCE error address detected by the module is the address of the memory channel corresponding to the memory module with the UCE error.

[0047] Specifically, in this embodiment Figure 1 The black solid line represents the routing path, through which normal memory access is performed. The blue dashed line represents the repair path, through which table insertion is performed, replacing the memorypage address and adding it to the routing table, and log report is submitted to management module 20. The red solid line represents the detection path, through which memory testing module 10 reports the UCE error address to management module 20 after detecting a memory UCE error.

[0048] Among them, HOST0~HOSTn are different computing units that initiate memory access requests; DIMM0 and DIMM1 are memory modules in the memory pool; PORT0~PORTn are different ports used to connect different memory channels; route_table0 and route_table1 are both routing tables; addr_trf is the address converter of management module 20, and log_manager is the log manager of management module 20.

[0049] Specifically, in this embodiment, the DDR controller is a double data rate synchronous DRAM (Dynamic Random Access Memory), connecting the processor and memory in the computer system. It handles memory access requests and coordinates read / write operations, achieving double bandwidth by synchronizing clock signals and transmitting data on both the rising and falling edges. Its operation includes: The DDR controller receives memory access requests from the processor, which include information such as logical address and read / write type.

[0050] The logical address is translated into a physical address, and the specific physical structure to be accessed is determined based on the physical address, including but not limited to memory slots (Socket), channels (Channel), ranks (Rank), rows (Row), and columns (Column).

[0051] The operation commands are generated and scheduled according to DRAM timing requirements, including read commands and write commands for memory.

[0052] In some embodiments, the memory test module 10 is an independent module that establishes a communication connection with the DDR controller through an internal interface. The memory test module 10 sends a test request to the DDR controller, the DDR controller processes the test request, coordinates the reading of memory data and sends it to the memory test module 10, and the memory test module 10 performs error address detection on the memory data.

[0053] In other embodiments, the memory testing module 10 is integrated by embedding BIST (Built-in Self-Test, memory testing) logic into key sub-modules of the DDR controller. These key sub-modules include, but are not limited to, an address decoding unit, a command scheduling unit, and a data comparison unit. The management module 20 is responsible for coordinating BIST operations on multiple DDR controllers and issuing commands to the memory testing modules 10 on each DDR controller via the system bus, including but not limited to configuration commands, start commands, and test result reporting commands.

[0054] Specifically, in this embodiment, the memory testing module 10 integrates a testing circuit inside the DDR controller, enabling the memory system to autonomously complete the detection of memory module error addresses without relying on external, expensive, and complex automatic testing equipment. This significantly reduces equipment dependence and testing time, thereby lowering the overall cost.

[0055] Specifically, the working process of the memory testing module 10 in this embodiment includes: During the system startup phase, the memory test module 10 receives the BIST startup command and then performs rapid initialization and test parameter configuration. The test parameters include: the memory range to be tested, the test algorithm, and the granularity setting. In this embodiment, the granularity is set to 64 bytes of cache line.

[0056] The memory test module 10 performs error address detection on the memory modules on one or more memory channels governed by its bound DDR controller.

[0057] The detected UCE error address is parsed, and the generated UCE error information is reported to the management module 20.

[0058] Specifically, the UCE error information described in this embodiment includes, but is not limited to: DIMM information: Slot, Device ID; Rank / Die information: Physical array (rank), Sub-rank physical array, Memory die; Bank information: Logical array group (Bank Group), Logical array (Bank); Row and column information: row address and column address; Topology information: Socket, Memory Controller (IMC), Channel; Internal information of the memory chip: Chip ID and Transfer Number.

[0059] Among them, the device identifier is the memory identifier of different manufacturers or types; the row and column information is used to jointly locate the specific location of the cacheline in the bank; the internal information of the chip is used to locate the DRAM and memory channel with UCE error information.

[0060] In this embodiment, the process of determining the repair status of the host memory address by the address routing module 30 includes: Send the memory access address to the host and perform routing table entry calculation.

[0061] Using the calculation results from the routing table entries, access the routing table entries and determine the repair status of the host memory address.

[0062] If the host memory address is a normally accessible memory address, then its repair status is "not repaired".

[0063] If the host memory address is a repaired memory page address, then its repair status is repaired.

[0064] In this embodiment, the routing table entry is an entry for address translation and address routing merging.

[0065] In this embodiment, the address routing module 30 obtains the host memory address HOST PHYSICS ADDRESS (HPA). If it is determined that the address is in an unrepaired state, the HPA is converted into a readable and writable DDR control channel address DDR CHANNELADDRESS (CPA), and the CPA is routed to the port where the memory module is located. If it is determined that the address is in a repaired state, the HPA is converted into the out-of-band address OOB ADDRESS of the out-of-band storage module 40 and routed to the port where the out-of-band storage module 40 is located.

[0066] In some embodiments, the routing table entries are calculated using hash calculation. The host memory address is calculated using a hash function, and a routing table index is generated. The corresponding entry in the routing table is accessed based on this index.

[0067] Specifically, in this embodiment, the complete host memory address is input, and a hash function is used to generate a value within a fixed range as an index for the routing table, thereby mapping a large range of address space to a small range of routing table indexes. The formula for calculating the fixed range of values ​​is: Index = Hash_Function(HPA) % Table_Size Where Hash_Function() is the hash function, HPA is the input host memory address, and Table_Size is the size of the routing table, that is, the value range of Index is: 0≤Index≤Table_Size.

[0068] In other embodiments, the routing table entries are calculated using direct address access, by extracting a portion of the host memory address as the routing table index, with the index value being HPA[Bit_Mask]. For example, the high-order bits of the host memory address can be extracted as input, and the index value can be HPA[63:20]. This value is then used to directly access the corresponding entry in the routing table.

[0069] In this embodiment, the management module 20 generates a memory repair log and writes the received repair information into the memory repair log. The memory repair log is used to query the repair status of the memory page address and the RAS log information of the memory repair system. It includes: UCE error address and information, repaired memory page address and information, the address of the repaired out-of-band storage module 40, repair result information, and exception cause information. The memory repair log can be set to the content shown in Table 1 below.

[0070] Table 1

[0071] Specifically, the memory repair log described in this embodiment records the repair status of the memory system, which can meet the serviceability requirements of the RAS system. Through this memory repair log, memory fault diagnosis, predictive maintenance, and memory resource monitoring can be performed, enhancing the memory-pooled RAS system and facilitating unified memory management in the data center.

[0072] In this embodiment, the out-of-band storage module 40 has an address code independent of the system memory pool. This address code is managed by the management module 20, and the address size is determined at the memory page level during system initialization.

[0073] The address routing module 30 routes memory access commands for damaged and repaired addresses to the out-of-band storage module 40 according to the routing table entries, and the out-of-band storage module 40 performs memory read or write operations according to the memory access commands.

[0074] Specifically, in this embodiment, the external storage module 40 is an independent storage hardware, which can be selected from DDR, flash, SSD or on-chip RAM space. It is separate from the memory modules (DIMMs) that make up the system memory pool. It is dedicated to memory repair and is therefore not affected by system memory pool failures or frequent access, thus improving the reliability of the memory system.

[0075] Specifically, this embodiment addresses the limitations of single-memory module structure by designing an out-of-band storage module 40 independent of the system memory pool, compared to traditional PPR memory repair solutions. PPR memory repair is performed on a row-by-row basis, and the number of usable ROWs is limited by the DDR's factory specifications. According to the DDR4 specification, one bank needs one spare ROW. When an unrepairable error within a bank involves two ROW units, the PPR repair technology fails. However, the memory repair technology in this embodiment is no longer limited by the DIMM module's factory specifications and can utilize separate out-of-band storage space as redundant resources for repair.

[0076] In this embodiment, the memory page of the out-of-band storage space of the out-of-band storage module 40 is defined as the memorypage. The page size of this memory page can be dynamically configured according to the memory damage situation, and its value is:

[0077] Where n is a non-negative integer.

[0078] The dynamic configuration range of its page size is:

[0079] in, This is the maximum capacity of the out-of-band storage space.

[0080] Specifically, in this embodiment, a smaller memory page size indicates a finer granularity of repairable memory, resulting in higher repair accuracy for fewer damaged bits and saving redundant memory. Conversely, a larger memory page size indicates a larger granularity of repairable memory, consuming more redundant memory, but requiring fewer routing table entries to repair the same memory space. Therefore, the choice of memory page size can be flexibly adjusted based on the memory damage situation. When the system has many large areas of damaged addresses with contiguous HPA, a larger memory page size should be chosen; when the damaged memory addresses are not contiguous and are damaged at a finer granularity, a smaller memory page size should be chosen.

[0081] Specifically, the flexible configurable memory page size of the external storage module 40 in this embodiment solves the problem of excessive redundant memory compared to the traditional ADDDC memory repair solution. The ADDDC memory repair solution performs memory redundancy backup on a RANK or BANK basis, requiring significant memory redundancy resources. In contrast, the memory repair technology in this embodiment performs memory repair on a dynamically adjustable memory page size basis, adapting to a wider range of scenarios and requiring less redundant memory resources for repair.

[0082] As can be seen, compared with traditional memory repair methods such as row (2K) and bank (GB level), this invention improves the utilization efficiency of redundant memory and compatibility with the operating system by repairing memory at the memorypage granularity. It can save the resources required for repair and solve the problem of excessively large traditional redundant memory. Moreover, by repairing memory at the dynamically adjustable memory page size, it enhances the repairability of damaged addresses and is applicable to a wider range of scenarios.

[0083] See Figure 2 This invention discloses a high-efficiency memory repair method based on pooled memory, applied to a high-efficiency memory repair system based on pooled memory, comprising: S1: The memory testing module performs error address detection on the memory modules it manages at the cacheline level and reports the detected UCE error addresses.

[0084] S2: After receiving the UCE error address, the management module converts it into the corresponding memory page address and checks whether the memory page address has been repaired. If it has not been repaired, a repair command for the memory page address is issued.

[0085] S3: After receiving the repair command, the address routing module uses a routing lookup algorithm to replace and add the memory page address to the routing table, and reports the repair result to the management module; if the entry for the memory page address can be added, the repair result is successful, otherwise it is a failure.

[0086] S4: The management module reports the repair results in the form of a memory repair log, and the memory system queries the RAS log information through the memory repair log.

[0087] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0088] The above embodiments are merely preferred embodiments of the present invention and should not be construed as limiting the scope of protection of the present invention. Any non-substantial changes and substitutions made by those skilled in the art based on the present invention shall fall within the scope of protection claimed by the present invention.

Claims

1. A high-efficiency memory repair system based on pooled memory, characterized in that, include: The system includes a memory testing module, a management module, an address routing module, and an out-of-band storage module. The memory testing module performs error address detection on the memory module at the cacheline level and reports the detected UCE error addresses to the management module. The management module is used to convert the UCE error address into its corresponding memory page address, query the repair status of the memory page address, generate a repair command for the memory page address based on its repair status, and send it to the address routing module. The address routing module uses a routing lookup algorithm to replace and add the memory page address to the routing table according to the repair command to complete the address repair. It also performs corresponding address translation and routes the memory to the memory module or the port where the out-of-band storage module is located based on the repair status of the accessed host memory address. The out-of-band storage module is an out-of-band memory pool independent of the system memory pool, used to repair damaged memory at the memory page level.

2. The high-efficiency memory repair system based on pooled memory according to claim 1, characterized in that: The memory testing module is integrated into the DDR controller of the memory system. Its detection range is limited to the memory modules on one or more memory channels managed by the DDR controller. The UCE error address detected by the module is the address of the memory channel corresponding to the memory module with the UCE error.

3. The high-efficiency memory repair system based on pooled memory according to claim 2, characterized in that: The memory testing module will report the UCE error information detected by the memory module to the management module. The UCE error information includes: the slot where the error occurred, the memory chip, the memory controller, the channel, the slot, the physical array, the sub-physical array, the logical array group, the logical array and the row address and column address, the device identifier, the memory chip identifier and the transfer number.

4. The high-efficiency memory repair system based on pooled memory according to claim 2, characterized in that: The process by which the address routing module determines the repair status of the host memory address includes: Send the memory access address to the host and perform routing table entry calculation; Access the routing table entry using the calculation result, and determine the repair status of the host memory address: If the host memory address is a normally accessible memory address, then its repair status is unrepaired; If the host memory address is a repaired memory page address, then its repair status is repaired.

5. The high-efficiency memory repair system based on pooled memory according to claim 4, characterized in that: The routing table entry is an entry for address translation and address routing merging. It is calculated using a hash function. The host memory address is calculated and a routing table index is generated. The corresponding entry in the routing table is accessed based on the index.

6. The high-efficiency memory repair system based on pooled memory according to claim 4, characterized in that: The address routing module obtains the host memory address. If it determines that the address is in an unrepaired state, it converts it into a readable and writable DDR control channel address and routes it to the port where the memory module is located. If it determines that the address is in a repaired state, it converts it into the address of the out-of-band storage module and routes it to the port where the out-of-band storage module is located.

7. The high-efficiency memory repair system based on pooled memory according to any one of claims 1-6, characterized in that: The management module generates a memory repair log, which is used to query the repair status of the memory page address and the RAS log information of the memory repair system. The log includes: UCE error address, repaired memory page address, address of the repaired out-of-band storage module, repair result information, and abnormal reason information.

8. The high-efficiency memory repair system based on pooled memory according to claim 6, characterized in that: The out-of-band storage module has an address code independent of the system memory pool. This address code is managed by the management module, and the address size is determined during system initialization with memory page granularity. The address routing module routes memory access commands for damaged and repaired addresses to the out-of-band storage module according to the routing table entries, and the out-of-band storage module performs memory read or write operations according to the memory access commands.

9. The high-efficiency memory repair system based on pooled memory according to claim 8, characterized in that: The memory page of the out-of-band storage space of the out-of-band storage module is defined as the memory page. The page size of this memory page can be dynamically configured according to the memory damage situation, and its value is: Where n is a non-negative integer; The dynamic configuration range of its page size is: in, This is the maximum capacity of the out-of-band storage space.

10. A high-efficiency memory repair method based on pooled memory, characterized in that, An application to a high-efficiency memory repair system based on pooled memory as described in any one of claims 1-9, comprising: S1: The memory testing module performs error address detection on the memory modules it manages at the cacheline level and reports the detected UCE error addresses; S2: After receiving the UCE error address, the management module converts it into the corresponding memory page address and checks whether the memory page address has been repaired. If it has not been repaired, a repair command for the memory page address is issued. S3: After receiving the repair command, the address routing module uses a routing lookup algorithm to replace and add the memory page address to the routing table, and reports the repair result to the management module; if the entry for the memory page address can be added, the repair result is successful, otherwise it is a failure. S4: The management module reports the repair results in the form of a memory repair log, and the memory system queries the RAS log information through the memory repair log.

Citation Information

Patent Citations

  • Data writing method and device of solid state disk and computer readable storage medium

    CN109753241A

  • COW snapshot technology for data storage and data disaster recovery

    CN111190770A

  • Out-of-band processing method and device capable of correcting errors of memory, equipment and medium

    CN115129508A

  • Memory system and method of operating memory system

    CN118363520A

  • An integrated circuit device and method for reading data from an SRAM memory

    EP3128428A1