A method for processing memory failure and related equipment

Through BMC monitoring of the number of correctable faults and PCLS times of the DRAM memory channel, directly isolating the fault bank or cell, solving the problem of waste of resources in DRAM memory fault processing and improving the efficiency of memory fault processing.

CN115712518BActive Publication Date: 2025-08-19XFUSION DIGITAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211303134.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-24
Publication Date
2025-08-19
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

In the DRAM memory failure handling, a large number of faults can be corrected in a short time, resulting in resource waste, especially after the number of cache line retention (PCLS) is exhausted, a more costly solution is required to cause resource waste.

Method used

The number of correctable faults and PCLS times of the memory channel are obtained through BMC. If the number of PCLS times exceeds, the fault bank will be directly isolated, otherwise the fault cell will be isolated to avoid waste of resources.

Benefits of technology

It effectively saves the number of PCLS usage, avoids resource waste, and improves the efficiency of memory failure handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115712518B_ABST
    Figure CN115712518B_ABST
Patent Text Reader

Abstract

The present application discloses a method for handling memory faults and related devices for avoiding unnecessary waste of resources. The method in an embodiment of the present application includes: the BMC obtains the number of correctable faults of the target memory channel and the number of times that a portion of the cache lines can retain PCLS. The number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time period. If the number of correctable faults is greater than the number of PCLS that the target memory channel can provide, the BMC sends a first indication message to the CPU, which is used to instruct the CPU to isolate the bank where the correctable fault is located.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computers, and in particular to a method for processing memory failures and related devices. Background Art

[0002] Dynamic random access memory (DRAM) is a type of memory widely used in storage and IT applications. However, as DRAM integration increases and process sizes shrink, DRAM failure rates are also increasing, even causing server downtime. When a correctable DRAM failure occurs, partial cache line sparing (PCLS) is often considered to mitigate the fault, as it offers the lowest isolation granularity and cost.

[0003] Each memory channel of each CPU can only provide a limited number of PCLS. Existing technologies wait until these limits are exhausted before eliminating correctable faults through isolation granularity and more costly solutions. However, if the number of correctable faults generated by a memory channel in a short period of time exceeds the number of PCLS that the memory channel can provide, existing technologies will result in a waste of resources. Summary of the Invention

[0004] The embodiments of the present application provide a method for handling memory failures and related devices to avoid unnecessary waste of resources.

[0005] The first aspect of the present application provides a method for handling memory failures:

[0006] In a server, each CPU includes multiple memory channels. The baseboard management controller (BMC) obtains the number of correctable faults in a target memory channel and the number of times a partial cache line can retain a cache line (PCLS). The correctable fault number is the total number of correctable faults generated by the target memory channel within a preset time period. If the number of correctable faults is greater than the number of PCLS that the target memory channel can retain, the BMC sends a first indication to the CPU, instructing the BIOS to isolate the bank containing the correctable fault.

[0007] In this application, when the number of correctable faults is greater than the number of PCLS that the target memory channel can provide, it means that not all correctable faults can be eliminated based on PCLS alone. Therefore, the bank where the correctable faults are located is directly isolated, thereby directly eliminating all correctable faults and saving the number of PCLS that the target memory channel can provide.

[0008] In a possible implementation, if the number of correctable faults is less than or equal to the number of PCLSs that can be provided by the target memory channel, the BMC sends second indication information to the CPU, where the second indication information is used to instruct the CPU to isolate the cell where the correctable fault is located.

[0009] In this application, when the number of correctable faults is less than or equal to the number of PCLS that the target memory channel can provide, it means that all correctable faults can be eliminated based on PCLS alone. Therefore, only the cells where the correctable faults are located are isolated, avoiding the isolation of normal cells.

[0010] In a possible implementation, the BMC obtains the number of correctable faults based on the number of correctable fault information from the BIOS within a preset time period, where the correctable fault information indicates that a correctable fault has occurred in the target memory channel.

[0011] In a possible implementation, the preset duration is 500 milliseconds.

[0012] In a possible implementation, the BMC and the CPU are set in the same computing device, and the target memory channel is any memory channel in the computing device.

[0013] The second aspect of the present application provides a method for handling memory failures:

[0014] If the number of correctable faults in the target memory channel is greater than the number of PCLSs that the target memory channel can provide, the CPU receives first indication information from the BMC, where the number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time period. The CPU isolates the bank where the correctable faults are located based on the first indication information.

[0015] In one possible implementation, if the number of correctable faults is less than or equal to the number of PCLSs that the target memory channel can provide, the CPU receives second indication information from the BMC and isolates the cell where the correctable fault is located according to the second indication information.

[0016] In a possible implementation, the preset duration is 500 milliseconds.

[0017] In a possible implementation, the BMC and the CPU are set in the same computing device, and the target memory channel is any memory channel in the computing device.

[0018] The third aspect of the present application provides a BMC:

[0019] It includes an acquisition unit for acquiring the number of correctable faults of the target memory channel and the number of PCLS that can be provided. The number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time length.

[0020] The sending unit is configured to send first indication information to a central processing unit (CPU) if the number of correctable faults is greater than the number of PCLSs that can be provided by the target memory channel. The first indication information is used to instruct the CPU to isolate the bank where the correctable faults are located.

[0021] In one possible implementation, the sending unit is further configured to send second indication information to the CPU if the number of correctable faults is less than or equal to the number of PCLSs that the target memory channel can provide, and the second indication information is configured to instruct the CPU to isolate the cell where the correctable fault is located.

[0022] In a possible implementation, the acquiring unit is specifically configured to acquire the number of correctable faults according to the number of correctable fault information from the CPU within a preset time period, where the correctable fault information indicates that a correctable fault has occurred in the target memory channel.

[0023] In a possible implementation, the preset duration is 500 milliseconds.

[0024] In a possible implementation, the BMC and the CPU are set in the same computing device, and the target memory channel is any memory channel in the computing device.

[0025] A fourth aspect of the present application provides a CPU:

[0026] The system includes a receiving unit configured to receive first indication information from a BMC if the number of correctable faults of a target memory channel is greater than the number of PCLSs that can be provided by the target memory channel, where the number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time period.

[0027] The processing unit is configured to isolate the bank where the correctable fault is located according to the first indication information.

[0028] In one possible implementation, the receiving unit is further configured to receive second indication information from the BMC if the number of correctable faults is less than or equal to the number of PCLSs that can be provided by the target memory channel;

[0029] The processing unit is further configured for the CPU to isolate the cell where the correctable fault is located according to the second indication information.

[0030] In a possible implementation, the preset duration is 500 milliseconds.

[0031] In a possible implementation, the BMC and the CPU are set in the same computing device, and the target memory channel is any memory channel in the computing device.

[0032] In a fifth aspect, the present application provides a computing device including a memory, a CPU, a storage chip, and a BMC. The CPU is connected to the memory, the storage chip, and the BMC. The storage chip stores a BIOS. The CPU is used to run the BIOS. The BMC is used to send a first indication message to the CPU based on the number of times the number of correctable faults of the target memory channel is greater than the number of available PCLS. The CPU is used to isolate the bank where the correctable faults are located based on the first indication message. The number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time length.

[0033] In a possible implementation, the BMC is further configured to send second indication information to the CPU based on the number of times the number of correctable faults is less than or equal to the number of PCLSs that can be provided by the target memory channel. The second indication information is configured to instruct the CPU to isolate the cell where the correctable faults are located.

[0034] In a possible implementation, the target memory channel is any memory channel in the computing device.

[0035] In a sixth aspect, the present application provides a BMC, comprising a processor coupled to a memory, the memory being used to store instructions. When the instructions are executed by the processor, the BMC executes the method in the first aspect.

[0036] In a seventh aspect, the present application provides a CPU, comprising a processor, the processor being coupled to a memory, the memory being used to store instructions, and when the instructions are executed by the processor, the CPU executes the method in the second aspect.

[0037] An eighth aspect of an embodiment of the present application provides a computer-readable storage medium having a computer program or instruction stored thereon. When the computer program or instruction is executed, the computer executes the method in the first or second aspect described above.

[0038] In a ninth aspect, the present application provides a computer program product, comprising computer instructions or programs, which, when executed, enable the computer to execute the method in the first or second aspect described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 A schematic diagram of bnak;

[0040] Figure 2 A schematic diagram of the system architecture used in this application;

[0041] Figure 3 A flowchart of a method for handling memory failures in this application;

[0042] Figure 4 A schematic diagram of reporting CE information for memory;

[0043] Figure 5 A structural diagram of the BMC in this application;

[0044] Figure 6 A schematic diagram of the structure of the CPU in this application;

[0045] Figure 7 This is another structural diagram of the BMC in this application. DETAILED DESCRIPTION

[0046] The following describes the embodiments of the present application in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present application, rather than all the embodiments. Those skilled in the art will appreciate that with the development of technology and the emergence of new scenarios, the technical solutions provided in this application are also applicable to similar technical problems.

[0047] The terms "first," "second," and the like in the specification and claims of this application and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having," as well as any variations thereof, are intended to cover non-exclusive inclusions, e.g., a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0048] DRAM is a type of memory commonly used in servers. As DRAM integration becomes increasingly higher and process technology becomes smaller, the failure rate of DRAM is also increasing. DRAM failures are divided into correctable failures and uncorrectable failures. Correctable failures are those that can be corrected within the fault tolerance range of the hardware platform. The system or application will not stop running due to correctable failures. Figure 1 DRAM can be divided into multiple banks, each of which is divided into rows and columns. Each row and column consists of multiple cells, which are the most basic unit in DRAM. When a correctable fault occurs in a DRAM, the DRAM reports a correctable error (CE) message. The CE message indicates the cell where the correctable fault is located through the row and column identifiers.

[0049] There are different solutions for eliminating correctable faults, including PCLS. PCLS is the solution with the lowest isolation granularity and cost. The principle of PCLS is to isolate the cell where the correctable fault is located. One PCLS can isolate one cell, which means it can eliminate one correctable fault. There are also some solutions with higher isolation granularity and cost, such as adaptive double device data correction (ADDDC Sparing). ADDDC Sparing works by isolating the entire bank where the correctable fault is located. Therefore, one ADDDC Sparing can eliminate a large number of correctable faults, but it also has a significant negative impact on memory performance.

[0050] The number of PCLS and ADDDC Sparing that each memory channel of each CPU can provide is limited, generally 16 times for PCLS and two times for ADDDC Sparing. In the prior art, when a correctable fault occurs, the correctable fault is usually eliminated first according to PCLS, and then the correctable fault is eliminated according to ADDDC Sparing after the number is exhausted. If the time interval between correctable faults in a memory channel is relatively long, the means of the prior art are considered acceptable. However, if the number of correctable faults generated by a memory channel in a short period of time exceeds the number of PCLS that the memory channel can provide, for example, if 50 correctable faults occur in a bank of a memory channel in a short period of time, in this case, if the correctable faults are eliminated first according to PCLS, only a part of the correctable faults can be eliminated, and the bank must be isolated later to eliminate all the correctable faults. Compared with directly isolating the bank, the means of the prior art waste the number of PCLS that the memory channel can provide.

[0051] The present application provides a method for handling memory failures and related devices to avoid unnecessary waste of resources.

[0052] This application can be applied to Figure 2 In the system architecture shown in Figure 2As shown, a computing device may include memory, a CPU, a baseboard management controller (BMC), and a storage chip. The computing device may be a server, a computer, a laptop, or a tablet. This embodiment is described using a server as the computing device. The CPU is connected to the memory, the BMC, and the storage chip, respectively. The storage chip stores the program code for the basic input and output system (BIOS). The BIOS is a set of programs embedded in a ROM chip on the motherboard of the server. It stores the computer's most important basic input and output programs, the power-on self-test program, and the system startup program. The CPU can read the program code in the storage chip to run the BIOS, providing the lowest-level and most direct hardware settings and control for the server. Of course, in another implementation, the BIOS program code can also be stored directly in the CPU. The BMC is a core component of the server management system defined by the Intelligent Platform Management Interface (IPIM). IPIM is a set of computer interface specifications defined for autonomous computer subsystems, used to provide management and monitoring functions for hardware and software such as the CPU, firmware, and operating system, independent of the host system.

[0053] See also Figure 3 The following is an introduction to the memory failure handling method in this application:

[0054] 301. The BMC obtains the number of correctable faults of the target memory channel and the number of available PCLS. The number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time period.

[0055] In this embodiment, CPU1 in the server runs BIOS as an example. Figure 4, assuming that a correctable fault occurs in the memory, the memory will send a CE message to the BIOS running on CPU1. The CE message indicates the CPU, memory channel, and cell where the correctable fault is located. After the BIOS running on CPU1 receives the CE message, it will send the CE message to the BMC. Therefore, the BMC can count the total number of correctable faults generated within the preset time length for each memory channel of each CPU in the server through the CE messages received within the preset time length. The above-mentioned preset time length is a relatively short time length, such as 500 milliseconds. In addition, the BMC can also obtain the number of PCLS that can be provided by each memory channel of each CPU. It should be understood that the above-mentioned number of PCLS that can be provided is the number of remaining PCLS that can be provided, not the maximum number of PCLS that can be provided. Exemplarily, the BMC can set a counter for the number of PCLS that can be provided by the memory channel. In the initial state, the counter is the maximum number of PCLS that can be provided by the memory channel. Every time the memory channel provides a PCLS, the value of the counter is reduced by 1.

[0056] It should be noted that the target memory channel can be any memory channel of any CPU.

[0057] 302. If the number of correctable faults is greater than the number of PCLSs that can be provided by the target memory channel, the BMC sends first indication information to the CPU, where the first indication information is used to instruct the CPU to isolate the bank where the correctable faults are located.

[0058] Take the target memory channel as CPU1's memory channel 1 as an example. Please continue to refer to Figure 4 If the total number of correctable faults generated by memory channel 1 within the preset time period is greater than the number of PCLSs that memory channel 1 can provide, the BMC determines that it is impossible to isolate all cells where the correctable faults generated by memory channel 1 within the preset time period are located at a cell granularity. Therefore, the BMC sends a first instruction message to the BIOS running on CPU 1. The first instruction message is used to instruct the BIOS running on CPU 1 to directly isolate the bank where the correctable faults generated by memory channel 1 within the preset time period are located.

[0059] For example, memory channel 1 generates 50 correctable faults within 500 milliseconds, and the number of PCLSs that memory channel 1 can provide is 16. Obviously, the 16 PCLSs that memory channel 1 can provide can only isolate the cells where 16 of the 50 correctable faults are located. If existing technologies are used, after the number of PCLSs that memory channel 1 can provide reaches zero, the bank where the 50 correctable faults are located must still be isolated. This not only exhausts all the PCLSs that memory channel 1 can provide, but also cannot avoid isolating the bank. Therefore, the BMC sends a first instruction message to the BIOS running on CPU 1. This first instruction message is used to instruct the BIOS running on CPU 1 to isolate the bank where the 50 correctable faults generated by memory channel 1 within the aforementioned 500 milliseconds are located. After receiving the first instruction message, the BIOS running on CPU 1 initiates a request to memory channel 1 and isolates the bank where the correctable faults are located based on the ADDDC Sparing resources provided by memory channel 1.

[0060] Please continue reading Figure 4 Alternatively, in another case, if the total number of correctable faults generated by memory channel 1 within the preset time period is less than or equal to the number of PCLSs that memory channel 1 can provide, the BMC determines that all cells where the correctable faults generated by memory channel 1 within the preset time period are located can be isolated at a cell granularity, and therefore the BMC sends second indication information to the BIOS running on CPU 1, where the second indication information is used to instruct the BIOS running on CPU 1 to isolate all cells where the correctable faults generated by memory channel 1 within the preset time period are located at a cell granularity.

[0061] For example, memory channel 1 generates 15 correctable faults within 500 milliseconds, and the number of PCLSs that memory channel 1 can provide is 16. Obviously, the 16 PCLSs that memory channel 1 can provide are sufficient to isolate the cells where these 15 correctable faults are located. Therefore, the BMC sends a second instruction to the BIOS running on CPU 1. This second instruction is used to instruct the BIOS running on CPU 1 to isolate all cells where the 15 correctable faults generated by memory channel 1 within the 500 milliseconds are located, based on the cell granularity. Subsequently, the BIOS running on CPU 1 initiates a request to memory channel 1 and, based on the PCLS resources provided by memory channel 1, isolates all cells where these correctable faults are located.

[0062] In this application, when the number of correctable faults is greater than the number of PCLS that the target memory channel can provide, it means that it is impossible to isolate the cells where all correctable faults are located based on PCLS alone. Therefore, the BMC instructs the CPU to directly isolate the bank where the correctable faults are located, thereby directly eliminating all correctable faults and saving the number of PCLS that the target memory channel can provide.

[0063] The above describes the method for handling memory failures in this application. The following describes the BMC and BIOS in this application:

[0064] See also Figure 5 The BMC500 in this application includes an acquisition unit 501 and a sending unit 502. The BMC500 is used to implement the aforementioned Figure 3 Operations performed by the BMC in the illustrated embodiment.

[0065] The acquisition unit 501 is configured to acquire the number of correctable faults of the target memory channel and the number of times that the partial cache line retains PCLS. The number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time period.

[0066] The sending unit 502 is configured to send first indication information to the CPU if the number of correctable faults is greater than the number of PCLSs that can be provided by the target memory channel. The first indication information is configured to instruct the CPU to isolate the bank where the correctable faults are located.

[0067] In one possible implementation,

[0068] The sending unit 502 is further configured to send second indication information to the CPU if the number of correctable faults is less than or equal to the number of PCLSs that the target memory channel can provide, wherein the second indication information is configured to instruct the CPU to isolate the cell where the correctable fault is located.

[0069] In one possible implementation,

[0070] The acquiring unit 501 is specifically configured to acquire the number of correctable faults according to the number of correctable fault information from the CPU within a preset time period, where the correctable fault information indicates that a correctable fault has occurred in the target memory channel.

[0071] In a possible implementation, the preset duration is 500 milliseconds.

[0072] In a possible implementation, the BMC 500 and the CPU are set in the same computing device, and the target memory channel is any memory channel in the computing device.

[0073] See also Figure 6The CPU 600 in this application includes a receiving unit 601 and a processing unit 602. The CPU 600 is used to execute the aforementioned Figure 3 The operations performed by the CPU in the embodiment shown in .

[0074] The receiving unit 601 is configured to receive first indication information from the BMC if the number of correctable faults of the target memory channel is greater than the number of PCLSs that the target memory channel can provide, where the number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time period.

[0075] The processing unit 602 is configured to isolate the bank where the correctable fault is located according to the first indication information.

[0076] In one possible implementation,

[0077] The receiving unit 601 is further configured to receive second indication information from the BMC if the number of correctable faults is less than or equal to the number of PCLSs that can be provided by the target memory channel.

[0078] The processing unit 602 is further configured to isolate the cell where the correctable fault is located according to the second indication information.

[0079] In a possible implementation, the preset duration is 500 milliseconds.

[0080] In a possible implementation, the BMC and the CPU 600 are provided in the same computing device, and the target memory channel is any memory channel in the computing device.

[0081] Figure 7 Schematic diagram of a BMC structure provided by the present application, which is used to implement the method executed by the BMC in the above embodiment. BMC 700 may include one or more central processing units (CPUs) 701 and a memory 705, wherein the memory 705 stores one or more application programs or data.

[0082] Memory 705 can be volatile or persistent storage. Programs stored in memory 705 can include one or more modules, each of which can include a series of instruction operations on the server. Furthermore, the central processing unit 701 can be configured to communicate with memory 705 and execute the series of instruction operations in memory 705 on the BMC 700. The BMC 700 can also include one or more power supplies 902, one or more wired or wireless network interfaces 703, one or more input / output interfaces 704, and / or one or more operating systems.

[0083] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0084] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0085] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0086] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0087] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk, and other media that can store program code.

Claims

1. A method for handling memory failure, characterized in that: include: The baseboard management controller BMC obtains the number of correctable faults of the target memory channel and the number of times that the available partial cache lines retain PCLS, where the number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time period; If the number of correctable faults is greater than the number of PCLSs that can be provided by the target memory channel, the BMC sends first indication information to a central processing unit (CPU), where the first indication information is used to instruct the CPU to isolate the bank where the correctable faults are located.

2. The method according to claim 1, characterized in that The method further comprises: If the number of correctable faults is less than or equal to the number of PCLSs that can be provided by the target memory channel, the BMC sends second indication information to the CPU, where the second indication information is used to instruct the CPU to isolate the cell where the correctable fault is located.

3. The method according to claim 1 or 2, characterized in that The BMC obtains the number of correctable faults of the target memory channel including: The BMC obtains the number of correctable faults according to the number of correctable fault information from the CPU within the preset time period, where the correctable fault information indicates that a correctable fault has occurred in the target memory channel.

4. The method according to any one of claims 1 to 3, characterized in that The preset duration is 500 milliseconds.

5. The method according to any one of claims 1 to 4, characterized in that The BMC and the CPU are set in the same computing device, and the target memory channel is any memory channel in the computing device.

6. A method for handling memory failure, characterized in that: include: If the number of correctable faults of the target memory channel is greater than the number of PCLSs that the target memory channel can provide, the CPU receives first indication information from the baseboard management controller (BMC), where the number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time period; The CPU isolates the bank where the correctable fault is located according to the first indication information.

7. The method according to claim 6, characterized in that The method further comprises: If the number of correctable faults is less than or equal to the number of PCLSs that can be provided by the target memory channel, the CPU receives second indication information from the BMC; The CPU isolates the cell where the correctable fault is located according to the second indication information.

8. A computing device, characterized in that The computing device includes a memory, a CPU, a storage chip, and a BMC. The CPU is connected to the memory, the storage chip, and the BMC. The storage chip stores a BIOS. The CPU is configured to run the BIOS. The BMC is configured to send first indication information to the CPU based on the number of correctable faults of a target memory channel being greater than the number of times that a partial cache line retains PCLS. The CPU is configured to isolate a bank where the correctable fault is located based on the first indication information. The number of correctable faults is the total number of correctable faults generated by the target memory channel within a preset time period.

9. The computing device according to claim 8, wherein: The BMC is further configured to send second indication information to the CPU based on the number of times the number of correctable faults is less than or equal to the number of PCLSs that can be provided by the target memory channel, wherein the second indication information is configured to instruct the CPU to isolate the cell where the correctable fault is located.

10. The computing device according to claim 8 or 9, characterized in that The target memory channel is any memory channel in the computing device.

Citation Information

Patent Citations

  • Memory fault processing method and device

    CN114064333A

  • Memory test device, memory test method and memory test program

    JP2013025452A