Memory error information processing method and computing device

The substrate management controller executes system management interrupts to the central processor when preset conditions are met, which solves the problem of triggering system management interrupts when collecting memory error information, and improves the business real-timeness of computing devices.

WO2025123614A1PCT designated stage expired Publication Date: 2025-06-19XFUSION DIGITAL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/097416
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-12
Filing Date
2024-06-05
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

When collecting memory error information, the prior art can easily trigger system management interrupts (SMIs), causing the operating system to hang up and affecting the real-time business of computing devices.

Method used

Use the Platform Environmental Control Interface (PECI) through the Baseboard Management Controller (BMC) to obtain memory error information from the central processor's registers and perform system management interrupts to the central processor when preset conditions are met to reset the error counter in the memory particles.

Benefits of technology

It reduces the number of interrupts in system management when the number of errors does not meet the preset conditions, improves the real-timeness of computing device services, and reduces the impact on the business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024097416_19062025_PF_FP_ABST
    Figure CN2024097416_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of computers. Disclosed are a memory error information processing method and a computing device, which ensure the real-time performance of computing device services. The method comprises: by means of a platform environment control interface (PECI), a baseboard management controller acquires memory error information from a first register of a central processing unit; and, when the memory error information satisfies a preset condition, the baseboard management controller executes a system management interrupt to the central processing unit, so that the central processing unit resets the count value of a second register in a memory die, wherein the count value of the second register is used for indicating the number of errors of the memory die.
Need to check novelty before this filing date? Find Prior Art

Description

Memory error information processing method and computing device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 12, 2023, with application number 202311704802.6 and application name “A memory error information processing method and computing device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a memory error information processing method and a computing device. Background Art

[0003] With the rapid development of the computing industry, computing systems are placing increasingly higher demands on memory capacity and transmission speeds, making the new generation, fifth-generation double data rate synchronous dynamic random access memory (DDR5), increasingly mainstream. Compared to the traditional fourth-generation double data rate synchronous dynamic random access memory (DDR4), DDR5 adds an error check and scrub (ECS) function, allowing single-bit memory errors to be read and corrected within the memory chip and written back to the data storage unit. Memory error information is used to indicate the error correction status of the ECS, and is therefore also referred to as ECS information. ECS information can, to a certain extent, reflect the concentration of minor faults within the memory chip.

[0004] In order to identify potential memory failures as early as possible, there is an urgent need for a solution that can obtain ECS information in a timely manner with minimal impact on the operation of the computing system.

[0005] Summary of the Invention

[0006] Embodiments of the present application provide a memory error information processing method and a computing device to solve the problem that a system management interrupt (SMI) is triggered during the ECS information collection process, causing the operating system to hang and affecting the real-time performance of the computing device business.

[0007] To implement the above technical solution, this application adopts the following technical solution:

[0008] In a first aspect, a method for processing memory error information is provided. A memory stick includes memory particles. The method can be applied to a baseboard management controller. The method includes: the baseboard management controller obtains memory error information from a first register of a central processing unit through a platform environment control interface (PECI). The memory error information is memory error information obtained when the memory particles complete an error check and error correction (ECS) inspection according to a first preset cycle. When the memory error information meets a preset condition, the baseboard management controller executes a system management interrupt to the central processing unit, so that the central processing unit resets the count value of a second register in the memory particle. The count value of the second register is used to indicate the number of times an error has occurred in the memory particle.

[0009] In this possible implementation, when the baseboard management controller determines that the memory error information meets the preset conditions, it executes a system management interrupt to the central processing unit so that the central processing unit resets the count value of the second register in the memory particle. When the baseboard management controller determines that the memory error information does not meet the preset conditions, there is no need to instruct the central processing unit to enter the system management interrupt step. Compared with the prior art, the baseboard management controller will enter the system management interrupt process as long as the central processing unit completes the error inspection of the memory bar. This reduces the number of times the central processing unit enters the system management interrupt when the number of errors generated in the dynamic random access memory does not meet the preset conditions, thereby improving the real-time performance of the computing device business. In addition, after entering the interrupt, the prior art not only needs to clear the count value of the register in the memory particle, but also needs to collect memory error information. In this possible implementation of the embodiment of the present application, the baseboard management controller actively reads the information of the first register in the central processing unit to obtain the memory error information, which is not restricted by whether the central processing unit enters the interrupt. This can reduce the time for the central processing unit to enter the interrupt and reduce the impact on the business.

[0010] In one possible implementation, the memory error information includes a count value of a second register representing the number of errors generated by the dynamic random access memory. The preset reset condition includes the count value in the second register being at least one of a first count value, a second count value, or a third count value, and the count value does not include zero; wherein the first count value is the total number of single-bit errors occurring in the memory cell after the memory cell completes an ECS inspection, and the first count value is greater than a first threshold; the second count value is the total number of rows in the memory cell that have single-bit errors occurring after the memory cell completes an ECS inspection, and the second count value is greater than a second threshold; and the third count value is the total number of single-bit errors in the row with the most single-bit errors occurring in the memory cell after the memory cell completes an ECS inspection, and the third count value is greater than a third threshold.

[0011] In one possible implementation, the memory error information also includes the address information of the row with the most single-bit errors in the memory particles. After the memory particles have completed a preset number of ECS inspections, the memory error information processing method provided by the present application also includes: when the baseboard management controller determines that the memory error information meets the preset repair conditions, the basic management controller sends a memory repair request to the central processing unit. The preset repair conditions include the existence of a target row of memory particles, and the number of times the target row with the same address appears reaches a fourth threshold. The target row is the row corresponding to the third count value in the second register. The address of the target row is determined based on the memory error information. The memory repair request is used to indicate the repair of the target row with the same address.

[0012] In this possible implementation, after the memory particle completes a preset number of ECS inspections, the baseboard management controller determines that there is a target row in the second register of the memory particle and the number of times the target row with the same address appears reaches a fourth threshold. In this case, the baseboard management controller sends a memory repair request to the central processing unit to instruct the repair of the target row with the same address. This can enable the central processing unit to identify and repair the target row with the same address that has caused errors multiple times as early as possible, thereby avoiding memory errors from evolving into serious faults.

[0013] In one possible implementation, the memory error information also includes address information of the row with the most single-bit errors in the memory cell. After the memory cell completes a preset number of ECS inspections, the memory error information processing method provided in this application further includes: if the baseboard management controller determines that the memory error information meets a preset alarm condition, the baseboard management controller generates an alarm. The preset alarm condition includes the first count value in the second register reaching a preset total error threshold and the target row not existing in the memory cell, or the target row existing and the number of occurrences of the target row with the same address less than a fourth threshold. The address of the target row is determined based on the memory error information. The alarm information indicates that the memory cell is in a sub-healthy state. In this possible implementation, after the memory cell completes a preset number of ECS inspections, if the baseboard management controller determines that the first count value in the second register of the memory error information reaches a preset total error threshold and the target row not existing in the memory cell, or the target row existing and the number of occurrences of the target row with the same address less than the fourth threshold, the baseboard management controller generates an alarm indicating that the memory cell is in a sub-healthy state. This can promptly notify the computing device that the number of errors in the memory error information has reached the preset total error threshold, thereby preventing serious malfunctions in the computing device.

[0014] In a second aspect, a memory error information processing device is provided, which is configured to execute any one of the memory error information processing methods provided in the first aspect.

[0015] In a possible implementation, the present application may divide the memory error information acquisition device into functional modules according to the method provided in the first aspect above. For example, each functional module may be divided according to each function, or two or more functions may be integrated into one processing module. Exemplarily, the present application may divide the memory error information acquisition device into a sending module, a reading module, and a generating module, etc., according to the function. The description of the possible technical solutions and beneficial effects executed by each of the functional modules divided above can refer to the technical solutions provided by the first aspect above or its corresponding possible implementation, and will not be repeated here.

[0016] In a third aspect, an embodiment of the present application provides a computing device, comprising a central processing unit (CPU), a memory module, and a baseboard management controller (BMC), wherein the memory module comprises memory chips; the CPU is electrically connected to the memory module and the BMC, respectively;

[0017] The central processing unit is used to send error inspection and correction ECS requests to the memory particles.

[0018] Memory particles are used to respond to ECS error inspection and correction requests, perform ECS inspection operations, and obtain memory error information.

[0019] The central processing unit is further configured to obtain memory error information from the memory particles and store the memory error information in the first register;

[0020] The baseboard management controller is configured to obtain memory error information from the first register.

[0021] In a possible implementation, the central processing unit is further configured to send a Patrol Scrub request to the memory module, so that the memory module performs a Patrol Scrub operation.

[0022] In one possible implementation, the central processing unit is further configured to reset the count value of the second register in the memory particle after the central processing unit enters a system interrupt. The central processing unit enters a system interrupt when the memory error information meets a preset condition, or after the central processing unit completes a Patrol Scrub inspection of the memory stick. The count value of the second register is used to indicate the number of times an error has occurred in the memory particle. In one possible implementation, the baseboard management controller is further configured to determine, before the central processing unit enters a system management interrupt, whether the memory error information meets a preset condition. When it is determined that the memory error information meets the preset condition, a system management interrupt is executed to the central processing unit.

[0023] In one possible implementation, the memory error information includes a count value of a second register in a memory cell. A preset condition includes the count value in the second register being at least one of a first count value, a second count value, or a third count value, and the count value does not include zero; wherein the first count value is the total number of single-bit errors that occur in the memory cell after the memory cell completes an ECS inspection, and the first count value is greater than a first threshold; the second count value is the total number of rows in the memory cell that have single-bit errors after the memory cell completes an ECS inspection, and the second count value is greater than a second threshold; and the third count value is the total number of single-bit errors in the row with the most single-bit errors in the memory cell after the memory cell completes an ECS inspection, and the third count value is greater than a third threshold.

[0024] The central processing unit is specifically configured to reset the count value after the central processing unit enters the system management interrupt.

[0025] In one possible implementation, the central processing unit is further configured to set a first preset period and a patrol mechanism. The ECS request includes the first preset period and the patrol mechanism. The patrol mechanism includes a patrol mode and an automatic patrol setting. The first preset period is used to instruct the memory cell to perform error patrol and error correction ECS according to the first patrol period. The automatic patrol setting is used to configure whether the ECS patrol of the memory cell is in automatic mode. The patrol mode includes a row mode and a codeword mode.

[0026] The memory particle is specifically configured to respond to an ECS request, perform an ECS inspection operation according to a first preset period and inspection mechanism, and obtain memory error information.

[0027] In a possible implementation, the central processing unit is specifically configured to periodically perform a patrol scrub operation on the memory bank according to a second preset period.

[0028] In a possible implementation, the baseboard management controller is specifically configured to periodically read the information of the first register according to a third preset period.

[0029] In a possible implementation, the memory cell is further configured to store the memory error information in a second register, and the memory cell includes the second register.

[0030] The central processing unit is specifically configured to read the memory error information by reading the information in the second register.

[0031] In a possible implementation, the baseboard management controller communicates with the central processing unit via a platform environment control interface (PECI).

[0032] In a fourth aspect, a baseboard management controller is provided. The baseboard management controller includes: an interface and a logic circuit, wherein the logic circuit is used to implement the memory error information processing method as described in the first aspect.

[0033] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which at least one computer program is stored. The computer program is loaded and executed by a processor to implement the memory error information processing method as described in the first aspect above.

[0034] In a sixth aspect, embodiments of the present application provide a computer program product or computer program, comprising computer instructions stored in a computer-readable storage medium. A processor reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing a terminal to perform the memory error information processing method provided in various optional implementations of the first aspect.

[0035] For the specific descriptions of the second to sixth aspects and their various implementations in this application, reference can be made to the detailed descriptions in the first aspect and its various implementations; and for the beneficial effects of the second to sixth aspects and their various implementations, reference can be made to the analysis of the beneficial effects in the first aspect and its various implementations, which will not be repeated here.

[0036] These and other aspects of the present application will become more readily apparent from the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] FIG1 is a structural diagram of the relationship between multiple physical granularities in a memory provided by an embodiment of the present application;

[0038] FIG2 is a schematic diagram of a memory failure according to an embodiment of the present application;

[0039] FIG3 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;

[0040] FIG4 is a schematic structural diagram of an ECS memory particle provided in an embodiment of the present application;

[0041] FIG5 is a specific flow chart of a method for processing memory error information provided by an embodiment of the present application;

[0042] FIG6 is a specific example diagram of a method for processing memory error information provided by an embodiment of the present application;

[0043] FIG7 is another specific flow chart of a method for processing memory error information provided by an embodiment of the present application;

[0044] FIG8 is another specific flow chart of a method for processing memory error information provided by an embodiment of the present application;

[0045] FIG9 is another specific flow chart of a memory error information processing method provided in an embodiment of the present application;

[0046] FIG10 is a schematic structural diagram of a memory error information acquisition device provided in an embodiment of the present application;

[0047] FIG11 is a schematic structural diagram of a baseboard management controller provided in an embodiment of the present application;

[0048] FIG12 is a schematic structural diagram of a chip system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] First, some concepts involved in a memory error information processing method and a computing device provided in an embodiment of the present application are explained.

[0050] Row error: A correctable error (CE) or uncorrectable error (UCE) that occurs in a row of memory cells. The physical granularity of memory is ranked from largest to smallest: Dimm, Rank, Device, Bank, Row / Column, Cell, and Bit. As shown in Figure 1, the relationship between these multiple physical granularities is as follows: each computing device may include multiple memory sticks (DIMMs), each of which has two memory ranks, located on the front and back sides of the memory, for example, the two memory ranks are memory rank 0 and memory rank 1. Each memory rank can be configured with multiple memory chips for storing data. Memory chips can also be called memory devices. Memory chips can be dynamic random access memory (DRAM), static random access memory (SRAM), etc. Each memory chip can be divided into multiple storage arrays (banks). In addition, multiple storage arrays can be grouped into a storage array group (bank group), where the number of storage arrays in each storage array group can be the same or different. A storage array is composed of a large number of data storage units (cells), which are arranged in a two-dimensional matrix. By specifying the row (row) and column (column) on the storage array, a data storage unit can be located on the storage array. The smallest unit of memory failure is a data storage unit on the storage array. Generally, in order to accurately identify memory failures, memory failures need to be located at the location of the failure on a memory chip. Failures on memory chips can be classified according to their severity, from most severe to least severe: storage array failure, row failure, column failure, and data storage unit failure.

[0051] For example, as shown in Figure 2, a in Figure 2a illustrates a memory array of a memory cell. A memory array failure refers to all failures occurring in a memory array of a memory cell. As shown in Figure 2b, a row failure refers to a failure occurring in a row of a memory array of a memory cell. As shown in Figure 2c, a column failure refers to a failure occurring in a column of a memory array of a memory cell. As shown in Figure 2d, a data storage unit failure refers to a failure occurring in a data storage unit of a memory array.

[0052] System Management Interrupt (SMI): When used, the central processing unit (CPU) enters the system management mode (SMM), and the CPU requires a memory area. Before entering SMM, the CPU will force the CPU to pause the executing program and store the data in the registers of the currently executing program in the memory area. The registers can be the registers of the CPU used to temporarily store the interrupted program. The CPU jumps to handle the interrupt event. After handling, it uses the RSM instruction to restore the data in the registers in the memory area to the CPU, so that the CPU jumps back to continue executing the originally interrupted program. The RSM instruction is used to fill the data in the memory area back to the CPU, so that the CPU can restore to the state before the system management interrupt and continue to execute the originally interrupted program.

[0053] Patrol Scrub: The integrated memory controller (IMC) periodically inspects memory. If errors are detected during the inspection, they are promptly corrected and written back. When the IMC completes the inspection, it triggers an SMI interrupt to report the results of the Patrol Scrub.

[0054] Error Check and Scrub (ECS): Memory modules that utilize the ECS mechanism are referred to as ECS memory modules, and the memory chips on an ECS memory module are known as ECS memory chips. ECS memory modules can include DDR5 memory modules or, as technology evolves, DDR6 memory modules. This application example uses a DDR5 memory module as an example.

[0055] DDR5 memory modules allow memory chips to read and correct single-bit errors internally, write the corrected data back to the memory chip's data storage unit, and record the number of errors corrected by the memory chip. Specifically, DDR5 memory modules can read the data in the memory chip's data storage unit through inspection and determine whether the data is erroneous. For each bit error that is corrected, the error count for the memory chip increases by 1.

[0056] Memory chips can include multiple data storage units, each of which can store at least one bit of data. A data storage unit is a physical unit (cell). If only one bit in a data storage unit flips (or only one bit in a data storage unit changes), it indicates that a single-bit error has occurred in the data storage unit. If more than one bit (such as two bits) in a data storage unit flips (or more than one bit in a data storage unit changes), it indicates that a multi-bit error has occurred in the data storage unit.

[0057] Memory sticks that use the ECC mechanism are called ECC memory sticks. Suppose that data 1 is stored in an ECC memory stick, and a 1-bit error occurs in data 1 (i.e., a single-bit error occurs in data 1). In this case, when the CPU reads data 1 from the ECC memory stick, the 1-bit error in data 1 is automatically corrected during the read process, making the data 1 read by the CPU correct. However, the 1-bit error in data 1 stored in the ECC memory stick is not corrected. If a 1-bit error occurs in a DDR5 memory stick, not only can the CPU ensure that the data 1 read is correct, but the 1-bit error in the DDR5 memory stick is also corrected, making the data 1 stored in the DDR5 memory stick also correct.

[0058] When DDR5 memory module failures are all single-bit errors, the CPU cannot detect the memory failure because the data it receives is corrected. Therefore, it cannot identify the distribution or type of memory failures. This delays timely diagnosis of memory failures and may affect service stability. Furthermore, the BMC and OS cannot obtain memory failure information and therefore cannot identify the distribution or type of memory failures. This delays timely diagnosis of memory failures and may affect service stability.

[0059] Memory register (MR): Deployed in memory particles and used to record error checking and correction information.

[0060] Platform environment control interface (PECI): A dedicated single-wire digital bus interface between the CPU and the baseboard management controller (BMC).

[0061] The solutions described in the embodiments of this application can be executed by a computing device, which can be a server. As shown in FIG3 , FIG3 is a schematic diagram of the structure of a computing device 300 provided in an embodiment of this application. The computing device 300 includes a central processing unit 301, a memory module 302, and a baseboard management controller 303.

[0062] The central processing unit 301 can run programs of one or more processing units. For example, the central processing unit 301 can run program code of a basic input / output system 3011, an integrated memory controller 3012, and an operating system 3013. The integrated memory controller 3012 may include registers 3014. To distinguish them from the registers in the memory module 302, the registers in the central processing unit 301 are referred to herein as first registers. The central processing unit 301 is in communication with the memory module 302. The central processing unit 301 can send requests to the memory module 302.

[0063] For example, the central processing unit 301 may be connected to the memory bank 302 via an I / O bus. The central processing unit 301 may be configured to send an ECS patrol request to the memory bank 302. The ECS patrol request instructs the memory chips on the memory bank 302 to perform error patrol and correction. The central processing unit 301 may also be configured to obtain memory error information from the slave memory bank 302 and store the memory error information in a first register of the memory controller 3012.

[0064] Memory stick 302 may include at least one memory chip, for example, memory chip 3021 and memory chip 3022, and the memory chips are connected by a link. The memory chips may be memory chips that apply the ECS mechanism, and memory chips that apply the ECS mechanism may be referred to as ECS memory chips. The ECS memory chips mentioned in the embodiments of the present application may be memory chips in an ECS memory stick, where the memory stick may be a DDR5 or DDR6 memory stick.

[0065] Exemplarily, the structural diagram of the ECS memory particle is shown in FIG4 . The ECS memory particle may include a data storage unit, which can be used to temporarily store the calculation data in the CPU and also to store the data exchanged between the CPU and an external memory such as a hard disk.

[0066] An ECS memory chip can include at least one memory register (MR) 3023. For example, as shown in Figure 4, an ECS memory chip can include seven MRs: MR14, MR15, MR16, MR17, MR18, MR19, and MR20. MRs can be used to record historical on-die error checking and correcting (ECC) information. On-die ECC is performed on the ECS memory chip to ensure data consistency within the ECS memory chip.

[0067] The MR in the ECS memory particle can at least count the number of errors that occur in the memory particle. Specifically, the MR in the ECS memory particle includes at least two counters, wherein the first counter can be an error counter (EC) and the second counter can be an errors per row counter (EpRC).

[0068] The inspection mode of ECS memory particles can be row mode (Row Mode) or codeword mode (CodeWord Mode).

[0069] When the ECS memory chip is inspected based on a row mode, the EC count value may be the total number of rows in which single-bit errors occur in the ECS memory chip.

[0070] When the ECS memory chip is inspected based on the codeword pattern, the EC count value may be the total number of single-bit errors occurring in the ECS memory chip.

[0071] Regardless of whether the ECS memory particles are inspected based on the row mode or the codeword mode, the EpRC count value is the total number of single-bit errors in the row with the most single-bit errors in the ECS memory particles.

[0072] To distinguish them from the registers in the CPU 301, the registers in the memory cell are referred to as second registers 3023. A memory cell can be divided into multiple storage arrays. When a memory cell stores data, the data is written bit by bit into one or more storage arrays. A storage array consists of a large number of data storage cells arranged in a two-dimensional matrix, forming rows and columns on the storage array. The address of a data storage cell can be represented by the row and column in the storage array where it is located.

[0073] For example, the memory bank 302 may be configured to respond to an ECS inspection request sent by the CPU 301, perform an ECS inspection operation, obtain memory error information, and store the memory error information in the second register. The memory bank 302 may be a DDR5 memory.

[0074] The baseboard management controller 303 can be used to monitor and manage components in a computing device. The computing device in the embodiment of the present application takes a server as an example. For example, the baseboard management controller 303 can monitor the working status (such as working temperature, working voltage, etc.) of various hardware devices in the server, such as memory bars, CPUs and other components. For another example, system configuration, firmware upgrades, fault diagnosis, etc. can be performed through the baseboard management controller 303. It should be noted that different equipment manufacturers may have different names for baseboard management controllers. For example, some baseboard management controllers are called fully automated integrated (integrated lights-out, iLO), and other baseboard management controllers are called integrated Dell remote access controller (integrated Dell remote access controller, iDRAC). Whether it is BMC, iLO or iDRAC, it can be understood as the baseboard management controller in the embodiment of the present invention. In the embodiment of the present application, the baseboard management controller is called BMC as an example.

[0075] In the related art, for the collection of error checking and error correction information, the central processing unit generally uses the Patrol Scrub inspection mechanism to perform periodic inspections on the DDR5, and after the inspection is completed, it triggers a system management interrupt to collect ECS information. Specifically, when the SMI interrupt is triggered, the basic input and output system (BIOS) of the CPU reads the register information of the integrated memory controller (IMC) to obtain ECS information and reset the ECS counter information, and reports the ECS information to the baseboard management controller. However, in this process, the BIOS reads the ECS information and resets the ECS counter information according to a fixed period, and the flexibility of obtaining ECS ​​information is relatively poor, which will lead to untimely information acquisition and untimely business migration, thereby causing business interruption. At the same time, the SMI interrupt will cause the system of the computing device to enter the system management (SMM) mode, causing the operating system (OS) of the computing device to hang, affecting the real-time performance of the computing device's business.

[0076] Based on this, the embodiment of the present application provides a memory error information processing method and computing device, the basic principle of which is: the baseboard management controller obtains memory error information from the first register of the central processing unit through the platform environment control interface PECI. When the memory error information meets the preset conditions, the baseboard management controller executes a system management interrupt to the central processing unit, so that the central processing unit resets the count value of the second register in the memory particle, wherein the count value of the second register is used to indicate the number of errors that have occurred in the memory particle. As can be seen, the memory error information processing method provided by the embodiment of the present application can execute a system management interrupt to the central processing unit when the baseboard management controller determines that the memory error information meets the preset conditions, so that the central processing unit resets the count value of the second register in the memory particle, and when the baseboard management controller determines that the memory error information does not meet the preset conditions, there is no need to instruct the central processing unit to enter the system management interrupt step. Compared with the prior art, the baseboard management controller will enter the system management interrupt process as soon as it receives the memory error information sent by the central processing unit, which reduces the number of times the central processing unit enters the system management interrupt when no errors have occurred in the dynamic random access memory, thereby improving the real-time performance of the computing device business.

[0077] Specifically, referring to Figure 3 , baseboard management controller 303 is connected to central processing unit 301 via a bus. Baseboard management controller 303 may include, but is not limited to, a memory error information collection module 3031 and a diagnostic module 3032. Memory error information collection module 3031 can be used to collect error information from memory chips, and diagnostic module 3032 can be used to diagnose memory error information.

[0078] For example, the baseboard management controller 303 and the central processing unit 301 may be connected via a PECI bus. The memory error information collection module 3031 of the baseboard management controller 303 may read information from the register 3014 via the PECI bus to obtain memory error information. The diagnosis module 3032 of the baseboard management controller 303 may obtain the memory error information read by the memory error information collection module 3031, perform diagnosis based on the obtained memory error information, and generate an alarm message or a memory repair request.

[0079] FIG5 is a flow chart of a method for processing memory error information provided by an embodiment of the present application. As shown in FIG5 , the method may include the following steps:

[0080] S501, the BMC obtains memory error information from the first register of the central processing unit through the platform environment control interface PECI, wherein the first register is the IMC register in the central processing unit and the second register is the MR register in the memory chip.

[0081] Specifically, the BMC reads the information of the first register via the PECI bus between the BMC and the CPU to obtain the memory error information.

[0082] The memory error information includes the count value of the second register.

[0083] In one possible implementation, after the memory cell completes an ECS inspection, if the total number of single-bit errors occurring in the memory cell is greater than a first threshold, the count value includes a first count value, and the first count value is the total number of single-bit errors occurring in the memory cell.

[0084] In another possible implementation, if the total number of rows in the memory cell where single-bit errors occur is greater than a second threshold, the count value includes a second count value, and the second count value is the total number of rows in the memory cell where single-bit errors occur.

[0085] In another possible implementation, if the total number of single-bit errors in the row with the most single-bit errors in the memory cell is greater than a third threshold, the count value includes a third count value, and the third count value is the total number of single-bit errors in the row with the most single-bit errors in the memory cell.

[0086] The count value of the second register is determined according to a patrol mechanism of error detection and error correction ECS request. The patrol mechanism includes a patrol mode. The patrol mode includes but is not limited to a codeword mode and a row mode.

[0087] In a possible implementation, when the inspection mode is the codeword mode, the count value of the second register includes a first count value and a third count value.

[0088] Exemplarily, when the inspection mode is the codeword mode, the first threshold value of the first count value is 8, and the third threshold value of the third count value is 4. As shown in FIG6 , the total number of single-bit errors occurring in the memory cells is 10, and the total number of single-bit errors occurring in the row with the most single-bit errors is 4. In this case, the total number of single-bit errors occurring in the memory cells is greater than the first threshold value, and the first count value is 10. The total number of single-bit errors occurring in the row with the most single-bit errors is not greater than the third threshold value, and the third count value is 0.

[0089] In another possible implementation, the patrol mode is a row mode, and the count value of the second register includes a second count value and a third count value.

[0090] Exemplarily, when the inspection mode is row mode, the second threshold value of the second count value is 4, and the third threshold value of the third count value is 4. As shown in FIG6 , the total number of rows in which single-bit errors occur in the memory cell is 4, and the total number of single-bit errors in the row with the most single-bit errors in the memory cell is 4. In this case, the total number of single-bit errors in the memory cell is not greater than the second threshold value, and the second count value is counted as 0. The total number of single-bit errors in the row with the most single-bit errors is not greater than the third threshold value, and the third count value is counted as 0.

[0091] S502: When the memory error information meets a preset condition, the baseboard management controller executes a system management interrupt to the central processing unit, so that the central processing unit resets the count value of the second register in the memory chip.

[0092] The count value of the second register is used to indicate the number of times errors occur in the memory cell.

[0093] Specifically, the BMC determines whether the memory error information satisfies a preset condition, and if the BMC determines that the memory error information satisfies the preset condition, executes a system management interrupt to the CPU. The preset condition includes that the count value of the second register is at least one of the first count value, the second count value, or the third count value, and the count value does not include zero.

[0094] For example, when the current inspection mode is the codeword mode, when the BMC determines that the count value of the second register is the first count value, it indicates that an error exists in the memory chip and has been detected and corrected. At this time, the BMC executes a system management interrupt to the CPU.

[0095] Alternatively, when the current patrol mode is the row mode, when the BMC determines that the count value of the second register is the second count value, the BMC executes a system management interrupt to the CPU.

[0096] The BMC's execution of a system management interrupt to the CPU includes, but is not limited to, direct triggering of an SMI interrupt using hardware signal pins and software triggering of an SMI interrupt. Direct triggering of an SMI interrupt using hardware signal pins involves the BMC controlling the level of a GPIO pin connected to the CPU to trigger an SMI interrupt, for example, by pulling the level of the GPIO pin from high to low or from low to high. Software triggering of an SMI interrupt involves the BMC writing a value to port 0xB2 of the CPU to trigger an SMI interrupt.

[0097] Specifically, after the CPU enters the system management interrupt, the CPU resets the count value of the second register in the memory particle.

[0098] MR14 stores a reset bit for indicating whether to perform a clearing operation on the ECS counter.

[0099] For example, after the computer system enters the SMM mode, the CPU configures MR14 and sets the reset bit of MR14 to 1 to clear the ECS counter in the MR register.

[0100] In the embodiment of the present application, when the BMC determines that the memory error information does not meet the preset conditions, the CPU executes a system management interrupt to the CPU. After the CPU enters the system management interrupt, the CPU resets the count value of the second register in the memory particle. When the BMC determines that the memory error information does not meet the preset conditions, it is not necessary to instruct the CPU to enter the system management interrupt. Compared with the related art, the CPU performs an ECS inspection on the memory and enters the system management interrupt process after the inspection is completed. This reduces the number of times the CPU enters the system management interrupt SMI when the number of errors in the memory particle is less than the first threshold, the second threshold, or the third threshold. In addition, after the CPU enters the system management interrupt SMI, the embodiment of the present application only performs the clearing action of the counter of the register of the memory particle, and the CPU can resume processing normal business, with a short interruption time. In the related art, after the inspection is completed, the CPU enters the system management interrupt SMI, and the CPU needs to read the register information of the integrated memory controller IMC to obtain the memory error information and perform the clearing action of the counter of the register of the memory particle before the CPU can jump out of the interrupt to process normal business, and the CPU interruption time is longer. Therefore, the embodiment of the present application can improve the reliability and stability of the computing device business.

[0101] In some embodiments, as shown in Figure 7, the memory error information processing method provided by the embodiment of the present application may further include: S701 to S707. Exemplarily, S701 to S707 may be executed before S501.

[0102] S701: The CPU sets a first preset period, an inspection mechanism, and a second preset period.

[0103] The first preset period and patrol mechanism are used to configure the memory chip to perform ECS patrol operations according to the first preset period and patrol mechanism, and the second period is used to configure the CPU to perform Patrol Scrub patrol operations according to the second period. It should be noted that completing the ECS patrol operation of the memory chip according to the first preset period means that the memory chip completes one ECS patrol operation.

[0104] S702: The CPU performs a Patrol Scrub operation according to a second preset period.

[0105] When executing the Patrol Scrub operation, the CPU may perform error patrol on the memory cells on the entire memory bank and patrol for link failures between the memory cells. The patrol of the memory cells may include patrol for multi-bit errors of the memory cells.

[0106] The second preset period may be greater than the first preset period, so that when the CPU completes the Patrol Scrub inspection on the memory stick, the memory chips on the memory stick have also completed the ECS inspection.

[0107] S703: The CPU sends an ECS inspection request to the memory bank.

[0108] Among them, the ECS patrol request includes: a first preset cycle and a patrol mechanism. The CPU sending an ECS patrol request to the memory stick includes the CPU sending an ECS patrol request to the memory particles on the memory stick. The first preset cycle is used to instruct the memory particles to perform ECS patrol according to the first patrol cycle. The patrol mechanism includes a patrol mode and an automatic patrol setting. The automatic patrol setting is used to configure whether the ECS patrol of the memory particles is in automatic mode, so as to indicate whether the memory particles automatically start the next round of error patrol and error correction ECS after the error patrol and error correction ECS. The patrol mode includes but is not limited to the row mode and the codeword mode. Among them, the row mode indicates that error patrol and error correction ECS are performed on each row in the memory particle, and the codeword mode indicates that error patrol and error correction ECS are performed on each data storage unit in the memory particle.

[0109] For example, in the automatic inspection mode, the memory chip automatically executes the next round of error inspection and correction after completing the error inspection and correction ECS. In the non-automatic inspection mode, the memory chip does not execute the next round of error inspection and correction ECS after executing the error inspection and correction ECS. It waits for the next ECS inspection request to be received and then executes the next round of error inspection and correction ECS according to the received ECS inspection request.

[0110] In some possible implementations, the CPU may periodically send an ECS inspection request to the memory chips on the memory module according to a preset period, where the ECS request includes the first preset period and inspection mechanism most recently set by the CPU. Alternatively, when one or more of the first preset period and inspection mechanism need to be updated, such as after the CPU has reconfigured one or more of the first preset period and inspection mechanism, the CPU may send an ECS inspection request to the memory chips on the memory module, where the ECS inspection request includes the first preset period and inspection mechanism most recently set by the CPU.

[0111] S704 : The memory chips on the memory bank respond to the ECS inspection request, perform an ECS inspection operation, and obtain memory error information.

[0112] Specifically, after receiving the ECS inspection request, the memory chip performs an ECS inspection operation according to the first preset period and inspection mechanism configured in the ECS request, and obtains memory error information.

[0113] For example, an ECS inspection request includes a first preset period of t1, an automatic inspection mode in the inspection mechanism of automatic mode, and an inspection mode of row mode. After receiving the ECS inspection request, the memory chip performs ECS inspection operations according to the t1 period, automatically starts the next round of error inspection and correction ECS after each error inspection and correction ECS, and performs error inspection and correction ECS according to the row mode.

[0114] S705: The memory chip stores the memory error information in the second register.

[0115] Specifically, after obtaining the memory error information, the memory cell stores the memory error information in a second register of the memory cell. The second register may be an MR register of the memory cell.

[0116] The MR registers of the memory particles may include but are not limited to MR14, MR15, MR16, MR17, MR18, MR19, and MR20. It should be noted that MR14, MR15, MR16, MR17, MR18, MR19, and MR20 are only identifiers of the MR registers, and do not limit the functions of the MR registers. Among them, MR14 stores the inspection mode; MR15 stores the inspection mechanism; MR16, MR17, and MR18 store the address information of all rows in the memory particles, including the row address information with the largest number of errors; MR19 stores the total number of rows with errors; MR20 stores the total number of errors in the memory particles, the target row, and the number of errors in the target row. The target row is the row corresponding to the EpRC count value in the second register exceeding the third threshold.

[0117] S706: The CPU obtains memory error information from the second register and stores the memory error information in the first register.

[0118] Specifically, the CPU reads the information in the second register of the memory particle to obtain the memory error information, and stores the memory error information in the first register of the CPU.

[0119] S707: The BMC sends a first request to the CPU.

[0120] The first request is used to request to read the information in the first register of the CPU. The information in the first register includes memory error information obtained when the ECS memory particles perform error checking and error correction ECS patrol and / or memory error information obtained when the CPU performs patrol scrub on the ECS memory bar. The memory error information may include but is not limited to information in the second register of the memory particles, wherein the information in the second register includes the values ​​of the EC counter and the EpRC counter in the second register and the error row address, and the error row address includes the row address of the row where the most single-bit errors occur in the memory particles.

[0121] Specifically, the BMC first configures a third preset period of the BMC to instruct the BMC to periodically send the first request to the CPU according to the third preset period.

[0122] In one possible implementation, the BMC periodically sends a first request to the CPU via the PECI bus between the BMC and the CPU according to a third preset period. The first request includes a read instruction and the address of the first register. The address of the first register is used by the BMC to find the first register. After receiving the read instruction and the address of the first register, the CPU indicates that the BMC needs to read the information in the first register at this time, and the CPU stores the information in the first register in the shared memory in the CPU. The shared memory is a memory that can be accessed by the BMC and the CPU. After receiving the read instruction and the address of the first register, the CPU can store the information in the first register in the shared memory, so that the BMC can quickly read the information in the first register, thereby realizing rapid information transmission between the CPU and the BMC through the shared memory.

[0123] In a possible implementation, the BMC may read information in the first register from the first register based on the address of the first register.

[0124] In another possible implementation, the BMC periodically sends the first request to the CPU via the PECI bus between the BMC and the CPU according to a third preset period.

[0125] In some embodiments, as shown in Figure 7, the memory error information processing method provided by the embodiment of the present application may further include: S708 and S709. Exemplarily, S708 and S709 may be executed after S502.

[0126] S708: The CPU completes the Patrol Scrub operation and executes a system management interrupt.

[0127] Specifically, after completing the Patrol Scrub operation, the CPU executes a system management interrupt.

[0128] S709: After the CPU enters the system management interrupt, the CPU resets the counter of the first register in the memory particle.

[0129] Specifically, after the CPU enters the system management interrupt, the CPU resets the value of the counter of the first register in the memory particle.

[0130] MR14 also stores a reset bit for indicating whether to perform a clearing operation on the ECS counter.

[0131] For example, after the computer system enters SMM mode, the CPU configures MR14 and sets the reset bit of MR14 to 1 to clear the ECS counter in the MR register. The ECS error information includes the address of the row with the largest number of errors recorded in MR16, MR17, and MR18; the ECS counter includes the total number of rows with errors stored in MR19 and the total number of errors in the memory cell, the target row, and the number of errors in the target row stored in MR20.

[0132] This embodiment of the present application triggers the system management interrupt of the CPU at the completion time of the CPU performing Patrol Scrub on the memory bar. The CPU's own patrol mechanism can be used to complete the reset of the counter of the first register of the memory particle, saving BMC resources.

[0133] In the embodiment of the present application, after the CPU enters the system management interrupt, the CPU only needs to reset the memory error information in the memory particles. Compared with the related art, when the Patrol Scrub inspection operation is completed, the CPU needs to read the Patrol Scrub inspection information and ECS information from the second register of the memory particles, and also needs to reset the memory error information in the memory particles and report the Patrol Scrub inspection information. The embodiment of the present application saves the time for the CPU to collect the Patrol Scrub inspection information, that is, shortens the duration of the SMI interrupt, and improves the real-time performance of the computing device business.

[0134] In some embodiments, as shown in FIG8 , a memory error information processing method provided by an embodiment of the present application may further include: S801 to S806. Exemplarily, S801 to S806 may be executed before S501.

[0135] S801: The CPU sets a first preset cycle and a patrol mechanism.

[0136] It should be noted that the process of the CPU setting the first preset period and the inspection mechanism is the same as the process of the CPU setting the first preset period and the inspection mechanism in the above S701, and this application will not elaborate on this.

[0137] S802: The CPU sends an ECS inspection request to the memory chips on the memory bank.

[0138] It should be noted that the process of the CPU sending the ECS inspection request to the memory stick is the same as the process of the CPU sending the ECS inspection request to the memory stick in the above S703, and this application will not elaborate on this.

[0139] S803: The memory chip responds to the ECS inspection request, performs an ECS inspection operation, and obtains memory error information.

[0140] It should be noted that the memory particles respond to the ECS inspection request, perform the ECS inspection operation, and obtain the memory error information. The process is the same as the memory particles respond to the ECS inspection request, perform the ECS inspection operation, and obtain the memory error information in the above S704. This application will not elaborate on this.

[0141] S804: The memory chip stores the memory error information in the second register.

[0142] It should be noted that the process of the memory chip storing the memory error information in the second register is the same as the process of the memory chip storing the memory error information in the second register in S705 above, and this application will not elaborate on this.

[0143] S805: The CPU obtains memory error information from the second register and stores the memory error information in the first register.

[0144] It should be noted that the CPU obtains memory error information from the memory particles and stores the memory error information in the first register, which is the same as the process in S706 above in which the CPU obtains memory error information from the memory particles and stores the memory error information in the first register. This application will not elaborate on this.

[0145] S806: The BMC sends a first request to the CPU.

[0146] It should be noted that the process of the BMC sending the first request to the CPU is the same as the process of the BMC sending the first request to the CPU in S707 above, which is not described in detail in this application.

[0147] In some embodiments, as shown in FIG9 , a memory error information processing method provided by an embodiment of the present application may further include: S901 or S902. Exemplarily, S901 or S902 may be executed after S708 or S807 is executed a preset number of times. For example, after the BMC executes S708 or S807 a preset number of times (n times), it indicates that the memory particles have completed the preset number of ECS inspections and obtained the preset number of memory error messages, and then S901 or S902 is executed.

[0148] S901: When the BMC determines that memory error information meets a preset repair condition, the BMC sends a memory repair request to the CPU.

[0149] The preset repair conditions include: the target row in S705 exists in the memory cell, and the number of occurrences of the target row at the same address reaches a fourth threshold; the memory error information also includes the address information of the row in the memory cell with the most single-bit errors; the address of the target row is determined based on the memory error information; and the memory repair request is used to indicate the repair of the target row.

[0150] Specifically, after the BMC executes S707 or S806 for a preset number of times and obtains memory error information for a preset number of times, if the number of occurrences of the target row with the same address reaches a fourth threshold and the address of the target row is the same, the BMC sends a memory repair request to the CPU.

[0151] For example, each time the ECS completes an inspection, the BMC can obtain memory error information in the first register once, and in each ECS inspection, if the target row appears, the BMC extracts the row address of the target row each time; then after the memory particle completes the preset 10 ECS inspections, the BMC saves the memory error information in the first register 10 times. At this time, the BMC can determine that the number of times the target row with the same address it has saved appears has reached a preset fourth threshold. For example, the fourth threshold can be 5.

[0152] Specifically, after the memory chip completes the first ECS inspection and determines that the 5th row of the first storage array of memory chip 1 is the target row, the BMC obtains the ECS inspection information of this time, including the row address of the target row, and saves it; if the memory chip determines that the 5th row of the first storage array of memory chip 1 is the target row five times from all its saved ECS inspection information after completing the 10th inspection, the BMC sends a memory repair request to the CPU. After receiving the memory repair request, the CPU repairs the error in the 5th row.

[0153] Because each storage cell in the memory array of a memory chip may have a single-bit error, it is common for a row to be the target row during a single ECS inspection. However, if the row is the target row during multiple ECS inspections, it indicates that the row is a faulty row and needs to be repaired.

[0154] S902: When the BMC determines that the memory error information meets a preset alarm condition, the BMC generates an alarm message.

[0155] The preset alarm conditions include: the count value of the EC counter in the second register reaches a preset total error threshold and the target row does not exist in the memory chip, or the number of occurrences of the target row at the same address is less than a fourth threshold. The alarm information is used to indicate that the memory chip is in a sub-healthy state.

[0156] Specifically, after the BMC executes S708 or S807 a preset number of times and receives a preset number of memory error messages, if the target row does not exist, or if the target row exists but the number of occurrences of the target row at the same address is less than a fourth threshold, the BMC generates an alarm message to remind the user that the memory is in a sub-healthy state. This sub-healthy state means that the memory chips are still operating normally, but there are many potential errors. For example, the memory connection is normal, but some chips have errors that affect performance, etc. This application does not limit the fault type of the sub-healthy state.

[0157] The above describes the solution provided by the embodiments of this application from a methodological perspective. The above-mentioned memory error information processing method can be applied to the above-mentioned computer system. In terms of hardware implementation, the memory error information acquisition system can be implemented as a chip deployed in the CPU. In terms of software, the computer system can include an operating system and a BIOS that communicates with the OS, wherein the OS is used to execute the memory error information processing method of the above-mentioned embodiment.

[0158] The above mainly introduces the solutions provided by the embodiments of the present application from the perspective of methods and systems. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0159] The embodiment of the present application can divide the memory error information acquisition device into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0160] Figure 10 shows a schematic diagram of the structure of a memory error information acquisition device 1000 provided in an embodiment of the present application. This memory error information acquisition device is used to execute the aforementioned memory error information processing method and can be applied to a computing device. For example, the memory error information processing method shown in any of Figures 5-9 can be executed. Exemplarily, the memory error information acquisition device 1000 may include a sending module 1001, a reading module 1002, and a generating module 1003.

[0161] The sending module 1001 is configured to send a first request to the central processing unit.

[0162] In a possible implementation, the sending module 1001 may also be configured to trigger a system management interrupt to the central processing unit when determining that the memory error information meets a preset reset condition.

[0163] In another possible implementation, the sending module 1001 may also be configured to send a memory repair request to the CPU when determining that the memory error information meets a preset repair condition.

[0164] For example, in conjunction with FIG5 , the sending module 1001 may be configured to execute S501 .

[0165] The reading module 1002 is configured to read the information in the first register to obtain memory error information.

[0166] For example, in conjunction with FIG5 , the reading module 1002 may be configured to execute S502 .

[0167] The generating module 1003 is configured to generate an alarm message when it is determined that the memory error message meets a preset alarm condition.

[0168] For the detailed description of the above optional methods, please refer to the above method embodiments, which will not be repeated here. In addition, the explanation and beneficial effects of any of the above memory error information acquisition devices can be referred to the above corresponding method embodiments, which will not be repeated here.

[0169] The solution shown in the embodiment of the present application can be executed by a baseboard management controller 1100. As shown in FIG11 , the baseboard management controller 1100 may include an interface 1101 and a logic circuit 1102.

[0170] The interface 1101 is used to support the baseboard management controller 303 to communicate with other hardware, for example, to support the baseboard management controller to communicate with the central processing unit.

[0171] Logic circuit 1102 is configured to execute the memory error information processing method provided in an embodiment of the present application, by sending a read instruction and address information of a first register to a central processing unit (CPU), thereby reading information from the first register and obtaining memory error information. The CPU includes the first register. The read instruction is configured to request reading information from the first register.

[0172] In an exemplary embodiment, a computer-readable storage medium is further provided, configured to store at least one instruction, at least one program, a code set, or an instruction set, wherein the at least one instruction, at least one program, the code set, or the instruction set is loaded and executed by a central processing unit to implement all or part of the steps in the above-mentioned memory error information processing method. For example, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0173] In an exemplary embodiment, a computer program product or computer program is also provided. The computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computing device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computing device to perform all or part of the steps of the method shown in any of the embodiments of Figures 5-9 above.

[0174] In some embodiments, the methods shown in the embodiments of the present application may be implemented as computer program instructions encoded in a machine-readable format on a computer-readable storage medium or encoded on other non-transitory media or products.

[0175] An embodiment of the present application further provides a chip system 1200 , as shown in FIG12 . The chip system 1200 includes at least one processor 1201 and at least one interface circuit 1202 .

[0176] As an example, when the chip system 1200 includes one processor and one interface circuit, the one processor may be the processor 1201 shown in the solid-line box in FIG12 (or the processor 1201 shown in the dotted-line box), and the one interface circuit may be the interface circuit 1202 shown in the solid-line box in FIG12 (or the interface circuit 1202 shown in the dotted-line box). When the chip system 1200 includes two processors and two interface circuits, the two processors include the processor 1201 shown in the solid-line box and the processor 1201 shown in the dotted-line box in FIG12, and the two interface circuits include the interface circuit 1202 shown in the solid-line box and the interface circuit 1202 shown in the dotted-line box in FIG12. This is not limited to this.

[0177] The processor 1201 and the interface circuit 1202 can be interconnected via lines. For example, the interface circuit 1202 can be used to receive signals. For another example, the interface circuit 1202 can be used to send signals to other devices (such as the processor 1201). For example, the interface circuit 1202 can read the computer instructions stored in the memory and send the computer instructions to the processor 1201. The processor 1201 executes the instruction and, in combination with the input and output devices, implements the various steps in the above-mentioned embodiments, such as implementing the various steps performed in the method embodiments shown in any of Figures 3 to 5. Of course, the chip system may also include other discrete devices, which is not specifically limited in the embodiments of the present application.

[0178] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0179] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0180] The units described as separate components may or may not be physically separated, and the components displayed as units may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, the functional units in the various embodiments of the present application may be integrated into a processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0181] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A memory error information processing method, characterized in that: Applied to a baseboard management controller, the method comprises: The baseboard management controller obtains memory error information from the first register of the central processing unit through the platform environment control interface PECI, wherein the memory error information is memory error information of the memory particle completing an error check and error correction ECS inspection according to the first preset period; When the memory error information meets a preset condition, the baseboard management controller executes a system management interrupt to the central processing unit so that the central processing unit resets the count value of the second register in the memory particle; wherein the count value of the second register is used to indicate the number of times an error occurs in the memory particle.

2. The memory error information processing method according to claim 1, characterized in that: The memory error information includes the count value of the second register; the preset condition includes that the count value in the second register is at least one of the first count value, the second count value or the third count value, and the count value does not include zero; wherein the first count value is the total number of single-bit errors occurring in the memory particle after the memory particle completes an ECS inspection, and the first count value is greater than a first threshold; the second count value is the total number of rows in which single-bit errors occur in the memory particle after the memory particle completes an ECS inspection, and the second count value is greater than a second threshold; the third count value is the total number of single-bit errors in the row with the most single-bit errors occurring in the memory particle after the memory particle completes an ECS inspection, and the third count value is greater than a third threshold.

3. The memory error information processing method according to claim 2, characterized in that: The memory error information also includes address information of the row in which the most single-bit errors occur in the memory cell; After the memory chip completes the ECS inspection for a preset number of times, the method further includes: When the baseboard management controller determines that the memory error information satisfies a preset repair condition, the baseboard management controller sends a memory repair request to the central processing unit; the preset repair condition includes that a target row exists in the memory particle, and the number of occurrences of the target row with the same address reaches a fourth threshold; the target row is the row corresponding to the third count value in the second register; the address of the target row is determined according to the memory error information; and the memory repair request is used to indicate the repair of the target row with the same address.

4. The memory error information processing method according to claim 2, characterized in that: The memory error information also includes address information of the row in which the most single-bit errors occur in the memory cell; After the memory chip completes the ECS inspection for a preset number of times, the method further includes: In the case where the baseboard management controller determines that the memory error information meets a preset alarm condition, the baseboard controller generates alarm information; The preset alarm conditions include: the first count value in the second register reaches a preset total error threshold and the target row does not exist in the memory particle; or the target row exists and the number of occurrences of the target row with the same address is less than a fourth threshold; the address of the target row is determined based on the memory error information; and the alarm information is used to indicate that the memory particle is in a sub-healthy state.

5. A computing device, comprising a central processing unit, a memory bar and a baseboard management controller; the memory bar comprises memory particles; the central processing unit comprises a first register, and the central processing unit is electrically connected to the memory bar and the baseboard management controller respectively, characterized in that: The central processing unit is used to send an error inspection and error correction ECS request to the memory particle; The memory particle is used to respond to the error inspection and correction ECS request, perform an ECS inspection operation, and obtain memory error information; The central processing unit is further used to obtain the memory error information from the memory particle and store the memory error information in a first register; The baseboard management controller is used to obtain the memory error information from the first register.

6. The computing device according to claim 5, characterized in that The central processor is further used to send a Patrol Scrub request to the memory bar so that the memory bar performs a Patrol Scrub operation.

7. The computing device according to claim 6, characterized in that The central processing unit is also used to reset the count value of the second register in the memory particle after the central processing unit enters a system interrupt; wherein the central processing unit enters a system interrupt when the memory error information meets a preset condition or after the central processing unit performs a Patrol Scrub on the memory bar; the count value of the second register is used to indicate the number of times an error occurs in the memory particle.

8. The computing device according to claim 7, characterized in that The baseboard management controller is further used to determine that the memory error information meets a preset condition before the central processing unit enters a system management interrupt; when it is determined that the memory error information meets the preset condition, execute a system management interrupt to the central processing unit.

9. The computing device according to any one of claims 7 to 8, characterized in that: The memory error information includes the count value of the second register in the memory particle; the preset condition includes that the count value in the second register is at least one of the first count value, the second count value or the third count value, and the count value does not include zero; wherein the first count value is the total number of single-bit errors occurring in the memory particle after the memory particle completes an ECS inspection, and the first count value is greater than a first threshold; the second count value is the total number of rows in which single-bit errors occur in the memory particle after the memory particle completes an ECS inspection, and the second count value is greater than a second threshold; the third count value is the total number of single-bit errors in the row with the most single-bit errors in the memory particle after the memory particle completes an ECS inspection, and the third count value is greater than a third threshold; The central processing unit is specifically used to reset the count value after the central processing unit enters a system management interrupt.

10. The computing device according to any one of claims 5 to 9, characterized in that: The central processor is also used to set a first preset period and a patrol mechanism; The ECS request includes the first preset period and the inspection mechanism; the inspection mechanism includes an inspection mode and an automatic inspection setting; The first preset cycle is used to instruct the memory particle to perform error inspection and error correction ECS according to the first inspection cycle; the automatic inspection setting is used to configure whether the ECS inspection of the memory particle is in automatic mode; the inspection mode includes a row mode and a codeword mode; The memory particle is specifically used to respond to the ECS request, perform an ECS inspection operation according to a first preset cycle and inspection mechanism, and obtain memory error information.

Citation Information

Patent Citations

  • Method and system for supervising DDR5 memory particle errors, storage medium and equipment

    CN115543678A

  • Adaptive error correction to improve system memory reliability, availability, and serviceability (RAS)

    CN116783654A

  • Memory error information processing method and computing device

    CN117909109A

  • Method and unit for handling interrupts in a system

    US20170329730A1