Computer system, memory exception processing method and storage medium
By completing memory exception logging on the operating system side, the processor recovery time delay caused by memory exception handling in traditional BIOS mode is solved, and faster system response is achieved.
Patent Information
- Application Number
- CN202311587135.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-05-27
AI Technical Summary
The traditional memory exception handling method causes the processor to recover from BIOS mode to operating system mode to be large, mainly because the BIOS needs to wait for the BMC to complete the IO operation during the process of recording memory exception logs.
The traditional BIOS's memory exception logging work is moved to the operating system (OS) side for completion. The correctable error (CE) count of memory is read through the target driver running on the OS, and the target program running on the OS clears the CE count in time when the CE count reaches less than the set CE number threshold, thereby avoiding triggering system management interrupts (SMI) to enter BIOS mode.
Reduces the time delay of the processor to resume operating system mode, reduces the overhead of SMI interrupts, and improves system response speed.
Smart Images

Figure CN120045363A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a computer system, a method for handling memory exceptions, and a storage medium. Background Art
[0002] Memory is one of the important components of a computing device. Memory exceptions are the most common exceptions in a hardware system, which greatly affect the reliability, availability, and serviceability (RAS) of the system. Correctable errors (CEs) in memory are one of the most common types of memory exceptions.
[0003] In traditional solutions, the number of CE exceptions in memory can be counted. When the number of CE exceptions reaches a preset threshold, the memory controller triggers a system management interrupt (SMI), and the computing device enters the basic input / output system (BIOS) for memory exception handling. This memory exception handling method causes a large time delay for the processor of the computing device to resume from the BIOS mode to the operating system mode. Summary of the Invention
[0004] Multiple aspects of this application provide a computer system, a method for handling memory exceptions, and a storage medium, so as to reduce the time delay for the processor to resume to the operating system mode.
[0005] An embodiment of this application provides a computer system, including: a processor and a memory; the processor includes: a target counter for counting correctable errors (CEs) in the memory; the processor runs an operating system (OS).
[0006] The OS runs a target driver, which is configured to, in response to an interrupt triggered by a CE occurring in the memory, obtain the number of CEs in the memory from the target counter; and in the case where the number of CEs in the memory reaches a set target number, send the number of CEs in the memory to a target program running on the OS; the target number is less than a set number threshold; the number threshold is the number threshold of correctable errors that triggers the processor to enter the basic input / output unit to record a memory exception log.
[0007] The target program running on the OS is used to obtain the number of CEs of the memory, and control the target driver to clear the CE count of the target counter; calculate the total number of CEs of the memory according to the number of CEs of the memory and the historical number of CEs of the memory; and record a memory exception log when the total number of CEs of the memory is greater than or equal to the CE quantity threshold.
[0008] An embodiment of the present application further provides a memory exception handling method, which is applicable to a processor on which an operating system OS runs; the processor includes: a target counter for counting correctable errors (CEs) of the memory; the method includes:
[0009] The target driver running in the OS responds to an interrupt triggered by a CE occurring in the memory, and obtains the number of CEs of the memory from the target counter;
[0010] When the number of CEs of the memory reaches a set target number, send the number of CEs of the memory to the target program running in the OS;
[0011] The target program obtains the number of CEs of the memory and clears the CE count of the target counter; the CE quantity threshold is greater than the target number; the target number is less than a set quantity threshold; the quantity threshold is the quantity threshold of correctable errors that trigger the processor to enter the basic input / output unit to record a memory exception log;
[0012] Calculate the total number of CEs of the memory according to the number of CEs of the memory and the historical number of CEs of the memory;
[0013] Record a memory exception log when the total number of CEs of the memory is greater than or equal to the CE quantity threshold.
[0014] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which when executed by one or more processors, cause the one or more processors to execute the steps in the above memory exception handling method.
[0015] In the embodiment of the present application, the work of the traditional BIOS to record memory exception logs is moved to the OS side to be completed. The target driver running on the OS reads the CE count of the memory, and when the CE count reaches a target number less than the set CE quantity threshold through the target program running on the OS, the CE count is cleared in a timely manner, so that the CE count of the target counter is always less than the set CE quantity threshold, which can avoid triggering the SMI when the target counter reaches the set CE quantity threshold, and thus can avoid the processor from entering the BIOS. The target program running on the OS in this embodiment can also record memory exception logs when the total CE quantity of the memory is greater than or equal to the CE quantity threshold. In short, the work of memory exception logs in this embodiment is completed in a closed loop by the OS without the participation of the BIOS, and there is no need to switch the processor to the BIOS mode, thereby reducing the time delay for the processor to resume the OS mode. Description of the Drawings
[0016] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0017] Figure 1 is a schematic structural diagram of a computer system provided by an embodiment of the present application;
[0018] Figure 2 is an internal structural schematic diagram of a memory provided by an embodiment of the present application;
[0019] Figure 3 is a schematic diagram of a memory exception handling process provided by an embodiment of the present application;
[0020] Figure 4 is a schematic flowchart of a memory exception handling method provided by an embodiment of the present application. Detailed Embodiments
[0021] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0022] Memory exceptions of computing devices can generally be divided into soft exceptions and hard exceptions. Hard exceptions generally refer to unrecoverable hardware errors, such as stuck-at bit, that is, the bit is always a fixed value. For example, if it is stuck-at 0, it means that no matter what value is written (0 or 1), the value of the bit is always 0. Soft exceptions are caused by random events, such as radiation or backplane radiation, and soft exceptions can usually be repaired by rewriting.
[0023] One of the popular RAS schemes in memory systems is to use the Error Checking and Correction (ECC) algorithm to repair memory data. The ECC error correction algorithm generates single-bit error correction and double-bit error detection (SECDED) codes for the actual memory data and stores the SECDED codes in the memory. The memory controller can correct single-bit errors and detect double-bit errors using SECDED. Single-bit errors are called correctable errors (CE) and double-bit errors are called uncorrectable errors (UE). Memory UE is the main cause of system downtime.
[0024] Memory CE can be used to predict UE or measure the health of memory. If a large number of CE abnormalities occur in the memory bar of a computing device, it indicates that the memory bar is aged and UE is likely to occur, causing the system to crash.
[0025] In traditional solutions, the number of CE exceptions in the memory can be counted. When the number of CE exceptions reaches a preset threshold, the memory controller triggers a system management interrupt (SMI), and the computing device enters the basic input / output system (BIOS) to handle the memory exception. Among them, BIOS is a set of programs fixed to the read-only memory (ROM) on the motherboard of the computing device. It stores the most important basic input and output programs of the computing device, the self-test program after power-on, and the system self-starting program. BIOS belongs to the firmware of the computing device, and generally does not support modification of the firmware.
[0026] Among them, the BIOS performs memory exception handling mainly by recording memory exception logs, including but not limited to: (1) The BIOS writes memory exception events into the System Event Log (SEL) of the Baseboard Management Controller (BMC); (2) The BIOS also records a Generic Hardware Error Source (GHES) log in the format of the Advanced Platform Error Interfaces (APEI) specification.
[0027] Among them, the SEL log is recorded and saved by the BMC. This log records some important events of the entire system, mainly including power-on / off and various hardware failures. The GHES log is the hardware failure information reported by the BIOS to the Operating System (OS). The OS receives the GHES information and prints a GHES log.
[0028] The inventors of this application found through research that in the above traditional solution, during the memory exception handling when the CPU enters the BIOS mode, all CPUs of the computing device enter the BIOS mode from the OS mode, and the operation of the OS and user applications is paused. The BIOS needs to wait for the memory exception log recording to be completed before returning the execution flow to the OS, that is, the CPU resumes the OS mode and resumes running the OS and user applications. This memory exception handling method results in a relatively long delay in the CPU resuming the OS mode.
[0029] Among them, the SMI interrupt triggered by the number of memory CEs exceeding the threshold will cause a large overall machine time overhead online. The overhead of a single SMI interrupt is about 100 - 200 ms on some platforms and 300 - 800 ms on some platforms, and the impact scope is all users on the entire server. The inventors of this application also found that the main source of the delay overhead is the BIOS recording memory exception logs. Among them, the operation of the BIOS writing the SEL log to the BMC is the main time overhead of the system. The BIOS needs to wait for the BMC to complete this IO operation before returning the execution flow to the OS. The operation of the BIOS writing the SEL log occupies more than 90% of the time overhead of memory exception handling.
[0030] Regarding the technical problem that the time delay for the processor to resume the OS mode due to the above-mentioned memory exception handling is relatively long, in some embodiments of the present application, the work of the traditional BIOS to record memory exception logs is moved to the OS side to be completed. The target driver running on the OS reads the CE count of the memory, and when the CE count reaches a target number less than the set CE quantity threshold through the target program running on the OS, the CE count is cleared in a timely manner, so that the CE count of the target counter is always less than the set CE quantity threshold, which can avoid triggering the SMI when the target counter reaches the set CE quantity threshold, and thus can avoid the processor from entering the BIOS. The target program running on the OS in this embodiment can also record memory exception logs when the total CE quantity of the memory is greater than or equal to the CE quantity threshold. In short, the memory exception log work in this embodiment is completed in a closed loop by the OS without the participation of the BIOS, so there is no need to switch the processor to the BIOS mode, and thus the time delay for the processor to resume the OS mode can be reduced.
[0031] The technical solutions provided by the embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0032] It should be noted that the same reference numerals represent the same object in the following drawings and embodiments. Therefore, once an object is defined in one drawing or embodiment, it does not need to be further discussed in the subsequent drawings and embodiments.
[0033] Figure 1 It is a schematic structural diagram of a computer system provided by an embodiment of the present application. As Figure 1 shown, the computer system includes: a processor 10 and a memory 20.
[0034] In this embodiment, the processor 10 is the processing unit of the computer system. The processor 10 can be a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), or a Microcontroller Unit (MCU); it can also be a programmable device such as a Field-Programmable Gate Array (FPGA), a Programmable Array Logic (PAL), a General Array Logic (GAL), or a Complex Programmable Logic Device (CPLD); or it can be an Advanced Reduced Instruction Set Compute (RISC) processor (Advanced RISC Machines, ARM) or a System on Chip (SoC), etc., but not limited thereto.
[0035] In the embodiments of this application, the specific implementation form of the memory is not limited. Optionally, the memory can be a Double Data Rate Synchronous Dynamic Random Access Memory (DDR SDRAM), also known as DDR memory. The DDR memory can be: DDR4 memory or DDR5 memory. Of course, the memory can also be a Phase Change Memory (PCM) or a High Bandwidth Memory (HBM), etc. The internal structure of the memory 20 will be described exemplarily below.
[0036] As Figure 2 shown, the channel is for the data bit width of the processor 10. If the processor 10 has a 64-bit data line and the memory chip is also 64 bits, then it is a single channel; if the processor 10 has a 64-bit data line, but the memory chip can support 128 bits, then the memory is a dual channel. Figure 2 The memory in [description] is a dual-channel memory. A Dual Inline Memory Module (DIMM) is a memory module where both the front and back sides of the printed circuit board have gold fingers that contact the memory slot on the motherboard, and this structure is called DIMM.
[0037] A physical array (Rank) refers to the memory chips connected to the same chip selector (Chip Select, CS). A memory chip includes several logical arrays (Bank). Cell (memory unit) is the basic unit of memory storage. Bank is a two-dimensional array composed of Cells. A Cell is used to store one bit of data.
[0038] Of course, the computer system may also include other storage media (not shown in the drawings), such as random access memory (ROM), disk, etc. In this embodiment, the ROM may include: BIOS ROM 30. The BIOS ROM 30 has BIOS 101 fixed therein.
[0039] BIOS101 is a startup program that cannot be tampered with and is solidified on the motherboard ROM chip. BIOS is responsible for the computer system self-test program and the system self-startup program, so it is the first program after the computer system is started. Due to its non-tamperability, the program is stored in the ROM chip, and the original settings can still be maintained after power failure. BIOS101 is the first program executed after the computer system is powered on, converting the system from a hardware state to a software state. In this embodiment, during the computer system startup process, the processor 10 loads the BIOS from the BIOS ROM 30. The BIOS detects and initializes the system's hardware devices through a series of steps, and loads the operating system (OS) 102. Accordingly, the processor 10 runs BIOS101 and OS102.
[0040] In this embodiment, the processor 10 further includes: a target counter 103 for counting the CE of the memory 20. The target counter 103 may be a register in the processor 10, such as a control status register (CSR). The CSR register may count the CE of the memory 20. Specifically, the memory 20 may include: one or more physical arrays (Ranks). Multiple refers to two or more. Figure 2 In the figure, the memory 20 includes 4 Ranks as an example, but it is not limited. Each Rank may correspond to a target counter (such as a CSR register) for counting the number of CEs of the Rank.
[0041] In the embodiments of the present application, the computer system may further include: a memory controller 40. The memory controller 40 integrates dedicated hardware to execute ECC and repair the CE of the memory 20. In the embodiments of the present application, when the memory controller determines that the number of CEs of the memory counted by the target counter 103 reaches or exceeds the set quantity threshold Q, the processor 10 is triggered to enter the BIOS 101. The BIOS 101 records the memory exception log. The set quantity threshold Q, which may also be referred to as the set CE quantity threshold Q, is the CE quantity threshold for triggering the processor to enter the BIOS.
[0042] Specifically, when the memory controller 40 monitors that the number of CEs counted by the target counter 103 reaches the set CE quantity threshold Q, it triggers a System Management Interrupt (SMI). The processor 10 responds to this SMI interrupt and enters the BIOS 101. The BIOS 101 records the memory exception log. Among them, the method for the BIOS 101 to record the memory exception log can refer to the foregoing related content and will not be elaborated here.
[0043] Since all processors 10 enter the BIOS mode and stop running the OS and user applications during the period when the BIOS 101 records the memory exception log. After the BIOS 101 finishes recording the memory exception log, the processor 10 enters the OS mode and resumes running the OS and user applications, and the time delay overhead for the processor 10 to resume the OS mode is relatively large.
[0044] In some embodiments of the present application, as Figure 1 shown, in order to reduce the time delay overhead for the processor 10 to resume the OS mode, a target program 104 is added to the processor 10. Among them, the target program 104 refers to software function modules, plugins, components, etc. added to implement the memory exception handling of the present application. The target program 104 runs in the OS 102. Among them, the kernel mode and the user mode are two different running states of the OS 102. In the kernel mode, the operating system program can be run to operate the hardware; while in the user mode, only application programs can be run. In this embodiment, in order to improve the flexibility of the target program 104, the target program 104 can be run in the user mode of the OS 102, that is, the target program 104 can be a user-mode program.
[0045] In this embodiment, the target driver 105 corresponding to the target counter 103 also runs in the OS 102. Among them, the target driver 105 can be a kernel driver in the operating system 102, which can collect and parse various fault-related information of the computing device, such as memory error information, etc., and print this information. In some embodiments, the target driver 105 can be an Error Detection And Correction Driver (EDAC Driver).
[0046] Based on the above target program 104 and target driver 105, the embodiments of the present application propose a new memory exception handling method to reduce the time delay overhead for the processor to resume the OS mode. The following will be described in detail in conjunction with Figure 1 and Figure 3 for specific illustration.
[0047] As Figure 1 and Figure 3 shown, when the memory controller 40 detects a CE in the memory, it can send an interrupt to the target driver 105 (corresponding to Figure 1 and Figure 3 step 1 "Interrupt" in Figure 3 ). As
[0048] shown in , this interrupt can be a Corrected Machine Check Interrupt (CMCI). Each time a CE occurs in the memory, such an interrupt will be triggered, and the time overhead of this interrupt (such as the CMCI interrupt) is only a few microseconds; when such an interrupt (such as the CMCI interrupt) is relatively frequent, the OS detects a CE storm, and the OS usually closes such an interrupt (such as the CMCI interrupt) and enters the Polling mode to prevent the frequent CMCI from affecting the system performance. The Polling mode means reporting once every few seconds to reduce the impact on the processor performance.
[0048] Specifically, as Figure 3 shown, when the memory controller 40 detects a CE in the memory, it can call the control function (Handler) corresponding to the interrupt; and use the function corresponding to the interrupt to send an interrupt to the target driver 105. In some embodiments, as Figure 3 shown in , if the interrupt is CMCI, the function corresponding to the interrupt is the CMCI control function (CMCI Handler). Correspondingly, when the memory controller 40 detects a CE in the memory, it can call the CMCI control function; and use the CMCI control function to send a CMCI interrupt to the target driver 105 (corresponding to Figure 3 step 1).
[0049] Correspondingly, the target driver 105 can, in response to the interrupt triggered by a CE in the memory, obtain the CE count X of the memory from the target counter 103 (corresponding toFigure 1 and Figure 3 In step 2) of this embodiment, in order to prevent the processor 10 from entering the BIOS mode, a target quantity P may be set, and the set CE quantity threshold Q for triggering the SMI. Among them, the CE quantity threshold Q is greater than the target quantity P. This is mainly because when the CE count of the target counter 103 reaches the CE quantity threshold Q, the memory controller will trigger the SMI, causing the processor 10 to enter the BIOS mode to record the memory exception log.
[0050] In the embodiment of the present application, the set CE quantity threshold Q for triggering the SMI is greater than the target quantity P. Preferably, the set CE quantity threshold Q for triggering the SMI is greater than the target quantity P and is a positive integer multiple of the target quantity P. That is, Q = N*P, N≥2, and N is an integer. For example, in some embodiments, the set CE quantity threshold Q for triggering the SMI = 5000, then P can be 1 / N of 5000; such as P = 2500, P = 1000 or P = 500, etc. The reason why the CE quantity threshold Q is a positive integer multiple of the target quantity P will be explained below and will not be elaborated here for the time being. Based on the pre-set target quantity P, when the CE quantity X of the memory read from the target counter 103 reaches the set target quantity P (i.e., X = P), the target driver 105 can send the CE quantity P of the memory read from the target counter 103 to the target program 104 (corresponding to Figure 1 and Figure 3 step 3) in.
[0051] Correspondingly, for the target program 104, the CE quantity P of the memory read from the target counter 103 can be obtained, and after obtaining the CE quantity P of the memory, the target driver 105 can be controlled to clear the CE count of the target counter 103 (corresponding to Figure 1 and Figure 3 steps 4 and 5) in, that is, controlling the target driver 105 to clear the CE count of the target counter 103 to zero. Specifically, after obtaining the CE quantity P of the memory, the target program 104 can send a CE count clear command to the target driver (corresponding to Figure 1 and Figure 3 step 4) in; the target driver 105 responds to the clear command and clears the CE count of the target counter 103 (corresponding to Figure 1 and Figure 3 step 5) in.
[0052] Among them, since the number of targets P is less than the set CE number threshold Q, therefore, after the number of CEs P in the memory of the target program 104, the target driver 105 is controlled to clear the CE count of the target counter 103, which can make the CE count of the target counter 103 always less than the CE number threshold Q, and thus can avoid the CE count of the target counter 103 reaching the CE number threshold Q. Furthermore, it can avoid the memory controller triggering the SMI to make the processor 10 enter the BIOS mode, and thus can avoid the processor 10 entering the BIOS mode.
[0053] After that, the target program 104 can calculate the total number of CEs Z in the memory according to the number of CEs P in the memory read from the target counter 103 and the historical number of CEs Y in the memory, that is, Z = P + Y. Among them, the historical number of CEs Y in the memory and the number of CEs P in the memory read from the target counter 103 can be stored in the memory space of the user state.
[0054] If the sum of the number of CEs P in the memory read from the target counter 103 and the historical number of CEs Y in the memory saved, that is, the total number of CEs Z in the memory is less than the above CE number threshold Q, then the target program 104 can save the number of CEs P in the memory sent by the target driver 105 this time, and use the sum Z of the number of CEs P in the memory received this time and the historical number of CEs Y as the new historical number of CEs Y in the memory; and keep the new historical number of CEs Y in the memory. After that, the target program 104 can continue to wait for the CE count of the target counter 103 sent by the target driver next time when the CE count of the target counter 103 reaches the set target number P, and add the CE count P of the target counter 103 received currently to the historical number of CEs Y of the previous time again until the sum of the CE count P of the target counter 103 received currently and the historical number of CEs of the previous time is greater than or equal to the above CE number threshold Q.
[0055] Furthermore, when the total number of CEs Z in the memory is greater than or equal to the above CE number threshold Q, the target program 104 can record a memory exception log (corresponding Figure 1 and Figure 3In step 6). After the target program 104 records the memory exception log, it can delete the saved historical CE count Y of the memory and the CE count P of the memory read from the target counter 103. That is, the target program 104 can, when the total CE count Z of the memory is greater than or equal to the set CE count threshold Q each time, promptly delete the saved historical CE count Y and the CE count P of the current target counter 103, so that the historical CE count Y of the memory saved by the target program 104 is less than the set CE count threshold Q. Only when the CE count P of a certain target counter 103 is accumulated, will it trigger the target program 104 to record the memory exception log, without the need to modify the existing BIOS code and configuration items.
[0056] In the embodiment of the present application, the work of the traditional BIOS to record the memory exception log is moved to the OS side to be completed. The target driver running on the OS reads the CE count of the memory, and when the CE count reaches the target count less than the set CE count threshold by the target program running on the OS, the CE count is promptly cleared, so that the CE count of the target counter is always less than the set CE count threshold, which can avoid the target counter reaching the set CE count threshold and triggering the SMI, and thus can avoid the processor entering the BIOS. The target program running on the OS in this embodiment can also record the memory exception log when the total CE count of the memory is greater than or equal to the CE count threshold. In short, the memory exception log work in this embodiment is completed in a closed loop by the OS without the participation of the BIOS, and thus there is no need to switch the processor to the BIOS mode, which can further reduce the time delay for the processor to resume the OS mode.
[0057] The inventor of the present application uses the memory exception handling method provided by the embodiment of the present application to test the OS mode recovery delay of the processor, and obtains that the memory exception handling method provided by the embodiment of the present application can reduce the OS mode recovery delay of the processor to less than 1 ms. Compared with the time overhead of hundreds of milliseconds in the traditional solution, this solution greatly reduces the time delay.
[0058] On the other hand, the memory exception handling method provided by the embodiment of the present application is completed entirely by the OS without the participation of the BIOS, without modifying the code and configuration items of the BIOS, and only needs to modify the software on the OS side to complete, without the need to restart the computing device, which can improve the flexibility of the solution deployment and is easy to be put on line for deployment.
[0059] In some embodiments, in order to further reduce the time delay for the processor to resume the OS mode, the target program 104 can record the memory exception log in an asynchronous manner when the total CE count of the memory is greater than or equal to the CE count threshold. In this way, the processor 10 does not need to wait for the memory exception log to be recorded and can continue to run the OS and user application programs, which helps to further reduce the time delay for the processor to resume the OS mode.
[0060] Alternatively, the processor 10 may use some of the processor cores to run the target program 104 and the target driver 105. In this way, when the target program 104 and the target driver 105 perform the above-mentioned memory exception handling operations, only some of the processor cores of the processor are occupied, and the other processor cores of the processor can continue to run the OS and other user applications, which can further reduce the overall time delay for the processor to resume the OS mode.
[0061] Alternatively, the processor 10 may use some of the processor cores to run the target program 104 and the target driver 105, and when the total number of CEs of the target program 104 in the memory is greater than or equal to the CE quantity threshold, the memory exception log is recorded in an asynchronous manner. In this way, on the one hand, when the target program 104 and the target driver 105 perform the above-mentioned memory exception handling operations, only some of the processor cores of the processor are occupied, and the other processor cores of the processor can continue to run the OS and other user applications, which can reduce the overall time delay for the processor to resume the OS mode. On the other hand, by recording the memory exception log in an asynchronous manner, the processor 10 does not need to wait for the memory exception log recording to be completed and can continue to run the OS and user applications, which helps to further reduce the time delay for the processor to resume the OS mode.
[0062] In some embodiments of the present application, the memory exception log includes: the SEL log and / or the GHES log. For the descriptions of the SEL log and the GHES log, reference may be made to the relevant content of the foregoing embodiments, which will not be elaborated herein. The SEL log is recorded and saved by the BMC. Therefore, as Figure 3 shown, the computer system may further include: a BMC 50. The BMC 50 is electrically connected to the processor 10.
[0063] The BMC 50 is a small operating system independent of the server system, and may be a chip integrated on the motherboard or inserted on the motherboard in the form of a serial interface or the like. The serial interface may be a high-speed serial computer expansion bus standard interface, such as a Peripheral Component Interconnect Express (PCIe) interface or the like. The BMC 50 has an independent firmware system, does not depend on other hardware on the system (such as the CPU, memory, etc.), and does not depend on the BIOS, OS, etc. The BMC 50 can interact with the BIOS and the OS. The BMC 50 usually includes a processor with firmware, a memory, and a network interface ( Figure 3 not shown in the figure).
[0064] Correspondingly, as Figure 3As shown, the target program 104 recording the memory exception log can be implemented as follows: when the total number Z of CEs in the memory is greater than or equal to the set CE quantity threshold Q, an SEL log can be written to the BMC 50 (corresponding to Figure 3 step 6) in it. Specifically, when the total number Z of CEs in the memory is greater than or equal to the set CE quantity threshold Q, the target program 104 can send a write command indicating that the BMC 50 writes the SEL log to the BMC 50. This write command includes: memory exception description information. Among them, the memory exception description information refers to the information used to describe the attributes of the memory exception, including but not limited to: the location information of the CE occurrence in the memory and the memory error type, etc.
[0065] Combined with Figure 2 the shown internal memory structure schematic diagram, a memory cell (Cell) is used to store 1 bit of data. Based on this, the location information of the CE occurrence in the memory can be accurate to the Cell granularity, that is, accurate to the memory row and memory column where the CE occurs. Correspondingly, the location information of the CE occurrence in the memory can include the socket (Socket), memory die (Die), channel (Channel), slot (Slot), physical rank (rank), sub-physical rank (Sub-Rank), logical array group (BankGroup), logical array (Bank), and row (Row) address and column (Column) address, etc., where the CE occurs.
[0066] The memory error type can reflect the memory error program to a certain extent and can be a CE, UE, or fatal error, etc.
[0067] Correspondingly, the BMC 50 can respond to the above write command and write the SEL log according to the memory exception description information. Specifically, the BMC 50 can respond to the above write command and write the memory exception description information into the SEL log according to the format of the SEL log.
[0068] Or, the target program 104 recording the memory exception log can be implemented as follows: when the total number Z of CEs in the memory is greater than or equal to the CE quantity threshold Q, the target program 104 controls the target driver to write a GHES log. Specifically, when the total number Z of CEs in the memory is greater than or equal to the set CE quantity threshold Q, the target program 104 can send a write command indicating that the target driver 105 writes the GHES log to the target driver 105. This write command includes: memory exception description information. Among them, the content of the memory exception description information here is the same as the memory exception description information included in the aforementioned write command for writing the SEL.
[0069] Accordingly, the target driver 105 can respond to the above write command and write to the GHES log according to the memory exception description information. Specifically, the target driver 105 can respond to the above write command and write the memory exception description information to the GHES log in the format of the GHES log.
[0070] Alternatively, the implementation of the target program 104 recording the memory exception log can be: when the total number Z of CEs in the memory is greater than or equal to the set CE quantity threshold Q, the SEL log can be written to the BMC 50; and, control the target driver 105 to write to the GHES log. For the specific implementation manners of the target driver 105 writing the SEL log to the BMC 50 and controlling the target driver 105 to write to the GHES log, reference can be made to the relevant content of the above embodiments, which will not be elaborated here.
[0071] In the traditional solution, the BIOS records the memory exception log when the CE count of the target counter reaches the set CE quantity threshold Q, that is, the BIOS writes the SEL log to the BMC and the BIOS records the GHES log when the CE count of the target counter reaches the set CE quantity threshold Q. Without changing the original code of the BIOS writing the SEL log and the GHES log, the above CE quantity threshold Q can be set to N times the target quantity P, where N≥2 and N is an integer.
[0072] The reason for setting the target quantity P as 1 / N of the CE quantity threshold Q in the above embodiments, where N≥2 and N is an integer, is that in this way, the historical CE quantity Y of the memory and the CE quantity P of the memory read from the target counter 103 can be accumulated, and the total CE quantity of the memory obtained is equal to the CE quantity threshold Q. Without changing the original logic of recording the memory exception log, without modifying the BIOS code and configuration items, and without modifying the BMC hardware and its code, the development difficulty can be reduced and it is easy to be put on the line for deployment.
[0073] Based on Figure 2 the internal structure of the memory 20 shown, the memory 20 includes at least one physical array (Rank). Generally, the memory 20 includes multiple physical arrays (Ranks). Multiple means 2 or more than 2. Figure 2 In the figure, the memory 20 includes 4 Ranks as an example for illustration, but it does not constitute a limitation. For the memory 20, the CE quantity can be counted from the granularity of the Rank. Accordingly, each Rank can correspond to a target counter 103 for counting the CEs in that Rank.
[0074] Accordingly, when a CE occurs in at least one Rank in the memory 20, the memory controller 40 may send an interrupt to the target driver 105. The interrupt may include: the identification of the Rank where the CE occurs. Among them, the interrupt may be a CMCI interrupt. Accordingly, the CMCI interrupt includes the identification of the Rank where the CE occurs.
[0075] Accordingly, for the target driver 105 to obtain the CE count of the memory from the target counter 103 in response to the interrupt triggered by a CE in the memory can be implemented as: in response to the interrupt sent by the memory controller 40 (such as a CMCI interrupt), determine the target Rank corresponding to the identification of the Rank where the CE occurs, and read the CE count X of the target Rank from the target counter 103 corresponding to the target Rank.
[0076] Further, when the CE count X of the target Rank counted by the target counter 103 corresponding to the target Rank reaches the set target count P (i.e., X = P), send the CE count P of the read target Rank and the memory location information A1 where the CE occurs to the target program 104. In this embodiment, the memory location information A1 where the CE occurs can be accurate to the Rank granularity.
[0077] Accordingly, for the target program 104, it can obtain the CE count P of the target Rank read from the target counter 103 and the memory location information A1 where the CE occurs, and after obtaining the CE count P of the target Rank, control the target driver 105 to clear the CE count of the target counter 103 corresponding to the target Rank, that is, control the target driver 105 to clear the CE count of the target counter 103 corresponding to the target Rank to zero. Among them, since the target count P is less than the set CE count threshold Q, therefore, after the CE count P of the target Rank, the target program 104 controls the target driver 105 to clear the CE count of the target counter 103 corresponding to the target Rank, which can make the CE count of the target counter 103 corresponding to the target Rank always less than the CE count threshold Q, and thus can avoid the memory controller triggering the SMI to cause the processor 10 to enter the BIOS mode, and thus can avoid the processor 10 entering the BIOS mode.
[0078] Further, accordingly, the target program 104 can obtain the historical CE count of the target Rank from the saved historical CE counts according to the memory location information A1 where the CE occurs. Specifically, the target program 104 can obtain, from the saved historical CE counts, the historical CE count whose memory location information where the CE occurs is the received memory location information A1 as the historical CE count Y of the target Rank according to the received memory location information A1 where the CE occurs.
[0079] Furthermore, the target program 104 can calculate the total number of CEs Z of the target Rank according to the number of CEs P of the memory read from the target counter 103 corresponding to the target Rank and the historical number of CEs Y of the target Rank, that is, Z = P + Y. Among them, the historical number of CEs Y of the target Rank and the number of CEs P of the target Rank read from the target counter 103 can be stored in the memory space of the user state.
[0080] If the sum of the number of CEs P of the target Rank read from the target counter 103 and the historical number of CEs Y of the target Rank stored, that is, the total number of CEs Z of the target Rank is less than the above CE quantity threshold Q, the target program 104 can save the number of CEs P of the target Rank sent by the target driver 105 this time, and use the sum Z of the number of CEs P of the target Rank received this time and the historical number of CEs Y of the target Rank as the new historical number of CEs Y of the target Rank; and maintain the new historical number of CEs Y of the target Rank. After that, the target program 104 can continue to wait for the CE count of the target counter 103 sent by the target driver next time when the CE count of the target counter 103 corresponding to the target Rank reaches the set target quantity P, and add the CE count P of the target counter 103 corresponding to the target Rank received currently and the historical number of CEs Y of the previous time again until the sum of the CE count P of the target counter 103 corresponding to the target Rank received currently and the historical number of CEs of the previous time is greater than or equal to the above CE quantity threshold Q.
[0081] Furthermore, the target program 104 can record a memory exception log when the total number of CEs Z of the target Rank is greater than or equal to the above CE quantity threshold Q.
[0082] After the target program 104 records the memory exception log, it can delete the historical number of CEs Y of the target Rank stored and the number of CEs P of the target Rank read from the target counter 103 above. That is, the target program 104 can delete the historical number of CEs Y of the target Rank stored and the CE count P of the target counter 103 corresponding to the target Rank of the current time in time when the total number of CEs Z of each target Rank is greater than or equal to the above set CE quantity threshold Q. This can make the historical number of CEs Y of the target Rank saved by the target program 104 less than the above set CE quantity threshold Q. Only when adding the CE count P of the target counter 103 corresponding to a certain target Rank, it will trigger the target program 104 to record a memory exception log, without changing the existing BIOS code and configuration items.
[0083] In some embodiments, in order to further reduce the time delay for the processor to resume the OS mode, when the total number of CEs in the target Rank is greater than or equal to the CE quantity threshold, the target program 104 may record the memory exception log in an asynchronous manner. In this way, the processor 10 does not need to wait for the memory exception log recording to complete and can continue to run the OS and user applications, which helps to further reduce the time delay for the processor to resume the OS mode.
[0084] Alternatively, the processor 10 may use some processor cores to run the target program 104 and the target driver 105. In this way, when the target program 104 and the target driver 105 perform the above-mentioned memory exception handling operations, only some processor cores of the processor are occupied, and the other processor cores of the processor can continue to run the OS and other user applications, which can further reduce the overall time delay for the processor to resume the OS mode.
[0085] Alternatively, the processor 10 may use some processor cores to run the target program 104 and the target driver 105, and when the total number of CEs in the target Rank is greater than or equal to the CE quantity threshold, the target program 104 records the memory exception log in an asynchronous manner. In this way, on the one hand, when the target program 104 and the target driver 105 perform the above-mentioned memory exception handling operations, only some processor cores of the processor are occupied, and the other processor cores of the processor can continue to run the OS and other user applications, which can reduce the overall time delay for the processor to resume the OS mode. On the other hand, by recording the memory exception log in an asynchronous manner, the processor 10 does not need to wait for the memory exception log recording to complete and can continue to run the OS and user applications, which helps to further reduce the time delay for the processor to resume the OS mode.
[0086] In some embodiments of the present application, the memory exception log includes: SEL log and / or GHES log. Correspondingly, the target program 104 recording the memory exception log can be implemented as: when the total number of CEs Z in the target Rank is greater than or equal to the set CE quantity threshold Q, the SEL log can be written to the BMC 50. Specifically, the target program 104 may send a write command indicating that the BMC 50 writes the SEL log to the BMC 50 when the total number of CEs Z in the target Rank is greater than or equal to the set CE quantity threshold Q. The write command includes: memory exception description information. Among them, the memory exception description information refers to the information used to describe the attributes of the memory exception, including but not limited to: the location information of the CE occurrence in the memory and the memory error type, etc.
[0087] Correspondingly, the BMC 50 may respond to the above write command and write the SEL log according to the memory exception description information. Specifically, the BMC 50 may respond to the above write command and write the memory exception description information into the SEL log according to the format of the SEL log.
[0088] Alternatively, the target program 104 recording the memory exception log can be implemented as follows: when the total number Z of CEs in the target Rank is greater than or equal to the CE quantity threshold Q, the target program 104 controls the target driver to write the GHES log. Specifically, when the total number Z of CEs in the memory is greater than or equal to the above-set CE quantity threshold Q, the target program 104 can send a write command to the target driver 105 to indicate that the target driver 105 writes the GHES log. The write command includes: memory exception description information.
[0089] Correspondingly, the target driver 105 can respond to the above write command and write the GHES log according to the memory exception description information. Specifically, the target driver 105 can respond to the above write command and write the memory exception description information into the GHES log according to the format of the GHES log.
[0090] Alternatively, the target program 104 recording the memory exception log can be implemented as follows: when the total number Z of CEs in the target Rank is greater than or equal to the above-set CE quantity threshold Q, the SEL log can be written to the BMC 50; and, the target driver 105 is controlled to write the GHES log. For the specific implementation manners of the target driver 105 writing the SEL log to the BMC 50 and controlling the target driver 105 to write the GHES log, reference can be made to the relevant content of the above embodiments, which will not be elaborated herein.
[0091] In the traditional solution, the BIOS records the memory exception log when the CE count of the target counter corresponding to the Rank reaches the set CE quantity threshold Q, that is, the BIOS writes the SEL log to the BMC and the BIOS records the GHES log when the CE count of the target counter corresponding to the Rank reaches the set CE quantity threshold Q. Without changing the original code of the BIOS writing the SEL log and the GHES log, the above CE quantity threshold Q can be set to N times the target quantity P, where N≥2 and is an integer.
[0092] The reason for setting the target quantity P as 1 / N of the CE quantity threshold Q in the above embodiments, where N≥2 and is an integer, is that in this way, the historical CE quantity Y of the memory and the CE quantity P of the memory read from the target counter 103 can be accumulated, and the total number of CEs of the Rank obtained is equal to the CE quantity threshold Q. Without changing the original logic of recording the memory exception log, without modifying the BIOS code and configuration items, and without modifying the BMC hardware and its code, the development difficulty can be reduced and it is easy to be deployed online.
[0093] It should be noted that the computer system provided in the above embodiments can be implemented as the system of a computing device. The computing device can be implemented as a terminal device such as a desktop computer, a laptop computer, a mobile phone, or an Internet of Things device; it can also be various server devices such as a traditional server, a cloud server, or a server cluster.
[0094] The computer system of the computing device may further include components such as a communication component, a power supply component, a display component, and an audio component. Figure 1 and Figure 3 Only some components are schematically shown, which does not mean that the computer system must include Figure 1 and Figure 3 all the components shown, nor does it mean that the computer system can only include Figure 1 and Figure 3 the components shown.
[0095] In addition to the above computer system, the embodiments of the present application also provide a memory exception handling method. The memory exception handling method provided by the embodiments of the present application will be exemplarily described below.
[0096] Figure 4 is a schematic flowchart of the memory exception handling method provided by the embodiments of the present application. As Figure 4 shown, the memory exception handling method mainly includes:
[0097] 401. The target driver running in the OS responds to the interrupt triggered by the occurrence of CE in the memory, and obtains the CE count of the memory from the target counter.
[0098] 402. When the CE count of the memory reaches the set target count, the target driver sends the CE count of the memory to the target program running in the OS.
[0099] 403. The target program obtains the CE count of the memory and clears the CE count of the target counter; the CE count threshold is greater than the target count.
[0100] 404. The target program calculates the total CE count of the memory according to the CE count of the memory and the historical CE count of the memory.
[0101] 405. When the total CE count of the memory is greater than or equal to the CE count threshold, the target program records a memory exception log.
[0102] In this embodiment, the processor runs BIOS and OS. The processor includes: a target counter for counting the CEs of the memory. Among them, the target counter can be a register in the processor 10, such as the CSR register number. Specifically, the memory can include: one or more physical ranks (Ranks). Multiple means two or more. Each Rank can correspond to a target counter (such as a CSR register) for counting the number of CEs of that Rank.
[0103] In an embodiment of the present application, the computer system may further include: a memory controller. When the memory controller determines that the number of CEs of the memory counted by the target counter reaches or exceeds the set CE number threshold Q, it triggers the processor to enter BIOS. BIOS records the memory exception log.
[0104] Since all processors enter the BIOS mode and stop running the OS and user applications during the period when BIOS records the memory exception log. After BIOS finishes recording the memory exception log, the processor enters the OS mode and resumes running the OS and user applications, and the time delay overhead of the processor 10 to resume the mode is relatively large.
[0105] In some embodiments of the present application, in order to reduce the time delay overhead of the processor to resume the OS mode, a target program is added to the processor. The target program runs in the OS. In order to improve the flexibility of the target program, the target program can run in the user state of the OS, that is, the target program can be a user state program.
[0106] In this embodiment, the OS also runs a target driver corresponding to the target counter. Among them, the target driver can be a kernel driver in the operating system. In some embodiments, the target driver can be an EDAC driver.
[0107] Based on the above-mentioned target program and target driver running in the OS, an embodiment of the present application proposes a new memory exception handling method to reduce the time delay overhead of the processor to resume the OS mode. The following is a specific description.
[0108] The memory controller can send an interrupt to the target driver when it detects that a CE occurs in the memory. This interrupt can be CMCI. Each time a CE occurs in the memory, this type of interrupt will be triggered, and the time overhead of this interrupt (such as CMCI interrupt) is only a few microseconds.
[0109] Accordingly, in step 401, the target driver can obtain the number of CEs X of the memory from the target counter in response to the interrupt triggered by the occurrence of a CE in the memory. In this embodiment, in order to prevent the processor from entering the BIOS mode, a target number P and the above-set CE number threshold Q for triggering SMI can be set. The CE number threshold Q is greater than the target number P. This is mainly because when the CE count of the target counter reaches the CE number threshold Q, the memory controller will trigger SMI, causing the processor to enter the BIOS mode to record the memory exception log.
[0110] In the embodiment of the present application, the above-set CE number threshold Q for triggering SMI is greater than the target number P. Preferably, the above-set CE number threshold Q for triggering SMI is greater than the target number P and is a positive integer multiple of the target number P. That is, Q = N*P, N≥2 and is an integer. The reason why the CE number threshold Q is a positive integer multiple of the target number P will be explained below and will not be elaborated here for the time being. Based on the pre-set target number P, in step 402, when the number of CEs X of the memory read from the target counter by the target driver reaches the set target number P (i.e., X = P), the target driver can send the number of CEs P of the memory read from the target counter to the target program.
[0111] Accordingly, for the target program, in step 403, the number of CEs P of the memory read from the target counter can be obtained, and after obtaining the number of CEs P of the memory, the target driver can be controlled to clear the CE count of the target counter, that is, the target driver can be controlled to clear the CE count of the target counter to zero. Specifically, after the target program obtains the number of CEs P of the memory, it can send a CE count clear command to the target driver; the target driver responds to the clear command and clears the CE count of the target counter.
[0112] Since the target number P is less than the set CE number threshold Q, after the target program obtains the number of CEs P of the memory and controls the target driver to clear the CE count of the target counter 103, the CE count of the target counter can always be less than the CE number threshold Q, which can avoid the CE count of the target counter 103 reaching the CE number threshold Q, and further avoid the memory controller triggering SMI to cause the processor to enter the BIOS mode, that is, avoid the processor entering the BIOS mode.
[0113] After that, in step 404, the target program can calculate the total number of CEs Z of the memory according to the number of CEs P of the memory read from the target counter and the historical CE number Y of the memory, that is, Z = P + Y. The historical CE number Y of the memory and the number of CEs P of the memory read from the target counter can be stored in the memory space of the user state.
[0114] If the CE quantity P of the memory read from the target counter and the accumulated sum of the historical CE quantity Y of the saved memory, that is, the total CE quantity Z of the memory, is less than the above CE quantity threshold Q, the target program can save the CE quantity P of the memory driven by the target for this transmission, and use the accumulated sum Z of the CE quantity P of the memory received this time and the historical CE quantity Y as the new historical CE quantity Y of the memory; and maintain the new historical CE quantity Y of the memory. After that, the target program can continue to wait for the CE count of the target counter sent by the next target drive when the CE count of the target counter reaches the set target quantity P, and accumulate the CE count P of the target counter received currently with the historical CE quantity Y of the previous time again until the accumulated sum of the CE count P of the target counter received in the current time and the historical CE quantity of the previous time is greater than or equal to the above CE quantity threshold Q.
[0115] Further, in step 405, the target program can record a memory exception log when the total CE quantity Z of the memory is greater than or equal to the above CE quantity threshold Q. After the target program records the memory exception log, it can delete the historical CE quantity Y of the saved memory and the CE quantity P of the memory read from the target counter. That is, the target program can timely delete the saved historical CE quantity Y and the CE count P of the target counter in the current time every time the total CE quantity Z of the memory is greater than or equal to the above set CE quantity threshold Q, so that the historical CE quantity Y of the memory saved by the target program is less than the above set CE quantity threshold Q. Only when the CE count P of a certain target counter is accumulated, will it trigger the target program to record a memory exception log, without changing the existing BIOS code and configuration items.
[0116] In the embodiment of the present application, the work of the traditional BIOS recording the memory exception log is moved to the OS side to be completed. The target drive running on the OS reads the CE count of the memory, and when the CE count reaches the target quantity less than the set CE quantity threshold by the target program running on the OS, the CE count is cleared in time, so that the CE count of the target counter is always less than the set CE quantity threshold, which can avoid the target counter reaching the set CE quantity threshold to trigger SMI and avoid the processor entering the BIOS. The target program running on the OS in this embodiment can also record a memory exception log when the total CE quantity of the memory is greater than or equal to the CE quantity threshold. In short, the work of the memory exception log in this embodiment is completed in a closed loop by the OS without the participation of the BIOS, so there is no need to switch the processor to the BIOS mode, and thus the time delay for the processor to resume the OS mode can be reduced.
[0117] The inventors of the present application used the memory exception handling method provided by the embodiments of the present application to test the OS mode recovery latency of the processor, and obtained that the memory exception handling method provided by the embodiments of the present application can reduce the OS mode recovery latency of the processor to less than 1 ms. Compared with the time overhead of hundreds of milliseconds in the traditional solution, this solution greatly reduces the time latency.
[0118] On the other hand, the memory exception handling method provided by the embodiments of the present application is completed by the OS throughout the process without the participation of the BIOS. It is not necessary to modify the code and configuration items of the BIOS. Only the software on the OS side needs to be modified to complete it, and it is not necessary to restart the computing device, which can improve the flexibility of the solution deployment and is easy to be put on line.
[0119] In some embodiments, in order to further reduce the time latency of the processor to recover the OS mode, when the total number of CEs in the memory is greater than or equal to the CE quantity threshold, the target program can record the memory exception log in an asynchronous manner. In this way, the processor can continue to run the OS and user applications without waiting for the memory exception log recording to be completed, which helps to further reduce the time latency of the processor to recover the OS mode.
[0120] Alternatively, the processor can use some processor cores to run the target program and the target driver. In this way, when the target program and the target driver perform the above-mentioned memory exception handling operations, only some processor cores of the processor are occupied, and the other processor cores of the processor can continue to run the OS and other user applications, which can further reduce the overall time latency of the processor to recover the OS mode.
[0121] Alternatively, the processor can use some processor cores to run the target program and the target driver, and when the total number of CEs in the memory of the target program is greater than or equal to the CE quantity threshold, record the memory exception log in an asynchronous manner. In this way, on the one hand, when the target program and the target driver perform the above-mentioned memory exception handling operations, only some processor cores of the processor are occupied, and the other processor cores of the processor can continue to run the OS and other user applications, which can reduce the overall time latency of the processor to recover the OS mode. On the other hand, by recording the memory exception log in an asynchronous manner, the processor can continue to run the OS and user applications without waiting for the memory exception log recording to be completed, which helps to further reduce the time latency of the processor to recover the OS mode.
[0122] In some embodiments of the present application, the memory exception log includes: SEL log and / or GHES log. For the description of the SEL log and the GHES log, reference can be made to the relevant content of the foregoing embodiments, which will not be elaborated here. The SEL log is recorded and saved by the BMC. Therefore, the computer system may further include: BMC. The BMC is electrically connected to the processor.
[0123] Accordingly, the target program's recording of the memory exception log can be implemented as follows: When the total number Z of CEs in the memory is greater than or equal to the set CE quantity threshold Q, an SEL log can be written to the BMC. Specifically, when the total number Z of CEs in the memory is greater than or equal to the set CE quantity threshold Q, the target program can send a write command to the BMC indicating that the BMC writes the SEL log. This write command includes: memory exception description information. Among them, the memory exception description information refers to the information used to describe the attributes of the memory exception, including but not limited to: the location information of the CE occurrence in the memory and the memory error type, etc.
[0124] Accordingly, the BMC can respond to the above write command and write the SEL log according to the memory exception description information. Specifically, the BMC can respond to the above write command and write the memory exception description information into the SEL log in accordance with the format of the SEL log.
[0125] Alternatively, the target program's recording of the memory exception log can be implemented as follows: When the total number Z of CEs in the memory is greater than or equal to the CE quantity threshold Q, the target program controls the target driver to write the GHES log. Specifically, when the total number Z of CEs in the memory is greater than or equal to the set CE quantity threshold Q, the target program can send a write command to the target driver indicating that the target driver writes the GHES log. This write command includes: memory exception description information.
[0126] Accordingly, the target driver can respond to the above write command and write the GHES log according to the memory exception description information. Specifically, the target driver can respond to the above write command and write the memory exception description information into the GHES log in accordance with the format of the GHES log.
[0127] Alternatively, the target program's recording of the memory exception log can be implemented as follows: When the total number Z of CEs in the memory is greater than or equal to the set CE quantity threshold Q, an SEL log can be written to the BMC; and, the target driver is controlled to write the GHES log. For the specific implementation methods of the target driver writing the SEL log to the BMC and controlling the target driver to write the GHES log, reference can be made to the relevant content of the above embodiments, which will not be elaborated here.
[0128] Since in the traditional solution, the BIOS records the memory exception log when the CE count of the target counter reaches the set CE quantity threshold Q, that is, the BIOS writes the SEL log to the BMC and the BIOS records the GHES log when the CE count of the target counter reaches the set CE quantity threshold Q. Without changing the original code for the BIOS to write the SEL log and the GHES log, the above CE quantity threshold Q can be set to N times the target quantity P, where N ≥ 2 and is an integer.
[0129] The reason for setting the target quantity P to 1 / N of the CE quantity threshold Q in the above embodiments, where N≥2 and N is an integer, is that in this way, the historical CE quantity Y of the memory can be accumulated with the CE quantity P of the memory read from the target counter 103, and the total CE quantity of the memory obtained is equal to the CE quantity threshold Q. Without changing the original logic of recording the memory exception log, without modifying the BIOS code and configuration items, and without modifying the BMC hardware and its code, the development difficulty can be reduced, and it is easy to be deployed online.
[0130] Based on Figure 2 From the internal structure of the memory 20 shown, the memory includes at least one physical array (Rank). Generally, a memory includes multiple physical arrays (Ranks). Multiple means 2 or more. For a memory, the CE quantity can be counted from the granularity of the Rank. Correspondingly, each Rank can correspond to a target counter for counting the CEs in that Rank.
[0131] Correspondingly, the memory controller can send an interruption to the target driver when at least one Rank in the memory 20 has a CE. The interruption can include: the identifier of the Rank where the CE occurs. Among them, the interruption can be a CMCI interruption. Correspondingly, the CMCI interruption includes the identifier of the Rank where the CE occurs.
[0132] Correspondingly, the implementation of step 401 above, where the target driver obtains the CE quantity of the memory from the target counter 103 in response to the interruption triggered by the CE of the memory, can be: in response to the interruption sent by the memory controller (such as a CMCI interruption), determine the target Rank corresponding to the identifier of the Rank where the CE occurs, and read the CE quantity X of the target Rank from the target counter corresponding to the target Rank.
[0133] Further, step 402 above can be implemented as: when the CE quantity X of the target Rank counted by the target counter corresponding to the target Rank reaches the above-set target quantity P (i.e., X = P), send the CE quantity P of the target Rank read and the memory location information A1 where the CE occurs to the target program. In this embodiment, the memory location information A1 where the CE occurs can be accurate to the Rank granularity.
[0134] Accordingly, step 403 described above can be implemented as follows: For the target program, the CE count P of the target Rank read from the target counter corresponding to the target Rank and the memory location information A1 where the CE occurred can be obtained. After obtaining the CE count P of the target Rank, the target driver is controlled to clear the CE count of the target counter corresponding to the target Rank, that is, the target driver is controlled to clear the CE count of the target counter corresponding to the target Rank to zero. Among them, since the target count P is less than the set CE count threshold Q, therefore, after the CE count P of the target Rank in the target program, controlling the target driver to clear the CE count of the target counter corresponding to the target Rank can make the CE count of the target counter corresponding to the target Rank always less than the CE count threshold Q, and thus it can be avoided that the CE count of the target counter corresponding to the target Rank reaches the CE count threshold Q, and further it can be avoided that the memory controller triggers the SMI to cause the processor to enter the BIOS mode, and thus it can be avoided that the processor enters the BIOS mode.
[0135] Furthermore, step 403 can be implemented as follows: The target program can obtain the historical CE count of the target Rank from the saved historical CE counts according to the memory location information A1 where the CE occurred. Specifically, the target program can obtain, from the saved historical CE counts, the historical CE count whose memory location information where the CE occurred is the received memory location information A1 of the CE as the historical CE count Y of the target Rank according to the received memory location information A1 where the CE occurred.
[0136] Furthermore, the target program can calculate the total CE count Z of the target Rank according to the CE count P of the memory read from the target counter corresponding to the target Rank and the historical CE count Y of the target Rank, that is, Z = P + Y. Among them, the historical CE count Y of the target Rank and the CE count P of the target Rank read from the target counter can be saved in the user-mode memory space.
[0137] If the cumulative sum of the number of CEs P of the target Rank read from the target counter and the historical number of CEs Y of the target Rank saved, that is, the total number of CEs Z of the target Rank is less than the above CE quantity threshold Q, the target program can save the number of CEs P of the target Rank that drives this transmission, and use the cumulative sum Z of the number of CEs P of the target Rank received this time and the historical number of CEs Y of the target Rank as the new historical number of CEs Y of the target Rank; and maintain the new historical number of CEs Y of the target Rank. After that, the target program can continue to wait for the CE count of the target counter sent when the CE count of the target counter corresponding to the target Rank reaches the set target quantity P for the next target drive, and again accumulate the CE count P of the target counter corresponding to the target Rank received currently and the historical number of CEs Y of the previous time until the cumulative sum of the CE count P of the target counter corresponding to the target Rank received in the current time and the historical number of CEs of the previous time is greater than or equal to the above CE quantity threshold Q.
[0138] Further, step 405 can be implemented as: the target program can record the memory exception log of the target Rank when the total number of CEs Z of the target Rank is greater than or equal to the above CE quantity threshold Q. After the target program records the memory exception log, it can delete the historical number of CEs Y of the target Rank saved and the number of CEs P of the target Rank read from the target counter above. That is, the target program can timely delete the historical number of CEs Y of the target Rank saved and the CE count P of the target counter corresponding to the target Rank in the current time each time the total number of CEs Z of the target Rank is greater than or equal to the above set CE quantity threshold Q, so that the historical number of CEs Y of the target Rank saved by the target program is less than the above set CE quantity threshold Q. Only when the CE count P of the target counter corresponding to a certain target Rank is accumulated, will it trigger the target program to record the memory exception log, without changing the existing BIOS code and configuration items.
[0139] In some embodiments, in order to further reduce the time delay for the processor to resume the OS mode, the target program can record the memory exception log of the target Rank in an asynchronous manner when the total number of CEs of the target Rank is greater than or equal to the CE quantity threshold. In this way, the processor can continue to run the OS and user application programs without waiting for the memory exception log to be recorded, which helps to further reduce the time delay for the processor to resume the OS mode.
[0140] Alternatively, the processor may utilize some of its processor cores to run the target program and the target driver. In this way, when the target program and the target driver perform the above-mentioned memory exception handling operations, only some of the processor cores of the processor are occupied, and the other processor cores of the processor can continue to run the OS and other user applications, which can further reduce the overall time delay for the processor to resume the OS mode.
[0141] Alternatively, the processor may utilize some of its processor cores to run the target program and the target driver, and when the total number of CEs of the target Rank in the target program is greater than or equal to the CE quantity threshold, the memory exception log of the target Rank is recorded in an asynchronous manner. In this way, on the one hand, when the target program and the target driver perform the above-mentioned memory exception handling operations, only some of the processor cores of the processor are occupied, and the other processor cores of the processor can continue to run the OS and other user applications, which can reduce the overall time delay for the processor to resume the OS mode. On the other hand, by recording the memory exception log in an asynchronous manner, the processor does not need to wait for the memory exception log to be recorded before it can continue to run the OS and user applications, which helps to further reduce the time delay for the processor to resume the OS mode.
[0142] In some embodiments of the present application, the memory exception log includes: SEL log and / or GHES log. Correspondingly, the target program's recording of the memory exception log can be implemented as: when the total number Z of CEs of the target Rank is greater than or equal to the above-mentioned set CE quantity threshold Q, the SEL log can be written to the BMC. Specifically, the target program can send a write command to the BMC indicating that the BMC 50 writes the SEL log when the total number Z of CEs of the target Rank is greater than or equal to the above-mentioned set CE quantity threshold Q. The write command includes: memory exception description information. Among them, the memory exception description information refers to the information used to describe the attributes of the memory exception, including but not limited to: the location information of the CE occurrence in the memory and the memory error type, etc.
[0143] Correspondingly, the BMC can respond to the above-mentioned write command and write the SEL log according to the memory exception description information. Specifically, the BMC can respond to the above-mentioned write command and write the memory exception description information into the SEL log according to the format of the SEL log.
[0144] Alternatively, the target program's recording of the memory exception log can be implemented as: the target program controls the target driver to write the GHES log when the total number Z of CEs of the target Rank is greater than or equal to the CE quantity threshold Q. Specifically, the target program can send a write command to the target driver indicating that the target driver writes the GHES log when the total number Z of CEs of the memory is greater than or equal to the above-mentioned set CE quantity threshold Q. The write command includes: memory exception description information.
[0145] Accordingly, the target driver can respond to the above write command and write to the GHES log according to the memory exception description information. Specifically, the target driver can respond to the above write command and write the memory exception description information to the GHES log in accordance with the format of the GHES log.
[0146] Alternatively, the implementation of the target program for recording memory exception logs can be as follows: when the total number Z of CEs of the target Rank is greater than or equal to the set CE quantity threshold Q, an SEL log can be written to the BMC 50; and, the target driver 105 can be controlled to write to the GHES log. For the specific implementation manners of the target driver 105 writing the SEL log to the BMC 50 and controlling the target driver 105 to write to the GHES log, reference can be made to the relevant content of the above embodiments, which will not be elaborated herein.
[0147] In the traditional solution, the BIOS records the memory exception log when the CE count of the target counter corresponding to the Rank reaches the set CE quantity threshold Q, that is, the BIOS writing the SEL log to the BMC and the BIOS recording the GHES log are triggered when the CE count of the target counter corresponding to the Rank reaches the set CE quantity threshold Q. Without changing the original code for the BIOS to write the SEL log and the GHES log, the above CE quantity threshold Q can be set to N times the target quantity P, where N≥2 and is an integer.
[0148] The reason for setting the target quantity P as 1 / N of the CE quantity threshold Q, where N≥2 and is an integer, in the above embodiments is that in this way, the historical CE quantity Y of the memory can be accumulated with the CE quantity P of the memory read from the target counter 103, and the total CE quantity of the Rank obtained is equal to the CE quantity threshold Q. Without changing the original logic for recording the memory exception log, without modifying the BIOS code and configuration items, and without modifying the BMC hardware and its code, the development difficulty can be reduced and it is easy to be put on the line for deployment.
[0149] Taking the set CE quantity threshold Q = 5000 and the target quantity P = 2500 as an example, the memory exception handling method provided by the embodiments of the present application will be exemplarily described below.
[0150] A certain machine has been continuously experiencing memory CE faults. When the CE count of the target counter of the target Rank in the memory reaches 2500, the target driver reports the CE count X (i.e., X = 2500) of the target counter of the target Rank to the target program. The target program receives the CE count X = 2500 of the target counter of the target Rank and controls the target driver to clear the CE count of the target counter of the target Rank. After a period of time, the CE count of this target Rank reaches 2500 again. The target driver reports the CE count X (i.e., X = 2500) of the target counter of the target Rank to the target program, and the target program receives the CE count X = 2500 of the target counter of the target Rank and controls the target driver to clear the CE count of the target counter of the target Rank.
[0151] On the target program side, a total of two 2500 counts are received. Therefore, when the total CE quantity of this target Rank reaches the set CE quantity threshold Q = 5000, the target program records a SEL log in the BMC. During the whole process, the highest value of the target counter corresponding to the target Rank is 2500, and the set CE quantity threshold is 5000. So, it will not trigger the SMI to cause the processor to enter the BIOS mode, effectively suppressing the generation of SMI interrupts, and also accurately recording the CE quantity and completing the recording of the SEL log, realizing the memory status monitoring of the computer system.
[0152] It should be noted that the execution subject of each step of the method provided in the above embodiments can be the same device, or different devices can be used as the execution subject of the method. For example, the execution subjects of steps 401 and 402 can be device A; or, the execution subject of step 401 can be device A, and the execution subject of step 402 can be device B; and so on.
[0153] In addition, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as 401, 402, etc. are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel.
[0154] Correspondingly, the embodiments of the present application also provide a computer-readable storage medium storing computer instructions. When the computer instructions are executed by one or more processors, one or more processors are caused to execute the steps in the above memory exception handling method.
[0155] In the embodiments of the present application, the memory is used to store computer programs and can be configured to store various other data to support operations on the device where it is located. Among them, the processor can execute the computer programs stored in the memory to implement corresponding control logic. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Electrically Programmable Read Only Memory (EPROM), Programmable Read Only Memory (PROM), Read Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0156] In the embodiments of the present application, the processor can be any hardware processing device capable of executing the above method logic. Optionally, the processor can be a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), or a Microcontroller Unit (MCU); it can also be a programmable device such as a Field-Programmable Gate Array (FPGA), a Programmable Array Logic (PAL), a General Array Logic (GAL), or a Complex Programmable Logic Device (CPLD); or an Advanced Reduced Instruction Set Compute (RISC) processor (Advanced RISC Machines, ARM) or a System on Chip (SoC), etc., but not limited thereto.
[0157] In an embodiment of the present application, the communication component is configured to facilitate communication between the device it is in and other devices in a wired or wireless manner. The device where the communication component is located can access a wireless network based on communication standards, such as Wireless Fidelity (WiFi), 2G or 3G, 4G, 5G, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component can also be implemented based on Near Field Communication (NFC) technology, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wide Band (UWB) technology, Bluetooth (BT) technology, or other technologies.
[0158] In an embodiment of the present application, the display component may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the display component includes a touch panel, the display component can be implemented as a touch screen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation.
[0159] In an embodiment of the present application, the power component is configured to provide power to various components of the device it is in. The power component may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device where the power component is located.
[0160] In an embodiment of the present application, the audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC). When the device where the audio component is located is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive external audio signals. The received audio signals can be further stored in the memory or transmitted via the communication component. In some embodiments, the audio component further includes a speaker for outputting audio signals. For example, for a device with a language interaction function, voice interaction with the user can be achieved through the audio component.
[0161] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0162] It should also be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.
[0163] Those skilled in the art should understand that the embodiments of this application can be provided as a method, a system, or a computer program product. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, read-only compact disc (CD-ROM), optical storage, etc.) containing computer-usable program code.
[0164] This application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of this application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0165] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.
[0166] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the process Figure 1 one process or multiple processes and / or blocks Figure 1 steps for the functions specified in one block or multiple blocks.
[0167] In a typical configuration, a computing device includes one or more processors (such as CPUs), input / output interfaces, network interfaces, and memory.
[0168] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.
[0169] The storage medium of a computer is a readable storage medium, also known as a readable medium. Readable storage media include permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD), or other optical storage, magnetic cassette tapes, disk storage, or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory computer-readable media, such as modulated data signals and carrier waves.
[0170] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the above elements.
[0171] The above content is only an embodiment of the present application and is not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A computer system, characterized in that, comprising: a processor and a memory; the processor includes: a target counter for counting correctable errors in the memory; the processor runs an operating system; the operating system runs a target driver, which is used to obtain the number of correctable errors in the memory from the target counter in response to an interrupt triggered by a correctable error occurring in the memory; when the number of correctable errors in the memory reaches a set target number, send the number of correctable errors in the memory to a target program running in the operating system; the target number is less than a set number threshold; the number threshold is the number threshold of correctable errors that triggers the processor to enter the basic input / output unit to record a memory exception log; the target program of the operating system is used to obtain the number of correctable errors in the memory and control the target driver to clear the count of the target counter; calculate the total number of correctable errors in the memory according to the number of correctable errors in the memory and the historical correctable errors in the memory; when the total number of correctable errors in the memory is greater than or equal to the number threshold, record a memory exception log.
2. A method for handling memory exceptions, applicable to a processor, characterized in that, the processor runs an operating system; the processor includes: a target counter for counting correctable errors in the memory; the method includes: a target driver running in the operating system obtains the number of correctable errors in the memory from the target counter in response to an interrupt triggered by a CE occurring in the memory; when the number of correctable errors in the memory reaches a set target number, send the number of correctable errors in the memory to a target program running in the operating system; the target program obtains the number of correctable errors in the memory and clears the count of the target counter; the target number is less than a set number threshold; the number threshold is the number threshold of correctable errors that triggers the processor to enter the basic input / output unit to record a memory exception log; calculate the total number of correctable errors in the memory according to the number of correctable errors in the memory and the historical correctable errors in the memory; when the total number of correctable errors in the memory is greater than or equal to the number threshold, record a memory exception log.
3. The method according to claim 2, characterized in that, the memory includes: at least one physical array, and each physical array corresponds to a target counter for counting correctable errors in the physical array; the interrupt triggered by a correctable error occurring in the memory is a correctable machine check interrupt; the target driver running in the operating system obtains the number of correctable errors in the memory from the target counter in response to an interrupt triggered by a correctable error occurring in the memory, including: the target driver obtains the identifier of the physical array where the correctable error occurs from the correctable machine check interrupt in response to the correctable machine check interrupt triggered by a correctable error occurring in the memory; Read the number of correctable errors of the target physical array corresponding to the identifier of the physical array where the correctable error occurs from the target counter corresponding to the target physical array; The case where the number of correctable errors in the memory reaches a set target number and sending the number of correctable errors in the memory to a target program running in the operating system includes: When the number of correctable errors of the target physical array reaches the target number, send the number of correctable errors of the target physical array and the memory location information where the correctable error occurs to the target program.
4. The method according to claim 3, wherein, The target program obtains the number of correctable errors in the memory and clears the count of the target counter, including: Obtain the number of correctable errors of the target physical array and the memory location information where the correctable error occurs; and control the target driver to clear the count of the target counter corresponding to the target physical array; The calculating the total number of correctable errors in the memory according to the number of correctable errors in the memory and the historical number of correctable errors in the memory includes: According to the memory location information where the correctable error occurs, obtain the historical number of correctable errors of the target physical array from the historical number of correctable errors; Calculate the total number of correctable errors of the target physical array according to the number of correctable errors of the target physical array and the historical number of correctable errors of the target physical array; The target program records a memory exception log when the total number of correctable errors in the memory reaches the number threshold, including: The target program records a memory exception log of the target physical array when the total number of correctable errors of the target physical array reaches the number threshold.
5. The method according to claim 4, wherein, further includes: After the target program records the memory exception log of the target physical array, delete the saved historical number of correctable errors of the target physical array and the number of correctable errors of the target physical array.
6. The method according to claim 4, wherein, The target program records a memory exception log of the target physical array when the total number of correctable errors of the target physical array reaches the number threshold, including: When the total number of correctable errors of the target physical array is greater than or equal to the number threshold, the target program writes a system event log of the target physical array to the motherboard controller electrically connected to the processor; and / or, When the total number of correctable errors of the target physical array reaches the number threshold, the target program controls the target driver to write a general hardware error source log; the memory exception log includes: the system event log and / or the general hardware error source log.
7. The method according to claim 6, wherein, When the total number of correctable errors in the target physical array is greater than or equal to the quantity threshold, the target program writes a system event log of the target physical array to the motherboard controller electrically connected to the processor, including: When the total number of correctable errors in the memory reaches the quantity threshold, send a first write command to the motherboard controller indicating that the motherboard controller writes the system event log; the first write command includes: first memory exception description information, so that in response to the first write command, the motherboard controller writes the system event log according to the first memory exception description information.
8. The method according to claim 6, characterized in that, When the total number of correctable errors in the target physical array reaches the quantity threshold, the target program controls the target driver to write a general hardware error source log, including: When the total number of correctable errors in the memory reaches the quantity threshold, the target program sends a second write command to the target driver indicating that the target driver writes the general hardware error source log; the second write command includes: second memory exception description information, so that in response to the second write command, the target driver writes the general hardware error source log according to the second memory exception description information.
9. The method according to claim 4, characterized in that, When the total number of correctable errors in the target physical array reaches the quantity threshold, the target program records a memory exception log of the target physical array, including: When the total number of correctable errors in the target physical array is greater than or equal to the quantity threshold, the target program records the memory exception log of the target physical array in an asynchronous manner.
10. The method according to any one of claims 2-9, characterized in that, The processor uses some processor cores to run the target program and the target driver.
11. The method according to any one of claims 2-9, characterized in that, The set quantity threshold is a positive integer multiple of the target quantity.
12. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by one or more processors, the one or more processors are caused to execute the steps in the method according to any one of claims 2-11.