Memory fault information processing method, memory fault prediction method and electronic equipment
The memory failure information is obtained through the preset interrupt handler of UEFI and uploaded to the target controller. Combined with the target failure rules, the problem of low memory failure prediction accuracy in the existing technology is solved, and the stable operation of the server is achieved.
Patent Information
- Application Number
- CN202510386807.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-04
AI Technical Summary
In the memory failure prediction, the prior art cannot accurately obtain all the information when CE Storm due to register information overwriting in memory failure prediction, resulting in low prediction accuracy and unrecoverable failures in time, affecting the stable operation of the server.
Triggered through the preset interrupt handler of UEFI, the memory is recoverable and unrecoverable fault information, and uploaded to the target controller, and predicted in combination with the target fault rules, including recording the number of recoverable faults, temperature and other information, and generating target information to predict unrecoverable faults.
It improves the accuracy and timeliness of memory failure prediction, reduces the risk of server downtime, and ensures stable hardware operation.
Smart Images

Figure CN120256184A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing, and in particular, to a method for processing memory failure information, a method for predicting memory failure, and an electronic device. Background Art
[0002] Memory failure is an important reason for server downtime. To ensure the stable operation of server hardware, it is necessary to predict the downtime failure of the server.
[0003] There are two types of failures during the operation of the memory, namely CE (Correctable Error, recoverable failure) and UE (Uncorrectable Error, non-recoverable failure). CE is a recoverable failure and does not affect the normal operation of the system for the time being. UE is a non-recoverable memory failure, which usually causes downtime.
[0004] Once a CE Strom (alarm storm) occurs in the memory, within a short period of time, a large number of CE events occur frequently, resulting in the following problems for the server: 1) System performance degradation: A large number of fault correction operations may occupy system resources, leading to a decrease in overall performance; 2) Hardware warning: CE Storm may indicate potential faults in hardware such as memory modules or motherboards, which need to be checked and maintained in a timely manner; 3) Potential data problems: Although CE itself is correctable, frequent failures may mean problems in the data transmission process, and long-term accumulation may lead to non-correctable failures (UE).
[0005] Currently, by polling the register recording CE, the recorded CE information is obtained, and a classifier model is used to extract and classify the features of the obtained CE information to predict whether UE will occur. However, since the information recorded in the register can be overwritten, when a CE Strom occurs, the CE information obtained by polling cannot reflect all the information of the CE Strom, resulting in a low prediction accuracy. Summary of the Invention
[0006] The first aspect of this application provides a method for processing memory failure information, including:
[0007] Triggered by a preset interrupt handler based on the Unified Extensible Firmware Interface (UEFI) to obtain the failure information of the memory, where the failure information characterizes at least one of a recoverable failure and a non-recoverable failure that occurs in the memory;
[0008] Generate target information according to the failure information;
[0009] Upload the target information to the target controller.
[0010] In a possible implementation, when the preset interrupt handler based on UEFI is triggered, fault information of the memory is obtained, including:
[0011] The first interrupt handler based on UEFI records the first occurrence count of recoverable faults in the memory;
[0012] When the first occurrence count reaches a preset count, the first interrupt handler is triggered to read the memory fault register to obtain first fault information.
[0013] In a possible implementation, when the preset interrupt handler based on UEFI is triggered, fault information of the memory is obtained, including:
[0014] When an unrecoverable fault occurs in the memory, the second interrupt handler of UEFI is triggered to read the memory fault register to obtain second fault information.
[0015] In a possible implementation, it further includes:
[0016] When the first occurrence count reaches a preset count, the trigger count of the first interrupt handler within a preset duration is statistically counted;
[0017] According to the trigger count and the preset count, a second occurrence count of recoverable faults in the memory within the preset duration is determined.
[0018] In a possible implementation, it further includes:
[0019] Based on receiving an interrupt trigger instruction from a target controller, the interrupt trigger instruction is generated by the target controller when it determines that an unrecoverable fault is about to occur based on the first fault information;
[0020] Through the first interrupt handler of UEFI, the target memory address to be repaired is obtained according to the interrupt trigger instruction;
[0021] The memory is repaired based on the target memory address, and repair information is fed back to the target controller, where the repair information includes relevant information on the repair of the memory corresponding to the target memory address.
[0022] In a possible implementation, when obtaining the fault information of the memory, it further includes:
[0023] Based on the triggering of the preset interrupt handler, the memory temperature and the processor temperature are obtained.
[0024] A second aspect of this application provides a memory fault prediction method, including:
[0025] Receiving target information sent by UEFI, the target information corresponding to the fault information of the memory, and the fault information indicates that a recoverable fault has occurred in the memory;
[0026] Determine a prediction result according to the target fault rule and the target information, where the prediction result is used to characterize whether an irrecoverable fault is about to occur in the memory.
[0027] In a possible implementation, the determining the prediction result according to the target fault rule and the target information includes:
[0028] Analyze the target information to obtain the information content included in the target information, where the information content includes first fault information, a second number of recoverable faults occurring within a preset duration, repair information, memory temperature, and processor temperature. The first fault information characterizes a recoverable fault occurring in the memory, the repair information includes information related to the repair of the memory, and the memory temperature and the processor temperature correspond to the first fault information;
[0029] Determine a prediction result according to the information content included in the target information and the target fault rule.
[0030] In a possible implementation, it further includes:
[0031] If the target information includes second fault information, update the target fault rule according to the target information, where the second fault information characterizes an irrecoverable fault occurring in the memory.
[0032] A third aspect of this application provides an electronic device, including: a UEFI module and a target controller;
[0033] The UEFI module is provided with a preset interrupt handling program. Triggered based on the preset interrupt handling program, obtain the fault information of the memory, where the fault information characterizes a recoverable fault occurring in the memory; obtain target information according to the fault information; upload the target information to the target controller;
[0034] The target controller is used to determine a prediction result according to the target fault rule and the target information, where the prediction result is used to characterize whether an irrecoverable fault is about to occur in the memory.
[0035] A fourth aspect of this application provides a computer program product, including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement the memory fault information processing method or the memory fault prediction method in the first aspect or any implementation manner of the first aspect.
[0036] A fifth aspect of this application provides an electronic device, including at least one processor and a memory connected to the processor, where:
[0037] The memory is used to store a computer program;
[0038] The processor is configured to execute the computer program, so that the electronic device can implement the memory fault information processing method or the memory fault prediction method according to the first aspect or any implementation manner of the first aspect as described above.
[0039] A sixth aspect of the present application provides a computer storage medium, which carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the memory fault information processing method or the memory fault prediction method according to the first aspect or any implementation manner of the first aspect as described above. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In combination with the accompanying drawings and with reference to the following specific embodiments, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic and the original components and elements are not necessarily drawn to scale.
[0041] Figure 1 is a schematic diagram of the structure of the memory used in the present application Figure 1 ;
[0042] Figure 2 is a schematic diagram of the structure of the memory used in the present application Figure 2 ;
[0043] Figure 3 is a schematic diagram of a failure of a row in the present application;
[0044] Figure 4 is a schematic diagram of a failure of a column in the present application;
[0045] Figure 5 is a schematic diagram of a failure of a cell in the present application;
[0046] Figure 6 is a schematic diagram generated based on failure data when a large number of errors occur in the present application;
[0047] Figure 7 is a schematic flowchart of a memory fault information processing method provided by an embodiment of the present application;
[0048] Figure 8 is a schematic flowchart of obtaining memory fault information by triggering a preset interrupt handler based on UEFI provided by an embodiment of the present application;
[0049] Figure 9 is a schematic flowchart of determining the number of recoverable fault occurrences within a preset duration provided by an embodiment of the present application;
[0050] Figure 10It is another process schematic diagram for obtaining memory failure information provided by an embodiment of the present application;
[0051] Figure 11 It is a process schematic diagram of a memory failure prediction method provided by an embodiment of the present application;
[0052] Figure 12 It is a process schematic diagram for determining a prediction result based on a target failure rule and the target information provided by an embodiment of the present application;
[0053] Figure 13 It is a schematic diagram of an application scenario provided by an embodiment of the present application;
[0054] Figure 14 It is a process schematic diagram for generating a target failure rule provided by an embodiment of the present application;
[0055] Figure 15 It is a schematic diagram of a structure of an electronic device provided by an embodiment of the present application;
[0056] Figure 16 It is another schematic diagram of a structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0057] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the embodiments of the present application are only used to explain the specific embodiments of the present application, rather than intended to limit the present application.
[0058] The embodiments of the present application will be described below with reference to the accompanying drawings. Those skilled in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0059] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0060] The present application can be applied to the memory field. Taking the memory in a server as an example below, the scenario of CE occurring in the memory of the server will be introduced.
[0061] The memory contains multiple DIMMs (Dual Inline Memory Modules), each DIMM consists of multiple ranks (independently addressable logical units on a memory module), one rank consists of multiple device / chip (memory chips), and one chip consists of many banks (the basic unit for actual data storage). Each bank is a storage array, and each cell in the storage array is located by row and column. And each cell can store multiple bits.
[0062] Figure 1 Schematic diagram of the memory structure adopted in this application Figure 1 , Figure 1 shows the relationship between the memory, DIMM, rank, device, and bank, Figure 2 Schematic diagram of the memory structure adopted in this application Figure 2 , Figure 2 shows the relationship between device, bank, cell, and bit.
[0063] The Figure 1-2 In this case, the memory contains 2 DIMMs, one DIMM contains 2 ranks, one rank includes 18 devices, one device contains n banks, one bank has 6×8 cells, and one cell contains 32 bits. The Figure 1 and Figure 2 are only used to illustrate the structural division of the memory, and do not limit the number of subordinate structures contained in each structure in the memory.
[0064] Memory may experience instantaneous errors (faults) due to the influence of cosmic rays, etc., or individual cells may have hard errors (errors caused by hardware failures or damages) due to circuit problems, etc., manifested as always returning a specific value when reading regardless of what data is written. It is also possible that a certain row or column has an error caused by the wordline or bitline, etc.
[0065] Figure 3 is a schematic diagram of a row failure in this application. In the Figure 3 , (a) shows some cells in the bank. The wordline connects a row of cells in series, and the bitline connects a column of cells in series. When an error is caused by the wordline, it is manifested as a specific value always being returned by the row of cells it connects in series. The Figure 3Figure (b) is a graph showing the values returned by each cell in the memory. In this graph, one point represents a cell that has an error, and these points form a row.
[0066] Figure 4 It is a schematic diagram of a column failure in this application. Figure 4 Figure (a) shows some cells in the bank. The wordlines connect a row of cells in series, and the bitlines connect a column of cells in series. When an error is caused by the bitline, it is manifested as a column of cells connected in series fixedly returning a specific value. The Figure 4 Figure (b) is a graph showing the values returned by each cell in the memory. In this graph, one point represents a cell that has an error, and these points form a column.
[0067] Figure 5 It is a schematic diagram of a cell failure in this application. Figure 5 Figure (a) shows some cells in the bank. The wordlines connect a row of cells in series, and the bitlines connect a column of cells in series. When an error is caused by a certain cell, it is manifested as that cell fixedly returning a specific value. The Figure 5 Figure (b) is a graph showing the values returned by each cell in the memory. One point in this graph represents a cell that has an error.
[0068] Figure 6 It is a schematic diagram generated based on the failure data when a large number of errors occur in this application. Each point represents an error occurring in the corresponding cell. Figure 6 In (a), the points are evenly distributed, which may be caused by multiple bitlines or multiple independent cells. Figure 6 In (b), there are fewer points in the upper left half area and more points in the upper right half area. Moreover, the points in the middle column are concentrated, indicating that it is caused by the bitline of the middle column of cells. And a large number of points are concentrated in the upper part of the upper right half area, indicating that the bitlines and wordlines in this area cause failures. Other areas may be caused by individual cells.
[0069] Refer to the above Figure 6 In the case where a large number of errors occur in the memory, the server still operates normally and does not generate a UE. However, whether a UE will occur subsequently needs to be predicted to take corresponding measures to prevent the server from crashing.
[0070] Therefore, the embodiments of this application provide a method for processing memory failure information. The method for processing memory failure information in the embodiments of this application will be introduced in detail below with reference to the accompanying drawings.
[0071] Reference Figure 7 , Figure 7 is a schematic flowchart of a method for processing memory fault information provided by an embodiment of the present application. As Figure 7 shown, a method for processing memory fault information provided by an embodiment of the present application may include steps 701 to 703, and these steps will be described in detail below respectively.
[0072] 701. Triggered by a preset interrupt handler based on UEFI, obtain the fault information of the memory, where the fault information represents at least one of a recoverable fault and an unrecoverable fault that occurs in the memory;
[0073] Among them, for the UEFI (Unified Extensible Firmware Interface) preset interrupt handler, when the trigger condition is met, the preset interrupt handler is triggered.
[0074] In an embodiment of the present application, it may be executed in UEFI. The preset interrupt handler of UEFI is used to trigger the acquisition of the fault information of the memory.
[0075] Among them, in an electronic device, a memory fault register is set for the memory. Each time a fault occurs, the information related to the fault is stored in the fault register. When obtaining the fault information of the memory, it can be read from the fault register to obtain the fault information.
[0076] In a possible implementation, the fault register may adopt a parity register, and an error is determined by reading the value of the register.
[0077] Among them, the fault information of the memory represents that a recoverable fault or an unrecoverable fault occurs in the memory.
[0078] Among them, a recoverable fault is a memory fault that can be recovered. The memory can self-repair the recoverable fault, and the repaired memory can operate normally. Therefore, when a recoverable fault occurs in the memory, it will not temporarily affect the normal operation of the system. However, when a large number of recoverable faults occur, an alarm storm is generated, which will lead to an unrecoverable fault and affect the normal operation of the device.
[0079] In a possible implementation, the preset interrupt handler may be triggered in the case of an alarm storm or in the case of an unrecoverable fault.
[0080] In a possible implementation, the preset interrupt handler may be set to be triggered when the number of recoverable faults reaches a preset number. Each time the preset interrupt handler is triggered, the action of obtaining the fault information of the memory is executed.
[0081] Among them, when an alarm storm occurs and the preset number of times is reached within a very short time, a preset interrupt handler is triggered to perform the action of obtaining the fault information of the memory.
[0082] As an example, the preset number of times is 64. Every time 64 recoverable faults occur, the preset interrupt handler is triggered once to obtain the fault information of the memory. Since the content recorded in the fault register can be overwritten, the fault information of the memory obtained based on the trigger of the preset interrupt handler is the fault information of the 64th recoverable fault. When an alarm storm occurs, thousands of faults may occur in a polling cycle. Reading the content recorded in the fault register according to the polling cycle is likely to result in a large number of missed alarm information. By triggering the preset interrupt handler according to the preset number of times to obtain the fault information of the memory, as much fault information as possible can be obtained.
[0083] Among them, the preset number of times can be set according to the actual situation, and the specific value of the preset number of times is not limited in this application.
[0084] Among them, an unrecoverable fault is an unrecoverable memory fault. Once an unrecoverable fault appears in the memory, it indicates that there is serious physical or logical damage in the memory, which will affect the normal operation of the device. Triggering the preset interrupt handler in the case of an unrecoverable fault can obtain the fault information in time when the unrecoverable fault occurs.
[0085] Among them, the fault information of the memory may include the memory address where the fault occurs, the fault type, and other relevant information about the fault.
[0086] As an example, the fault information may include a timestamp, a slot, a DIMM (Dual In-line Memory Module, a series of modules composed of dynamic random access memory (DRAM)) serial number, relevant information of the hardware register when the fault occurs, detailed location information of the fault, information of the vendor, information of the CPU (central processing unit), etc.
[0087] The process of obtaining the fault information will be described in detail in the subsequent embodiments.
[0088] 702. Generate target information based on the fault information;
[0089] Among them, each content in the obtained fault information is processed to obtain the target information, and the target information is information containing each content.
[0090] Among them, IPMI (Intelligent Platform Management Interface) or the redfish standard can be used to transmit information between the target controller. The redfishe is a RESTful API standard formulated by the DMTF (Distributed Management Task Force). This RESTful API standard is an interface design specification based on the REST (Representational State Transfer) architectural style, used to build the API (Application Programming Interface) of network applications. Its core goal is to achieve efficient, scalable, and loosely coupled communication between different systems.
[0091] Correspondingly, communication between UEFI and the target controller can be achieved by using IPMI or redfish.
[0092] In a possible implementation, each content in the fault information can be format-converted into a standard format for transmission, or into a format that the target controller can parse.
[0093] Among them, when processing the fault information, other forms of optimization can also be performed. In this application, the specific optimization method for obtaining the target information from the fault information is not limited.
[0094] 703. Upload the target information to the target controller.
[0095] Among them, the obtained target information is uploaded to the target controller, and the target controller analyzes and processes the target information to obtain the fault situation in the memory.
[0096] In a possible implementation, the target controller can be a BMC (Baseboard Management Controller) or an XCC (XClarity Controller, a management controller in a server), etc.
[0097] In a possible implementation, a target fault rule is set in the target controller. By using the target fault rule to analyze the target information, a prediction result of whether an irrecoverable fault is about to occur in the memory can be obtained, and it can also be obtained whether an irrecoverable fault has occurred in the memory.
[0098] The subsequent memory fault prediction method executed at the target controller end will elaborate on the prediction process in detail and will not be described in detail here.
[0099] In this embodiment, based on the trigger of the preset interrupt handler of UEFI, fault information of the memory is obtained, and the fault information characterizes at least one of a recoverable fault and an unrecoverable fault that occurs in the memory; according to the fault information, target information is generated; and the target information is uploaded to the target controller. By triggering the preset interrupt handler of UEFI, fault information of the memory is obtained, and the generated target information is uploaded to the template controller. The preset interrupt handler can set the number of occurrences of a recoverable fault as a trigger condition, and when an alarm storm occurs, as much fault information as possible can be obtained. The preset interrupt handler can also set the occurrence of an unrecoverable fault as a trigger condition, and corresponding fault information can be obtained when an unrecoverable fault occurs, so as to transmit as much fault information as possible to the target controller.
[0100] Figure 8 It is a schematic flow chart of obtaining the fault information of the memory based on the trigger of the preset interrupt handler of UEFI provided by the embodiment of the present application, which may include steps 801 to 802, and the following will describe these steps in detail.
[0101] 801. Based on the first interrupt handler of UEFI, record the first number of occurrences of recoverable faults in the memory;
[0102] Among them, the preset interrupt handler of the UEFI adopts the first interrupt handler, and the first interrupt handler is triggered when the preset number of occurrences of a recoverable fault occurs.
[0103] Among them, the first interrupt handler can adopt a UEFI SMI (System Management Interrupt) handler, and the handler is a program for responding to the SMI and can implement the SMI function.
[0104] Among them, after the device is started, the UEFI SMI handler also starts to run, and it records the number of occurrences of recoverable faults in the memory.
[0105] In a possible implementation, a counting register can be set. When a recoverable fault occurs in the memory, the value recorded in the counting register is incremented by 1. Correspondingly, the first interrupt handler records the number of recoverable faults recorded by the counter.
[0106] 802. When the first number reaches the preset number, trigger the first interrupt handler to read the memory fault register to obtain the first fault information.
[0107] Wherein, when the number of occurrences of recoverable faults recorded by the first interrupt handler reaches a preset number, the first interrupt handler reads the memory fault register to obtain the fault information recorded in the memory fault register, and also obtains other fault information corresponding to the fault information recorded in the fault register, such as the memory address and time when a memory error occurs.
[0108] Wherein, the first fault information is fault information related to recoverable faults.
[0109] Wherein, the preset number can be set according to the situation, for example, it can be set to 32, 64, 128, etc. The smaller the value of the preset number, the lower the threshold for triggering the first interrupt handler, and the more fault information can be collected during an alarm storm.
[0110] Wherein, when the number of recoverable faults recorded by the first interrupt handler reaches the preset number, the memory fault register is read, and moreover, the recording times of the first interrupt handler are reset to zero, and the recording times of subsequent recoverable faults start from 0 again.
[0111] As an example, the preset number is set to 32. Every time the number of occurrences of recoverable faults recorded in the memory reaches 32, the memory fault register is read to obtain the first fault information; and when it reaches 32, the recording times of the first interrupt handler are reset to zero, and the recording of the number of occurrences of recoverable faults in the memory starts again, and so on in a cycle.
[0112] In a possible implementation, after triggering the first interrupt handler to obtain the first fault information, the first fault information is generated into target information and uploaded to the target controller, so as to realize the real-time upload of the information content in the first fault information to the target controller, so that the target controller can use the fault information of recoverable faults obtained in real time for non-recoverable fault prediction.
[0113] In this embodiment, the first interrupt handler based on UEFI records the first number of occurrences of recoverable faults in the memory; when the first number reaches the preset number, the first interrupt handler is triggered to read the memory fault register to obtain the first fault information. By setting the first interrupt handler for obtaining the fault information of recoverable faults, the first interrupt handler records the number of occurrences of recoverable faults in the memory, and when the number reaches the preset number, the first interrupt handler reads the memory fault register to obtain the first fault information, so as to obtain the corresponding fault information when a recoverable fault occurs, and obtain as much fault information of recoverable faults as possible according to the preset number as the trigger acquisition condition.
[0114] In a possible implementation, when a preset interrupt handler based on UEFI is triggered, fault information of the memory is obtained, including:
[0115] When an irrecoverable fault occurs in the memory, the second interrupt handler of UEFI is triggered to read the memory fault register to obtain second fault information.
[0116] Among them, the UEFI also sets a second interrupt handler, which is used to be triggered when an irrecoverable fault occurs in the memory, so as to obtain the data in the memory fault register and get the second fault information corresponding to the irrecoverable fault.
[0117] Among them, the second interrupt handler can adopt a UEFI FEH (Firmware Execution Harness) handler, and this handler is a program used to respond to the FEH and can implement the FEH function.
[0118] Among them, when an irrecoverable fault UE occurs in the memory, the UEFI FEH handler is triggered to obtain the fault information of the irrecoverable fault.
[0119] Among them, when generating target information from the second fault information, it can carry relevant information of the UEFI FEH, so that the target controller can set the information source of the target information to the UEFI FEH, providing a basis for subsequent processing using the second fault information.
[0120] Moreover, the second interrupt handler is not affected by the first interrupt handler, nor by the count in the count register used to record recoverable faults, nor by the polling period of polling the fault register.
[0121] In a possible implementation, during the instantaneous process of system crash caused by an irrecoverable fault, before the server restarts, the detailed information of the irrecoverable fault is captured. Moreover, the number of errors occurring can also be calculated through the counter of the FEH, and the detailed information of the irrecoverable fault and the number of errors occurring are used to generate target information to obtain the detailed information of the irrecoverable fault information.
[0122] In a possible implementation, after triggering the second interrupt handler to obtain the second fault information, the second fault information is generated into target information and uploaded to the target controller, so as to realize real-time uploading of the information content in the second fault information to the target controller, so that the target controller can use the fault information of the irrecoverable fault obtained in real time for subsequent processing.
[0123] Among them, the second interrupt handler is further provided with a counter, which is used to calculate the number of occurrences of the non-recoverable fault, so as to upload the number of occurrences of the non-recoverable fault to the target controller together, so that the target controller can update the target fault rule in combination with this number of times. The target fault rule is a rule for predicting whether a non-recoverable fault will occur by using the relevant information of the recoverable fault.
[0124] In this embodiment, when a non-recoverable fault occurs in the memory, the second interrupt handler of UEFI is triggered to read the memory fault register to obtain the second fault information. By setting the second interrupt handler for obtaining the fault information of the non-recoverable fault, the second interrupt handler reads the memory fault register to obtain the second fault information, so as to obtain the corresponding fault information when a non-recoverable fault occurs. This obtaining process is parallel to the fault information of the recoverable fault, and the two do not affect each other.
[0125] Figure 9 It is a schematic flow chart for determining the number of occurrences of recoverable faults within a preset duration provided by an embodiment of the present application, which may include steps 901 to 902. The following will describe these steps in detail.
[0126] 901. When the first number reaches the preset number, count the number of times the first interrupt handler is triggered within the preset duration;
[0127] In addition to Figure 8 the steps for obtaining the memory fault information provided above, the obtaining of the memory fault information may further include obtaining the number of occurrences of the recoverable fault within the preset duration.
[0128] Among them, when the first number of occurrences of the recoverable fault in the memory reaches the preset number, count the number of times the first interrupt handler is triggered within the preset duration.
[0129] Among them, the preset duration can be set according to the actual situation. For example, it can be a duration less than the polling period, or it can be the duration of the polling period, or it can be a duration greater than the polling period. The present application does not limit the value of the preset duration.
[0130] Among them, the preset duration is generally less than the duration of a warning storm, and can determine the second number of recoverable faults multiple times during a warning storm.
[0131] 902. Determine the second number of recoverable faults that occur in the memory within the preset duration according to the number of trigger times and the preset number.
[0132] Among them, according to the number of trigger times and the preset number of times for triggering the first interrupt handler, the second number of recoverable faults that occur in the memory within the preset duration can be determined.
[0133] Wherein, the preset number is multiplied by the trigger number to obtain the second number.
[0134] In a possible implementation, the second number and the fault information read from the fault register can be used as the first fault information, and the first fault information is used to generate target information, and the target information is uploaded to the target controller.
[0135] As an example, the preset duration is 2 seconds, the trigger number of the first interrupt handler within 2 seconds is counted. If the trigger number is 10 times and the preset number is 64, then the number of recoverable faults occurring within 2 seconds is 640 times.
[0136] In this embodiment, when the first number reaches the preset number, the trigger number of the first interrupt handler within the preset duration is counted; according to the trigger number and the preset number, the second number of recoverable faults occurring in the memory within the preset duration is determined, and the second number of recoverable faults occurring in the memory within the preset duration is also counted, so as to upload the second number and the first fault information to the target controller, providing an information basis for the target controller to make a prediction according to the second number and the first fault information.
[0137] Figure 10 It is another schematic flowchart for obtaining memory fault information provided by an embodiment of the present application, which may include steps 1001 to 1003, and the following will describe these steps in detail.
[0138] 1001. Based on receiving an interrupt trigger instruction from the target controller, the interrupt trigger instruction is generated when the target controller determines that an irrecoverable fault is about to occur based on the first fault information;
[0139] Wherein, after receiving the first fault information, when the target controller predicts that the memory is about to have an irrecoverable fault, it generates an interrupt trigger instruction, and the interrupt trigger instruction may carry the memory information to be repaired.
[0140] Wherein, the memory information may be the memory address to be repaired, so as to trigger the repair of the memory.
[0141] 1002. Obtain the target memory address to be repaired through the first interrupt handler of the UEFI according to the interrupt trigger instruction;
[0142] Wherein, when the target controller predicts that an irrecoverable fault occurs in the memory, it can also determine the component structure that needs to be repaired in the memory, and send the address information of the component structure to the first interrupt handler of the UEFI together in the interrupt trigger instruction.
[0143] As an example, when the component structure with an irrecoverable memory failure is a certain row, the target controller sends a signal to the UEFI SMI through a preset pin to trigger the SMI interrupt. Correspondingly, the SMI handler obtains information about the row to be repaired from the target controller.
[0144] Among them, the first interrupt handler receives the interrupt trigger instruction and parses the interrupt trigger instruction to obtain the target memory address to be repaired.
[0145] 1003. Repair the memory based on the target memory address, and feedback the repair information to the target controller. The repair information includes relevant information about the repair of the memory corresponding to the target memory address.
[0146] Among them, after the UEFI parses the target memory address, it triggers the repair process, and the UEFI generates repair information corresponding to this repair and feedbacks the repair information to the target controller. The repair information may include the memory address of this repair, the serial number of the memory, and the time, etc.
[0147] In a possible implementation, the repair process may adopt a PPR (Post Package Repair) process or a page retire process. The specific implementation process of the repair process is not limited in this application.
[0148] Among them, after the repair information is fed back to the target controller, the target controller can eliminate the recoverable failure corresponding to the first failure information based on the repair information. This is because after replacing the row with an error with a spare row, the errors that occurred on the previous row with an error are all invalid, and it is necessary to revoke the influence of the previous errors on that row to avoid the influence of the repaired failure information on subsequent predictions.
[0149] In this embodiment, based on receiving the interrupt trigger instruction from the target controller, the interrupt trigger instruction is generated when the target controller determines that an irrecoverable failure is about to occur based on the first failure information; the first interrupt handler of the UEFI obtains the target memory address to be repaired according to the interrupt trigger instruction; repairs the memory based on the target memory address, and feedbacks the repair information to the target controller. The repair information includes relevant information about the repair of the memory corresponding to the target memory address. When the target controller predicts that an irrecoverable failure is about to occur, it triggers the first interrupt handler of the UEFI. The UEFI obtains the target memory address to be repaired based on the received interrupt trigger instruction to trigger the repair of the memory at the target memory address, and feedbacks the repair information to the target controller, so that the target controller can timely understand this repair and provide more accurate information for subsequent prediction of irrecoverable failures.
[0150] In a possible implementation, obtaining the fault information of the memory further includes:
[0151] Triggered based on the preset interrupt handler, obtain the memory temperature and the processor temperature.
[0152] Wherein, the preset interrupt handler includes a first interrupt handler and a second interrupt handler.
[0153] Wherein, each time the first interrupt handler is triggered, obtain the memory temperature and the processor temperature, so that when the number of recoverable faults that occur in the memory reaches a preset number, the corresponding memory temperature and processor temperature are obtained together.
[0154] Wherein, when an irrecoverable fault occurs, the memory temperature and the processor temperature may reach a very high temperature. Therefore, combining the memory temperature and the processor temperature to predict whether an irrecoverable fault is about to occur in the memory can make the prediction more accurate.
[0155] Correspondingly, when the first interrupt handler is triggered, also trigger obtaining the memory temperature and the processor temperature, which can timely understand the temperatures of the memory and the processor when an alarm storm occurs, and provide another dimension of basis for the subsequent target controller to predict whether an irrecoverable fault is about to occur in the memory.
[0156] In a possible implementation, the processor temperature and the memory temperature can be detected by sensors set at the memory and the CPU, or obtained from the operating parameters. In this application, the specific sources of the memory temperature and the processor temperature are not limited.
[0157] Wherein, each time the second interrupt handler is triggered, also obtain the memory temperature and the processor temperature, and use the temperature of the memory when an irrecoverable fault occurs and the processor temperature as the basis for subsequent analysis of whether an irreparable fault is about to occur for the first fault information.
[0158] In this embodiment, triggered based on the preset interrupt handler, obtain the memory temperature and the processor temperature, which realizes timely acquisition of the memory temperature and the processor temperature, and realizes transmitting as much fault-related information as possible to the target controller, improving the accuracy of the target controller's prediction of whether an irrecoverable fault occurs in the memory.
[0159] The above introduces a memory fault prediction method provided by the embodiments of this application. This method is applied to the UEFI side. The following will introduce the memory fault prediction method applied to the controller side.
[0160] Figure 11 It is a schematic flowchart of a memory fault prediction method provided by the embodiments of this application, asFigure 11 As shown in Figure 11 , a memory fault prediction method provided by an embodiment of the present application may include steps 1101 to 1102, and the following will describe these steps in detail.
[0161] 1101. Receive target information sent by UEFI. The target information corresponds to the fault information of the memory, and the fault information characterizes that a recoverable fault has occurred in the memory.
[0162] Among them, the memory fault prediction method in the embodiment of the present application is applied to a target controller.
[0163] In a possible implementation, the target controller may be a BMC (Baseboard Management Controller) or an XCC (XClarity Controller, a management controller in a server), etc.
[0164] Among them, the target information is transmitted between the target controller and UEFI through IPMI or redfish.
[0165] Among them, after receiving the target information, the target controller may parse the target information to obtain the memory fault information contained therein.
[0166] Among them, the fault information characterizes that a recoverable fault has occurred in the memory.
[0167] In a possible implementation, when the memory controller receives the target information, it may add a corresponding identifier to the target information according to the source of the target information.
[0168] For example, if the target information is obtained by the first interrupt handler of UEFI and uploaded to the target controller, then after receiving the target information, the target controller adds the identifier of the first interrupt handler, such as the UEFI SMI identifier, to the target information.
[0169] Among them, the fault information obtained by the first interrupt handler is the fault information of a recoverable fault occurring in the memory. Correspondingly, the fault information can be used to predict whether an irrecoverable fault will occur in the memory soon.
[0170] 1102. Determine a prediction result according to the target fault rule and the target information. The prediction result is used to characterize whether an irrecoverable fault will occur in the memory soon.
[0171] Among them, a target fault rule is set in the target controller, and the target fault rule is a rule corresponding to an irrecoverable fault about to occur in the memory.
[0172] Among them, the target fault rule can be formed by analyzing historical data through a heuristic algorithm, combining interpretable faults, learning parameter variables, and forming heuristic rules; calculating precision and recall through historical data in combination with heuristic rules, and then selecting heuristic rules to form a rule cluster according to different memory models, processor signals, server models, etc.
[0173] In a possible implementation, the target information is used to screen the rule cluster to determine whether there is a rule in the rule cluster that matches the target information. If there is a rule that matches the target information, it is determined that an irrecoverable fault is about to occur, and the probability of the occurrence of the irrecoverable fault can be calculated.
[0174] Among them, the interpretable fault is determined by communicating with experts.
[0175] In a possible implementation, expert experience can be combined to define a general interpretable memory fault model in the industry, and it is judged whether an interpretable memory fault occurs by analyzing data.
[0176] As an example, analyzing data to determine the occurrence of memory faults can include row errors, column errors, single-bit errors on multiple memory grains, etc.
[0177] In a possible implementation, expert experience can also be combined to define variables for each interpretable fault.
[0178] As an example, variables of interpretable faults can include the distance (Span) where the memory error occurs, etc.
[0179] Among them, the interpretable memory fault model can include the following models: the same row of multiple devices has an error; the same row of multiple devices has an error and the total number of cells with errors is greater than 9; the same column of multiple devices has an error; if the same bank of multiple devices has an error; only one of the 32 bits of a cell has a fault, all the error bits in a cell are concentrated in one column, all the error bits in a cell are concentrated in one row, all the error bits are concentrated in the upper half 16 bits or the lower half 16 bits, all the error bits in a cell are concentrated in a 2×2 block, all the error bits in a cell are concentrated in two non-adjacent columns, etc. The above only lists some of the interpretable memory fault models. In specific implementations, other models can be defined according to actual situations, and the specific content of the interpretable memory fault model is not limited in this application.
[0180] In a possible implementation, combining expert experience to define a general interpretable memory fault model in the industry is determined based on historical irrecoverable fault information, which can be obtained by the second interrupt handler of UEFI when an irrecoverable memory fault occurs.
[0181] For the specific explanation of how the second interrupt handler obtains the irrecoverable fault information, please refer to the explanation in the foregoing embodiments of the memory fault information processing method.
[0182] In a possible implementation, when determining the prediction result, in addition to the target information, relevant information of the memory and information such as the platform of the electronic device (such as a server) where it is located can also be obtained as the basis for prediction.
[0183] In this embodiment, the target information sent by UEFI is received. The target information corresponds to the fault information of the memory, and the fault information indicates that a recoverable fault has occurred in the memory. According to the target fault rule and the target information, a prediction result is determined, and the prediction result is used to indicate whether an irrecoverable fault is about to occur in the memory. By obtaining the information that a recoverable fault has occurred in the memory through UEFI and using the target fault rule and this information to predict that an irrecoverable fault is about to occur in the memory, it realizes predicting whether an irrecoverable fault is about to occur in the memory using the recoverable fault information.
[0184] Figure 12 The flowchart for determining the prediction result according to the target fault rule and the target information provided by the embodiments of the present application may include steps 1201 to 1202, and the following will describe these steps in detail.
[0185] 1201. Analyze the target information to obtain the information content included in the target information. The information content includes first fault information, a second number of recoverable faults occurring within a preset duration, repair information, memory temperature, and processor temperature. The first fault information indicates that a recoverable fault has occurred in the memory, the repair information includes relevant information on memory repair, and the memory temperature and the processor temperature correspond to the first fault information.
[0186] Among them, the first fault information may include: timestamp, slot, memory serial number, information of the fault register when the fault occurs, detailed location information of the fault in the memory, information of the service provider, and CPU information, etc.
[0187] Among them, the first fault information is the information of a recoverable fault in the memory obtained by the first interrupt handler in UEFI, and the specific obtaining method can refer to the explanation in the foregoing embodiments of the memory fault information processing method.
[0188] Among them, the second number is calculated based on the number of times the first interrupt handler in UEFI is triggered and a preset number of times being triggered.
[0189] Among them, the memory temperature and the processor temperature are obtained when the first interrupt handler is triggered.
[0190] Among them, the repair information is information related to repairing the memory. When the target controller predicts that a UE is about to occur based on the first fault information of the memory combined with the memory temperature and the processor temperature, and there is a structure (such as a row) that needs to be repaired, it triggers the first interrupt handler to interrupt through the GPIO pin. The first interrupt handler obtains the target memory address of the memory to be repaired from the target controller, starts the repair, and uploads the repair information to the target controller. The target controller records the repair information in the local database and uses the repair information to revoke the fault of the repaired structure to eliminate the fault information corresponding to the repaired structure and prevent the historical fault information of the repaired structure from affecting the prediction.
[0191] 1202. Determine the prediction result according to the information content included in the target information and the target fault rule.
[0192] Among them, according to the received target information, use the information content it contains to predict whether the memory is about to have an irrecoverable fault.
[0193] First, the repair information can be used to revoke the fault information of the repaired structure in the first fault information to eliminate the fault information corresponding to the repaired structure and prevent the historical fault information of the repaired structure from affecting the prediction.
[0194] Then, use the processed fault information and the target fault rule to predict that the memory release is about to have an irrecoverable fault.
[0195] In a possible implementation, a memory platform is also set in the electronic device, and structures such as the memory and the memory fault register are set on the memory platform. The memory platform can poll the memory fault register according to the polling period, obtain the fault information, and upload the polled fault information to the target controller. Correspondingly, the target controller will also analyze the polled fault information and the fault information obtained by the first interrupt handler interruption together to increase the amount of data for analysis.
[0196] Among them, the target controller receives the target information in real time and processes it in real time.
[0197] As an example, if the target information of the row1 failure is received from 12:00 to 14:00 and the repair information indicating that row1 has been repaired and replaced is received at 15:00, then when the repair information is received at 15:00, the impact of the row1 failure received from 12:00 to 14:00 should be revoked.
[0198] Among them, after receiving the target information, the target controller also classifies and stores the target information in the database.
[0199] In a possible implementation, the received first failure information is stored corresponding to the memory temperature, processor temperature, etc.; the repair information is recorded separately.
[0200] As an example, 64 errors on a certain platform will trigger an SMI (i.e., the first interrupt handler) interrupt. In the past 1 second, 100 SMI interrupts have occurred. Then, 64 * 100 = 6400 memory errors have occurred in the past second. Mark the source of the memory failure as the UEFI SMI Handler, the memory error as CE, the number of memory errors as 6400, the memory temperature as 56 °C, and the CPU temperature as 120 °C.
[0201] In a possible implementation, the target controller can also directly obtain the memory temperature, processor temperature, etc. through the platform without the UEFI side obtaining the memory temperature and processor temperature.
[0202] Among them, when the target controller receives the information uploaded by the first interrupt handler in the UEFI, it triggers the acquisition of the memory temperature and processor temperature.
[0203] In this embodiment, the target information is analyzed to obtain the information content included in the target information. The information content includes the first failure information, the second number of recoverable failures occurring within a preset duration, the repair information, the memory temperature, and the processor temperature. The first failure information represents a recoverable failure occurring in the memory. The repair information includes information related to the memory repair. The memory temperature and the processor temperature correspond to the first failure information. According to the information content included in the target information and the target failure rule, the prediction result is determined. By analyzing the target information, the information content included in the target information is obtained, and then combined with the target failure rule to determine whether the memory is about to have an irrecoverable failure, and the prediction result is obtained. In this process, analysis and prediction are carried out from multiple dimensions, and the obtained prediction result has a high accuracy.
[0204] In a possible implementation, the memory failure prediction method further includes:
[0205] If the target information includes second failure information, the target failure rule is updated according to the target information. The second failure information represents an irrecoverable failure occurring in the memory.
[0206] Among them, the target information includes a second fault, indicating that the UEFI also uploads the second fault information of the irrecoverable fault that occurs in the memory to the target controller, so that the target controller can update the target fault rule by combining the fault information of the irrecoverable fault that occurs in the memory and the relevant information.
[0207] Among them, after receiving the target information, the target controller also classifies and stores the target information in the database.
[0208] In a possible implementation, the received second fault information and the corresponding memory temperature, processor temperature, etc. are stored correspondingly.
[0209] In a possible implementation, when an irrecoverable fault occurs, the second interrupt handler of the UEFI will be triggered, record the detailed information of the irrecoverable fault in time, and transmit it to the target controller, and the target controller can mark the target information uploaded from the second interrupt handler.
[0210] Among them, the target controller also obtains the memory temperature and processor temperature when the irrecoverable fault occurs, and updates the target fault rule by using the memory temperature and processor temperature in combination with the second fault information.
[0211] In this embodiment, if the target information includes the second fault information indicating an irrecoverable fault in the memory, the target fault rule is updated by using the target information, so as to realize the update and maintenance of the target fault rule by using the relevant information of the irrecoverable fault that occurs in the memory, so as to improve the prediction accuracy of using the target fault rule.
[0212] Figure 13 It is a schematic diagram of an application scenario provided in an embodiment of the present application. The application scenario includes a UEFI module 1301, a target controller 1302, and a platform 1303. The platform 1303 can manage a memory 1304.
[0213] Among them, the UEFI module 1301 includes a first interrupt handler (UEFI SMI handler) and a second interrupt handler (UEFI FEH handler). The first interrupt handler is set to be triggered when a preset number of recoverable faults occur in the memory. When the first interrupt handler is triggered, it reads the relevant information of the recoverable fault recorded in the fault register. The second interrupt handler is set to be triggered when an irrecoverable fault occurs in the memory. The second interrupt handler is used to read the relevant information of the irrecoverable fault recorded in the fault register. The UEFI module processes the fault information obtained by the two interrupt handlers into target information and uploads it to the target controller 1302 in real time.
[0214] Among them, the target controller 1302 includes a prediction module 13021 and a database 13022. The database is used to store the obtained information related to faults. The information related to faults includes the information related to recoverable faults and the information related to non-recoverable faults. A target fault rule is preset in the prediction module 13021.
[0215] Among them, when the target controller receives the information related to a memory fault, it sends the information to the prediction module 13021 and the database 13022 respectively. The prediction module 13021 is used to predict whether an irrecoverable memory fault is about to occur based on the information related to recoverable faults; the prediction module 13021 can also update the target fault rule based on the information related to non-recoverable faults.
[0216] Among them, when the target controller receives fault information, it accesses other services of the target controller through a D-Bus (an inter-process communication (IPC) system) interface to obtain the memory temperature and the processor temperature to assist in processing the corresponding fault information.
[0217] The prediction module can be set as a dynamic library, which can implement memory prediction fault analysis (Memory Predict Failure Analysis-lib) and storage of the received data.
[0218] Among them, when the target controller obtains the information related to a non-recoverable fault, it sends it to the prediction module through the interface with the prediction module, so that the information of the non-recoverable fault can be recorded and sent in time, so as to be processed in time, without being affected by the polling period and the preset number of times of the first interrupt handler.
[0219] Among them, the platform 1303 uploads the fault information obtained by polling to the target controller; the number of recoverable faults occurring in the platform is recorded, and when the preset number is met, the first interrupt handler is triggered to interrupt. When the first interrupt handler is triggered, it reads the fault register to obtain the information related to the recoverable fault recorded in the fault register at the triggering moment; moreover, when an irrecoverable memory fault occurs in the platform 1303, the second interrupt handler is triggered to interrupt. When the second interrupt handler is triggered, it reads the fault register to obtain the information related to the non-recoverable fault.
[0220] Among them, a display interface for the prediction result of the preset module can be set, and the prediction result is displayed in the display interface. The prediction result can include whether an irrecoverable fault is about to occur, the prediction probability, the matching fault rule, etc., and can also select to display detailed fault information, etc., so that the user can timely and detailedly understand the prediction result of this prediction and the basis for the prediction, etc.
[0221] Figure 14 It is a schematic flow chart for generating a target fault rule provided by an embodiment of the present application, including a heuristic rule part 1041 and a repair part 1402.
[0222] Among them, in the heuristic rule part 1401, the original data is analyzed to determine the fault types therein, such as row fault (RowFault), column fault (ColumnFault), or bank fault, etc., and various faults are extracted as parameters; the features are combined according to the heuristic algorithm to determine various fault combinations to obtain unrecoverable faults, and a memory fault rule model is obtained.
[0223] Among them, in the repair part 1402, the influence in the original data is eliminated through the repair information PPR, the memory fault rule model is obtained by learning from experts, the memory fault rule model is trained using on-site fault data, the rules that meet the requirements of precision and recall are screened, and the screened rules are fed back to the memory fault rule model and the model parameters are adjusted.
[0224] Among them, several rules are included in the memory fault rule model. During the actual prediction process, the rules of the memory fault rule model can be updated using the unrecoverable fault information in the actual process.
[0225] An embodiment of the present application also provides a process for determining a memory fault rule model:
[0226] 1. Collect the repair information PPR and eliminate the influence of PPR on the current memory fault.
[0227] For example, PPR has repaired row 0x354, and all the errors that occurred in this row 0x354 in history are repaired and reset to the normal value.
[0228] 2. Adopt expert experience to define a general interpretable memory fault model in the industry, and determine interpretable memory faults by analyzing historical fault information to form a set of basic fault types.
[0229] Among them, the set of basic fault types can be represented as F = {F1, F2, …, Fn}, where n is an integer greater than 1.
[0230] For example: F = {RowFault, ColumnFault, …}
[0231] 3. Adopt expert experience to define parameter variables for each fault.
[0232] Among them, variables for randomness can be defined. For example, the distance (ColSpan) at which a memory fault occurs with an error on a row. For instance, a memory fault occurs with a random error on a row, such as an intermittent error, and it is not necessary for all columns of a row to have an error.
[0233] Through feature extraction in machine learning, features can be extracted according to each industrially interpretable error, the feature importance can be found, and by sorting according to the degree of importance, parameters for each interpretable fault can be defined.
[0234] Among them, memory faults can be defined according to the memory addresses where recoverable faults occur. For each fault type, Fi is defined, and the fault variable combination is defined as:
[0235] V i ={v i1 ,v i2 ,…, v imi}, where i is an integer greater than 0.
[0236] For example: V RowFault ={α,β}, V ColumnFault ={γ}, and the value spaces of these variables are α∈{0,1}, β∈{0,1}, γ∈{0,1}
[0237] 4. Through a heuristic algorithm, combine interpretable faults, learn parameter variables from historical non-recoverable fault data to form heuristic rules, and calculate precision and recall through historical non-recoverable fault data.
[0238] First, define the rules.
[0239] Each rule can be defined as
[0240] Among them, , S is a subset of the fault types, C i is the variable condition of the fault type F i , for example, C i can be α = 1 and β = 0
[0241] Then, generate the rules.
[0242] For a single fault rule, for each fault type F i , all possible variable combinations Ci can be generated.
[0243] For example: C RowFault,1 : α = 1; C RowFault,2 : β = 1; C RowFault,3 : α = 1 ∧ β = 1.
[0244] For multiple fault combination rules, various faults can be combined.
[0245] For example: R1: (RowFault AND α = 1) ∧ (ColumnFault AND γ = 1), where γ is a variable.
[0246] For the combination strategy, a recursive or iterative method can be adopted to generate all possible rule combinations, or generate a partial rule set according to prior knowledge and combination complexity constraints.
[0247] 5. Evaluate the precision and recall of each rule.
[0248] To calculate the precision and recall of each rule, the following formulas can be used:
[0249]
[0250]
[0251] Among them, TP (True Positive, true positive example) refers to the number of samples that are actually positive examples and are correctly predicted as positive examples by the model; FP (False Positive, false positive example) refers to the number of samples that are actually negative examples but are wrongly predicted as positive examples by the model; FN (False Negative, false negative example) refers to the number of samples that are actually positive examples but are wrongly predicted as negative examples by the model. A positive example is a rule that can characterize the occurrence of an irrecoverable fault, and a negative example is a rule that cannot characterize the occurrence of an irrecoverable fault.
[0252] 6. Use the precision and recall of each rule to screen out the rules with high weights.
[0253] A threshold can be preset to screen out the rules whose precision and recall are greater than the set threshold as the target fault rules.
[0254]
[0255] Among them, R represents a rule, and are the set thresholds.
[0256] Through the above formal heuristic rule mining algorithm, the basic memory fault types containing multiple variables can be systematically combined and evaluated, and the rules that can efficiently predict irrecoverable faults can be mined.
[0257] The above introduces a memory fault prediction method and a memory fault information processing method provided by the embodiments of the present application. Next, an electronic device for executing the above memory fault prediction method and memory fault information processing method will be introduced.
[0258] Please refer to Figure 15 , Figure 15 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 15 shown, the electronic device 1500 includes: a UEFI module 1501 and a target controller 1502;
[0259] The UEFI module 1501 is provided with a preset interrupt handler. Triggered based on the preset interrupt handler, fault information of the memory is obtained. The fault information characterizes that a recoverable fault has occurred in the memory; according to the fault information, target information is obtained; and the target information is uploaded to the target controller;
[0260] The target controller 1502 is configured to determine a prediction result according to a target fault rule and the target information, and the prediction result is used to characterize whether an irrecoverable fault is about to occur in the memory.
[0261] Among them, for the function explanation of the UEFI module, please refer to the explanation in the foregoing embodiment of the memory fault information processing method. For the function explanation of the target controller, please refer to the explanation in the foregoing embodiment of the memory fault information processing method, which will not be elaborated here.
[0262] In this embodiment, the UEFI module receives the target information sent by the UEFI. The target information corresponds to the fault information of the memory, and the fault information characterizes that a recoverable fault has occurred in the memory; the target controller determines a prediction result according to the target fault rule and the target information, and the prediction result is used to characterize whether an irrecoverable fault is about to occur in the memory. By obtaining the information that a recoverable fault has occurred in the memory through the UEFI, and using the target fault rule and this information to predict that an irrecoverable fault is about to occur in the memory release, it realizes predicting whether an irrecoverable fault is about to occur in the memory by using the recoverable fault information.
[0263] An embodiment of the present application also provides an electronic device. Refer to Figure 16 shown, which shows a schematic structural diagram of an electronic device suitable for implementing the electronic device in the embodiment of the present application. The electronic device in the embodiment of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like. Figure 16 The electronic device shown is only an example and should not bring any limitation to the functions and usage scope of the embodiment of the present application.
[0264] As Figure 16As shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 1601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1602 or a program loaded from a storage device 1608 into a random access memory (RAM) 1603. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 1603. The processing device 1601, the ROM 1602, and the RAM 1603 are connected to each other through a bus 1604. An input / output (I / O) interface 1605 is also connected to the bus 1604.
[0265] Generally, the following devices may be connected to the I / O interface 1605: an input device 1606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1608 including, for example, a memory card, a hard disk, etc.; and a communication device 1609. The communication device 1609 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 16 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0266] In an embodiment of the present application, there is also provided a computer program product including computer-readable instructions. When the computer-readable instructions run on an electronic device, the electronic device is enabled to implement any one of the memory fault prediction methods / a memory fault information processing method provided in the embodiment of the present application.
[0267] In an embodiment of the present application, there is also provided a computer-readable storage medium. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can be enabled to implement any one of the memory fault prediction methods / a memory fault information processing method provided in the embodiment of the present application.
[0268] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the present application, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0269] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits, etc. However, for the present application, in more cases, software program implementation is a better embodiment. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk, or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0270] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0271] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a dedicated computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, training device, or data center to another website, computer, training device, or data center in a wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that a computer can store, or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. A method for processing memory fault information, comprising: Triggered by a preset interrupt handler based on the Unified Extensible Firmware Interface (UEFI) to obtain the fault information of the memory, where the fault information characterizes at least one of a recoverable fault and an unrecoverable fault occurring in the memory; Generate target information according to the fault information; Upload the target information to a target controller.
2. The method for processing memory fault information according to claim 1, wherein the triggered by the preset interrupt handler based on UEFI to obtain the fault information of the memory comprises: Based on the first interrupt handler of UEFI, record the first occurrence count of recoverable faults in the memory; When the first occurrence count reaches a preset count, trigger the first interrupt handler to read the memory fault register to obtain the first fault information.
3. The method for processing memory fault information according to claim 1, wherein the triggered by the preset interrupt handler based on UEFI to obtain the fault information of the memory comprises: When an unrecoverable fault occurs in the memory, trigger the second interrupt handler of UEFI to read the memory fault register to obtain the second fault information.
4. The method for processing memory fault information according to claim 2, further comprising: When the first occurrence count reaches a preset count, count the trigger times of the first interrupt handler within a preset duration; Determine the second occurrence count of recoverable faults occurring in the memory within the preset duration according to the trigger times and the preset count.
5. The method for processing memory fault information according to claim 2, further comprising: Based on receiving an interrupt trigger instruction from the target controller, where the interrupt trigger instruction is generated by the target controller when it determines that an unrecoverable fault is about to occur based on the first fault information; Obtain the target memory address to be repaired through the first interrupt handler of UEFI according to the interrupt trigger instruction; Repair the memory based on the target memory address and feedback the repair information to the target controller, where the repair information includes relevant information on repairing the memory corresponding to the target memory address.
6. The method for processing memory fault information according to any one of claims 2-3, wherein obtaining the fault information of the memory further comprises: Based on the trigger of the preset interrupt handler, obtain the memory temperature and the processor temperature.
7. A method for predicting memory faults, comprising: Receive the target information sent by UEFI, where the target information corresponds to the fault information of the memory, and the fault information characterizes a recoverable fault occurring in the memory; Determine a prediction result according to the target fault rule and the target information, where the prediction result is used to characterize whether an unrecoverable fault is about to occur in the memory.
8. The method for predicting memory faults according to claim 7, wherein determining the prediction result according to the target fault rule and the target information comprises: Analyze the target information to obtain the information content included in the target information. The information content includes first fault information, a second number of recoverable faults occurring within a preset duration, repair information, memory temperature, and processor temperature. The first fault information indicates that a recoverable fault has occurred in the memory. The repair information includes information related to the repair of the memory. The memory temperature and the processor temperature correspond to the first fault information; Determine a prediction result based on the information content included in the target information and the target fault rule.
9. The memory fault prediction method according to claim 8 further includes: If the target information includes second fault information, update the target fault rule according to the target information. The second fault information indicates that an irrecoverable fault has occurred in the memory.
10. An electronic device, comprising: UEFI module and target controller; The UEFI module is provided with a preset interrupt handler. Based on the trigger of the preset interrupt handler, obtain the fault information of the memory. The fault information indicates that a recoverable fault has occurred in the memory. Obtain the target information according to the fault information; Upload the target information to the target controller; The target controller is configured to determine a prediction result according to the target fault rule and the target information. The prediction result is used to indicate whether an irrecoverable fault is about to occur in the memory.
Citation Information
Cited By
Fault processing method and device, medium and program product
CN120524182A