Method, apparatus, device, and storage medium for handling memory faults
By analyzing historical fault information, real-time detection and replacing fault rows with redundant rows or redundant banks, the system downtime caused by cold reset is solved, and fault repair without cold reset is achieved, and system stability is improved.
Patent Information
- Application Number
- CN202011179463.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-06-20
- Filing Date
- 2020-10-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2040-10-29
AI Technical Summary
In the prior art, memory failures require cold reset of equipment to be repaired, resulting in business interruption and affecting system stability.
By analyzing historical fault information, detect memory failures in real time and replacing fault rows with redundant rows or redundant banks, fault repair without cold reset is achieved.
Repair memory failures in a timely manner, prevent system downtime, reduce business impacts, and improve system stability.
Smart Images

Figure CN113821364B_ABST
Abstract
Description
[0001] This application claims the priority of a Chinese patent application with the application number 202010569797.2 and the invention title "A Method for Memory Fault Handling" filed on June 20, 2020, the entire content of which is incorporated herein by reference. Technical Field
[0002] Embodiments of this application relate to the field of computer technologies, and particularly to a method, apparatus, device, and storage medium for handling memory faults. Background Art
[0003] Memory is one of the important components of a device. Generally, memory includes multiple banks (also known as memory matrices), and each bank includes multiple memory rows. During the use of memory, faults often occur for various reasons, and memory faults caused by memory row faults account for a high proportion. Therefore, the repair of memory row faults can be an important repair means for memory faults.
[0004] In the related art, there are redundant rows on each bank in memory. After the device performs a cold reset, that is, after the device crashes and restarts or the user manually restarts the device, the device will perform a memory self-check. If it is detected that a memory fault occurs on a memory row, it is considered that a memory row fault has occurred, and the memory row where the fault occurs is called a fault row. At this time, it can be determined whether the number of memory faults of the correctable error (CE) type occurring on the fault row recorded in the fault log reaches a threshold. If the threshold is reached, it is determined that the condition for starting hard postpackage repair (hPPR) is currently met, hPPR is started, and the fault row is replaced with the redundant row on the bank where the fault row is located, thereby achieving the repair of the memory row fault.
[0005] However, in the related art, hPPR needs to be started after the device performs a cold reset to replace the fault row, which will affect the service. If memory faults are serious during the operation of the device and are not repaired all the time, it will cause the device to crash, which will seriously affect the service. Summary of the Invention
[0006] Embodiments of this application provide a method, apparatus, device, and storage medium for handling memory faults, which can repair memory faults in a timely manner, prevent the system from crashing, and reduce the impact on the service. The technical solutions are as follows:
[0007] In a first aspect, a method for handling memory faults is provided. The method includes:
[0008] Start the fault analysis of the memory at the first moment; the fault analysis includes: obtaining the current fault analysis result of the memory by analyzing historical fault information, where the historical fault information is the fault information accumulated by the memory within a historical time period, and the historical time period is the time period before the first moment or the time period before and including the first moment; start the fault repair of the memory according to the current fault analysis result of the memory.
[0009] In the embodiment of the present application, the fault analysis result is obtained by analyzing historical fault information, and then the memory is repaired according to the fault analysis result. This solution can analyze the memory fault more precisely, and can start the fault repair of the memory without cold reset, prevent system downtime, and reduce the impact on services.
[0010] Optionally, the first moment is the moment before the UCE fault occurs in the computer system. That is, start the fault analysis of the memory during the operation of the computer system, and the operation period of the computer system refers to the period when the computer system is working properly.
[0011] Optionally, the first moment includes: the moment periodically started according to a preset condition; and / or, the moment when it is determined that a memory fault occurs in the memory after the computer system runs.
[0012] That is, when the computer device detects that a memory fault occurs, it starts to analyze the historical fault information to obtain the fault analysis result. Or, the computer device periodically analyzes the historical fault information to obtain the fault analysis result. Or, the computer device periodically analyzes the historical fault information to obtain the fault analysis result, and if a memory fault is detected within the periodic interval, it analyzes the historical fault information to obtain the fault analysis result, and starts a new periodic analysis based on the time when the memory fault is detected this time. Or, the computer device periodically analyzes the historical fault information to obtain the fault analysis result, and if a memory fault is detected within the periodic interval, it analyzes the historical fault information to obtain the fault analysis result, but does not start a new periodic analysis based on the time when the memory fault is detected this time, that is, it does not affect the periodic analysis.
[0013] It should be noted that the computer device periodically analyzing the historical fault information can timely predict the severity of the memory fault and repair the memory fault in time.
[0014] Optionally, in the embodiment of the present application, the historical fault information is analyzed by a fault analysis model to determine the fault analysis result. That is, the computer device obtains the current fault analysis result of the memory by analyzing the historical fault information, including: inputting the historical fault information into the fault analysis model to obtain the current fault analysis result of the memory, and the fault analysis model is an intelligent computing analysis model.
[0015] It should be noted that analyzing historical fault information through the fault analysis model is only one implementation manner for analyzing historical fault information provided by the embodiments of the present application. The computer device can also analyze historical fault information through other implementation manners, such as a data statistics-based manner, and the embodiments of the present application do not limit this. Next, the implementation manner in which the computer device obtains the fault analysis result through the fault analysis model or through other means will be introduced.
[0016] In the embodiments of the present application, if the fault analysis result includes a fault mode, then the computer device starts the fault repair of the memory according to the current fault analysis result of the memory, including: when the fault mode is a memory row fault, starting the fault repair of the memory, where the fault repair includes: replacing the faulty row with a redundant row and repairing the data on the redundant row.
[0017] In the embodiments of the present application, the computer device obtains the current fault analysis result of the memory, including: obtaining a first statistical feature according to the historical fault information, where the first statistical feature represents the number of faulty bits that have occurred in the first memory row during a historical time period, and the first memory row is any memory row. When the first statistical feature is greater than a first threshold, it is determined that the fault mode is a memory row fault, and the first threshold represents the number of faulty bits that each memory row can tolerate.
[0018] Optionally, assuming that the computer device analyzes the historical fault information through the fault analysis model, then the fault analysis model includes a first threshold, and the computer device inputs the historical fault information into the fault analysis model, and the fault analysis model obtains the first statistical feature according to the historical fault information.
[0019] Optionally, if the fault analysis result further includes a fault level, then the computer device starts the fault repair of the memory according to the current fault analysis result of the memory, including: when the fault mode is a memory row fault and the fault level is a high-risk level, starting the fault repair of the memory.
[0020] Optionally, the computer device obtaining the current fault analysis result of the memory further includes: obtaining a second statistical feature and / or a third statistical feature according to the historical fault information, where the second statistical feature represents the number of faults of each fault type that have occurred in the first memory row during a historical time period, and the third statistical feature represents the number of error corrections that have occurred in the first memory row during a historical time period; when the second statistical feature is greater than a second threshold, or when the third statistical feature is greater than a third threshold, or when the second statistical feature is greater than the second threshold and the third statistical feature is greater than the third threshold, it is determined that the fault level is a high-risk level. Wherein, the second threshold represents the number of faults of each fault type that each memory row can tolerate, and the third threshold represents the number of error corrections that each memory row can tolerate.
[0021] Optionally, assuming that the computer device analyzes historical fault information through a fault analysis model, the fault analysis model further includes a second threshold and / or a third threshold. The computer device inputs the historical fault information into the fault analysis model, and the fault analysis model obtains a second statistical feature and / or a third statistical feature based on the historical fault information.
[0022] It should be noted that the historical fault information further includes the fault type and fault correction information of memory faults that occurred during the historical time period. Among them, the fault type includes the CE type and the UCE type. Optionally, the CE type includes the inspection CE type, the read CE type, etc. The fault correction information includes information such as the amount of error correction data (also known as error correction data, with a unit such as bit) and error correction codes for error correction of each memory fault sent.
[0023] Optionally, risk mode options are displayed on the interaction interface, and the risk mode options include a memory high-risk mode option and a memory low-risk mode option. That is, the computer device provides an interaction interface, and the user can select a risk mode through the interaction interface.
[0024] Optionally, the first threshold, the second threshold, and the third threshold are variables set according to the risk mode.
[0025] Optionally, the first threshold of the memory high-risk mode is less than the first threshold of the memory low-risk mode; and / or, the second threshold of the memory high-risk mode is less than the second threshold of the memory low-risk mode; and / or, the third threshold of the memory high-risk mode is less than the third threshold of the memory low-risk mode.
[0026] Optionally, the duration of the historical time period is a variable set according to the risk mode, and the duration of the historical time period of the memory high-risk mode is less than the duration of the historical time period of the memory low-risk mode.
[0027] As can be seen from the above, the user can flexibly select a risk mode according to needs. For example, if the user's business risk is relatively high, the high-risk mode can be selected. In this way, the first threshold and / or the second threshold and / or the third threshold are lower and / or the historical time period is shorter. The computer device analyzes the historical fault information within a shorter time period to obtain the first statistical feature, the second statistical feature, and / or the third statistical feature, and compares the obtained data with smaller thresholds to analyze whether it is a memory row fault and a high-risk level. In this way, the computer device can ensure timely identification of less serious memory row faults. If the user's business risk is relatively low, the low-risk mode can be selected, which can ensure high recognition, that is, timely identification of more serious memory row faults.
[0028] In an embodiment of the present application, the computer device provides an interaction interface for the user to select a risk mode. The computer device determines the duration of the fault information to be analyzed and / or the threshold value when making a threshold judgment according to the risk mode selected by the user. By counting the fault information within the corresponding duration and performing a threshold comparison, when it is identified that the fault mode is a memory row fault, the memory fault is repaired in a timely manner. In this way, by integrating the risk mode selected by the user with the method of threshold comparison, while accurately predicting the memory row fault, the computing pressure on the computer device is reduced.
[0029] In an embodiment of the present application, as can be seen from the foregoing, the fault analysis result includes a fault mode. Then, the computer device starts the fault repair of the memory according to the current fault analysis result of the memory, including: when the fault mode is a memory row fault, starting the fault repair of the memory, where the fault repair includes: replacing the faulty row with a redundant row and repairing the data on the redundant row. That is to say, when the computer device determines that the fault mode is a memory row fault, it replaces the faulty row with a redundant row in the memory and repairs the faulty data.
[0030] Alternatively, as can be seen from the foregoing, the fault analysis result further includes a fault level. Then, the computer device starts the fault repair of the memory according to the current fault analysis result of the memory, including: when the fault mode is a memory row fault and the fault level is a high-risk level, starting the fault repair of the memory. That is to say, when the computer device determines that the fault mode is a memory row fault and the fault level is a high-risk level, it replaces the faulty row with a redundant row in the memory and repairs the faulty data.
[0031] Optionally, the redundant row and the faulty row are located on the same bank in the memory. That is to say, the computer device replaces the faulty row with a redundant row on the bank where the faulty row is located.
[0032] Optionally, the computer device repairs the data on the redundant row, including: performing a read operation on the redundant row; if the data read from the redundant row is incorrect data, correcting the incorrect data and writing the corrected data back to the redundant row to achieve the repair of the data on the redundant row. That is to say, in an embodiment of the present application, the faulty data is repaired through the read operation of the redundant row and the data write-back.
[0033] Optionally, a read operation is performed on the redundant row. If the data read from the redundant row is incorrect data, the incorrect data is corrected, and the corrected data is written back to the redundant row, including: dividing the redundant row into M segments, each segment including one or more storage units, where M is an integer greater than 1; setting i = 1, and performing a read operation on the i-th segment of the redundant row; if the data read from the i-th segment of the redundant row is incorrect data, the incorrect data is corrected, and the corrected data is written back to the i-th segment; if i is not equal to M, then set i = i + 1, and return to perform a read operation on the i-th segment of the redundant row until i is equal to M. That is, the computer device repairs the data on the redundant row by means of segment-by-segment successive reading, correction, and writing back.
[0034] Optionally, after the data read from the redundant row is incorrect data, the method further includes: generating a correctable error CE; suppressing the CE.
[0035] In the embodiments of the present application, after the data read from the redundant row is incorrect data, a CE is generated in the computer device, and the computer device suppresses the CE. That is, since the computer device detects incorrect data when reading the redundant row, the computer device will consider that a CE has been detected. Since this CE is not caused by a memory failure of the computer, it is necessary to suppress the CE, that is, not process the CE, or in other words, the computer device does not record the CE.
[0036] Optionally, after the data on the redundant row is repaired, the method further includes: releasing the suppression operation of the CE.
[0037] The CE generated by the computer device after repairing the redundant row is caused by a real memory failure. Therefore, it is necessary to process the CE, that is, release the suppression operation of the CE and record the CE.
[0038] The foregoing describes the implementation manner in which the computer device starts the memory failure repair after obtaining the failure analysis result by analyzing the failure information of the first memory row in the historical time period: when the failure mode is a memory row failure, or when the failure mode is a memory row failure and the failure level is a high-risk level, start the memory failure repair, and the failure repair is to replace the failed row with a redundant row and repair the data on the redundant row. In other embodiments, the computer device obtains the failure analysis result by analyzing the failure information of the second bank in the historical time period. Correspondingly, the implementation manner in which the computer device starts the memory failure repair is: when the failure mode is a memory bank failure, or when the failure mode is a memory bank failure and the failure level is a high-risk level, start the memory failure repair, and the failure repair is to replace the failed bank with a redundant bank and repair the data on the redundant bank.
[0039] That is, if the fault analysis result includes a fault mode, the computer device starts the fault repair of the memory according to the current fault analysis result of the memory, including: when the fault mode is a memory bank fault, starting the fault repair of the memory, where the fault repair includes: replacing the faulty bank with a redundant bank and repairing the data on the redundant bank.
[0040] Alternatively, if the fault analysis result includes a fault mode and a fault level, the computer device starts the fault repair of the memory according to the current fault analysis result of the memory, including: when the fault mode is a memory bank fault and the fault level is a high-risk level, starting the fault repair of the memory, where the fault repair includes: replacing the faulty bank with a redundant bank and repairing the data on the redundant bank.
[0041] Optionally, the redundant bank and the faulty bank are on the same channel in the memory.
[0042] It should be noted that the difference between this embodiment and the foregoing embodiment is that the second bank in this embodiment and the first memory in the foregoing embodiment are concepts of the same level. In the foregoing embodiment, the historical fault information is analyzed at the granularity of memory rows to obtain the fault analysis result, and in this embodiment, the historical fault information is analyzed at the granularity of banks to obtain the fault analysis result. In the foregoing embodiment, the faulty row is replaced with a redundant row, and the redundant row and the faulty row are on the same bank. In this embodiment, the faulty bank is replaced with a redundant bank, and the redundant bank and the faulty bank are on the same channel in the memory.
[0043] In a second aspect, a memory fault processing device is provided, and the memory fault processing device has the function of implementing the behavior of the memory fault processing method in the first aspect above. The memory fault processing device includes one or more modules, and the one or more modules are used to implement the memory fault processing method provided in the first aspect.
[0044] That is, a memory fault processing device is provided, and the device includes:
[0045] An analysis module, configured to start the fault analysis of the memory at a first moment; the fault analysis includes: obtaining the current fault analysis result of the memory by analyzing historical fault information, where the historical fault information is the fault information accumulated by the memory during a historical time period, and the historical time period is a time period before the first moment or a time period before and including the first moment;
[0046] A processing module, configured to start the fault repair of the memory according to the current fault analysis result of the memory.
[0047] Optionally, the first moment is the moment before an uncorrectable error (UCE) fault occurs in the computer system.
[0048] Optionally, the first moment includes:
[0049] The moment periodically started according to preset conditions; and / or, the moment when it is determined that a memory fault has occurred in the memory after the computer system runs.
[0050] Optionally, the analysis module includes:
[0051] An analysis sub-module for inputting historical fault information into a fault analysis model to obtain the current fault analysis result of the memory, where the fault analysis model is an intelligent computing analysis model.
[0052] Optionally, if the fault analysis result includes a fault mode, the processing module includes:
[0053] A first repair sub-module for starting the fault repair of the memory when the fault mode is a memory row fault, where the fault repair includes: replacing the faulty row with a redundant row and repairing the data on the redundant row.
[0054] Optionally, the analysis module is specifically configured to:
[0055] Obtain a first statistical feature according to the historical fault information, where the first statistical feature represents the number of faulty bits that have occurred in the first memory row within a historical time period, and the first memory row is any memory row;
[0056] When the first statistical feature is greater than a first threshold, determine that the fault mode is a memory row fault, where the first threshold represents the number of faulty bits that each memory row can tolerate.
[0057] Optionally, if the fault analysis result further includes a fault level, the processing module includes:
[0058] A second repair sub-module for starting the fault repair of the memory when the fault mode is a memory row fault and the fault level is a high-risk level.
[0059] Optionally, the analysis module is further specifically configured to:
[0060] Obtain a second statistical feature and / or a third statistical feature according to the historical fault information, where the second statistical feature represents the number of faults of each fault type that have occurred in the first memory row within a historical time period, and the third statistical feature represents the number of error corrections that have occurred in the first memory row within a historical time period;
[0061] When the second statistical feature is greater than the second threshold, or when the third statistical feature is greater than the third threshold, or when the second statistical feature is greater than the second threshold and the third statistical feature is greater than the third threshold, determine that the fault level is a high-risk level. The second threshold represents the number of faults of each fault type that each memory row can tolerate, and the third threshold represents the number of error corrections that each memory row can tolerate.
[0062] Optionally, the apparatus further includes:
[0063] An interaction module, configured to display risk mode options on an interaction interface, where the risk mode options include a memory high-risk mode option and a memory low-risk mode option.
[0064] Optionally, the first threshold, the second threshold, and the third threshold are variables set according to the risk mode.
[0065] Optionally, the first repair sub-module is specifically configured to:
[0066] Perform a read operation on the redundant row;
[0067] If the data read from the redundant row is incorrect data, correct the incorrect data and write the corrected data back to the redundant row to implement the repair of the data on the redundant row.
[0068] Optionally, the apparatus further includes:
[0069] A generation module, configured to generate a correctable error (CE) after the data read from the redundant row is incorrect data;
[0070] A suppression module, configured to suppress the CE.
[0071] Optionally, the apparatus further includes:
[0072] A release module, configured to release the suppression operation of the CE after the repair of the data on the redundant row is completed.
[0073] Optionally, if the fault analysis result includes a fault mode, the processing module includes:
[0074] A third repair sub-module, configured to start the fault repair of the memory when the fault mode is a memory bank fault, where the fault repair includes: replacing the faulty bank with a redundant bank and repairing the data on the redundant bank.
[0075] Optionally, the redundant bank and the faulty bank are located on the same channel in the memory.
[0076] In a third aspect, a computer device is provided. A computer program is stored in the computer device. When the computer program is run by the computer device, the method for handling memory faults provided in the first aspect is implemented.
[0077] Optionally, the computer device includes a processor and a memory. The memory is used to store a program for executing the method for handling memory faults provided in the first aspect, and to store data involved in implementing the method for handling memory faults provided in the first aspect. The processor is configured to execute the program stored in the memory to implement the method for handling memory faults provided in the first aspect. The operating device of the storage device may further include a communication bus, which is used to establish a connection between the processor and the memory.
[0078] In a fourth aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium. When the instructions are run on a computer, the computer is caused to execute the method for handling memory faults provided in the first aspect.
[0079] In a fifth aspect, a computer program product containing instructions is provided. When the computer program product is run on a computer, the computer is caused to execute the method for handling memory faults described in the first aspect.
[0080] The technical effects obtained in the second, third, fourth, and fifth aspects are similar to those obtained by the corresponding technical means in the first aspect, and will not be elaborated here.
[0081] The technical solutions provided in the embodiments of the present application can at least bring the following beneficial effects:
[0082] In the embodiments of the present application, a fault analysis result is obtained by analyzing historical fault information, and then the memory is repaired according to the fault analysis result. This solution can analyze memory faults more accurately. In addition, this solution can start the repair of memory faults without a cold reset, that is, it can repair memory faults in a timely manner, prevent system downtime, and reduce the impact on services. Description of the Drawings
[0083] Figure 1 is a flowchart of a method for handling memory faults provided in an embodiment of the present application;
[0084] Figure 2 is a schematic diagram of data repair for redundant rows provided in an embodiment of the present application;
[0085] Figure 3 is a flowchart of another method for handling memory faults provided in an embodiment of the present application;
[0086] Figure 4It is a flowchart of another method for handling memory faults provided by an embodiment of the present application;
[0087] Figure 5 It is a flowchart of another method for handling memory faults provided by an embodiment of the present application;
[0088] Figure 6 It is a schematic structural diagram of a device for handling memory faults provided by an embodiment of the present application;
[0089] Figure 7 It is a schematic structural diagram of another device for handling memory faults provided by an embodiment of the present application;
[0090] Figure 8 It is a schematic structural diagram of another device for handling memory faults provided by an embodiment of the present application;
[0091] Figure 9 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0092] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0093] Figure 1 It is a flowchart of a method for handling memory faults provided by an embodiment of the present application, and this method is applied to a computer device. Please refer to Figure 1 and the method includes the following steps.
[0094] Step 101: Start fault analysis of the memory at a first moment, and the fault analysis includes: obtaining the current fault analysis result of the memory by analyzing historical fault information.
[0095] In the embodiments of the present application, the basic storage unit of a memory (such as a dynamic random access memory (DRAM)) usually consists of a transistor and a capacitor, and the amount of charge carried on the capacitor determines whether the basic storage unit is '0' or '1'. Due to ionizing particles in the external environment or semiconductor hardware defects in the internal transistor, memory errors may occur, that is, memory faults occur.
[0096] After a memory failure occurs, the memory itself has an error correction algorithm (such as error checking and correcting (ECC)) to correct errors. The corrected errors are called corrected errors (CE). The error correction algorithm has a certain error correction ability, but the ability is limited. If the error exceeds the error correction ability of the error correction algorithm, uncorrected errors (UCE) will be generated, resulting in device downtime.
[0097] In order to repair memory failures in a timely manner, reduce the generation of UCE, reduce device downtime and restart, and mitigate the impact on services, the computer device in the embodiments of this application analyzes the fault information of memory failures that occurred in a historical time period to obtain a fault analysis result, and then determines whether to handle the memory failure and how to handle the memory failure according to the fault analysis result.
[0098] In the embodiments of this application, the computer device starts a fault analysis of the memory at a first moment. The fault analysis includes obtaining the current fault analysis result of the memory by analyzing historical fault information. The historical fault information is the fault information accumulated by the memory in a historical time period, and the historical time period is the time period before the first moment or the time period before and including the first moment.
[0099] Optionally, the first moment is the moment before a UCE fault occurs in the computer system. That is, a fault analysis of the memory is started during the operation of the computer system, and the operation period of the computer system refers to the period when the computer system is working properly.
[0100] Optionally, the first moment includes: the moment periodically started according to a preset condition; and / or, the moment when it is determined that a memory failure has occurred after the computer system starts running.
[0101] That is, when the computer device detects a memory failure, it starts to analyze the historical fault information to obtain a fault analysis result. Or, the computer device periodically analyzes the historical fault information to obtain a fault analysis result. Or, the computer device periodically analyzes the historical fault information to obtain a fault analysis result, and if a memory failure is detected within the periodic interval, it analyzes the historical fault information to obtain a fault analysis result and starts a new periodic analysis based on the time when the memory failure is detected this time. Or, the computer device periodically analyzes the historical fault information to obtain a fault analysis result, and if a memory failure is detected within the periodic interval, it analyzes the historical fault information to obtain a fault analysis result, but does not start a new periodic analysis based on the time when the memory failure is detected this time, that is, it does not affect the periodic analysis.
[0102] It should be noted that the computer device periodically analyzes historical fault information to predict the severity of memory faults in a timely manner and repair memory faults in a timely manner.
[0103] Optionally, in the embodiments of the present application, a fault analysis result is obtained by using a fault analysis model to intelligently analyze historical fault information. That is, the computer device inputs the historical fault information into the fault analysis model to obtain the current fault analysis result of the memory. The fault analysis model is an intelligent computing analysis model.
[0104] It should be noted that analyzing historical fault information through a fault analysis model is only one implementation manner for analyzing historical fault information provided by the embodiments of the present application. The computer device can also analyze historical fault information through other implementation manners, such as a data statistics-based manner. The embodiments of the present application do not limit the analysis methods adopted. Next, the implementation manners for the computer device to obtain a fault analysis result through a fault analysis model or through other means are introduced.
[0105] In the embodiments of the present application, if the fault analysis result includes a fault mode, when the fault mode is a memory row fault, the computer device starts to repair the memory fault. Among them, the fault repair includes: replacing the faulty row with a redundant row and repairing the data on the redundant row. That is, when the computer device determines through analyzing historical fault information that the current fault mode of the memory is a memory row fault, it performs memory row replacement and data repair.
[0106] In the embodiments of the present application, the historical fault information includes the fault location and fault time of memory faults that occurred during a historical time period. The computer device statistically analyzes the fault location and fault time included in the historical fault information to analyze the memory fault information and determine the fault mode.
[0107] Among them, the fault location refers to the physical address where the memory fault occurs. It should be noted that each memory fault occurs on a cell. When a memory fault is detected, which bank and which memory row the cell where the current memory fault occurs is located in, or which bank, which row, and which column it is located in, is the fault location of the current memory fault. The fault time refers to the time when the memory fault occurs.
[0108] It should be noted that the computer device stores a memory fault log, and the memory fault log records the fault information of memory faults that occurred during a historical time period, that is, it stores historical fault information.
[0109] In an embodiment of the present application, the computer device obtaining the current memory failure analysis result includes: obtaining a first statistical feature according to historical failure information, where the first statistical feature represents the number of failure bits that have occurred in a first memory row during a historical time period, the first memory row being any memory row. When the first statistical feature is greater than a first threshold, it is determined that the failure mode is a memory row failure, and the first threshold represents the number of failure bits that each memory row can tolerate.
[0110] Optionally, assuming that the computer device analyzes historical failure information through a failure analysis model, then the failure analysis model includes a first threshold. The computer device inputs the historical failure information into the failure analysis model, and the failure analysis model obtains the first statistical feature according to the historical failure information. That is, the computer device statistically counts the number of failure bits that have occurred in the first memory row during a historical time period to obtain the first statistical feature, and determines the failure mode through threshold judgment.
[0111] It should be noted that the memory includes multiple banks, each bank includes multiple memory rows, and each memory row includes multiple cells. A cell in the memory where a memory failure has occurred is a failure bit. During the historical time period, a memory failure may not have occurred on a cell, may have occurred once, or may have occurred more than once. The historical failure information includes the failure time and failure location of each memory failure that has occurred during the historical time period. The computer device statistically counts the number of memory failures with different failure locations among the memory failures in the first memory row during the historical time period to obtain the first statistical feature. If the first statistical feature is greater than the first threshold, indicating that multiple cells on the first memory row have had memory failures, then the computer device determines that the current failure mode of the memory is a memory row failure.
[0112] In addition, as can be seen from the foregoing, the computer device periodically starts memory failure analysis, or starts memory failure analysis when a memory failure occurs. Based on this, there are multiple situations for the computer device to determine the first memory row to be statistically counted, which will be introduced next.
[0113] In the case where memory failure analysis is started due to detecting a memory failure, the computer device determines the first memory row according to the failure location of the memory failure that has occurred this time. The first memory row refers to the memory row where the memory failure that has occurred this time is located. Alternatively, the computer device determines the first bank according to the failure location of the memory failure that has occurred this time, and determines a memory row included in the first bank as the first memory row. The first bank refers to the bank where the memory failure that has occurred this time is located, and the first memory row refers to one of the memory rows included in the first bank. Alternatively, the computer device determines a memory row included in the memory as the first memory row, that is, the first memory row refers to one of the memory rows included in the memory.
[0114] In the case where the computer device periodically initiates memory fault analysis, the computer device determines a first memory row based on the fault location of the most recent memory fault. The first memory row refers to the memory row where the most recent memory fault occurred. Alternatively, the computer device determines a first bank based on the fault location of the most recent memory fault, and determines a memory row included in the first bank as the first memory row. The first bank refers to the bank where the most recent memory fault occurred, and the first memory row refers to one of the memory rows included in the first bank. Alternatively, the computer device determines a memory row included in the memory as the first memory row, that is, the first memory row refers to one of the memory rows included in the memory.
[0115] It should be noted that in the case where the first memory row refers to one of the memory rows included in the first bank or the memory, for the other memory rows in the first bank or the memory except the first memory row, the computer device also statistically obtains the data corresponding to each memory row in the other memory rows in the same way as for the first memory row, and determines the first statistical feature based on the statistically obtained data.
[0116] In the case where the first memory row refers to the memory row where the current or most recent memory fault occurred, the computer device statistically analyzes the fault information about the first memory row in the historical fault information to obtain a quantity, and directly uses the statistically obtained quantity as the first statistical feature, that is, obtains a first statistical feature. In the case where the first memory row refers to one of the memory rows included in the first bank or the memory, the computer device statistically analyzes the fault information about multiple first memory rows in the historical fault information to obtain multiple quantities, each quantity corresponding to a memory row. The computer device uses the maximum value of the multiple statistically obtained quantities as the first statistical feature, or uses each of the multiple quantities as a first statistical feature to obtain multiple first statistical features, each first statistical feature corresponding to a memory row.
[0117] In the embodiments of the present application, after obtaining the first statistical feature, the computer device compares the first statistical feature with a first threshold to determine the current fault mode of the memory. For example, in the case of obtaining one first statistical feature, when the first statistical feature is greater than the first threshold, it is determined that the memory mode is a memory row fault. In the case of obtaining multiple first statistical features, when at least one of the multiple first statistical features is greater than the first threshold, it is determined that the memory mode is a memory row fault.
[0118] Optionally, the fault analysis result further includes a fault level. Then, when the fault mode is a memory row fault and the fault level is a high-risk level, the computer device initiates fault repair of the memory. Next, the implementation manner of the computer device determining the current fault level of the memory by analyzing the historical fault information is introduced.
[0119] In an embodiment of the present application, for the computer device to obtain the current fault analysis result of the memory, it further includes: obtaining a second statistical feature and / or a third statistical feature according to historical fault information, where the second statistical feature represents the number of faults of each fault type that occurred in the first memory row within a historical time period, and the third statistical feature represents the number of error corrections that occurred in the first memory row within the historical time period. When the second statistical feature is greater than a second threshold, or when the third statistical feature is greater than a third threshold, or when the second statistical feature is greater than the second threshold and the third statistical feature is greater than the third threshold, it is determined that the fault level is a high-risk level. Herein, the second threshold represents the number of faults of each fault type that each memory row can tolerate, and the third threshold represents the number of error corrections that each memory row can tolerate.
[0120] Optionally, assuming that the computer device analyzes the historical fault information through a fault analysis model, then the fault analysis model further includes the second threshold and / or the third threshold. The computer device inputs the historical fault information into the fault analysis model, and the fault analysis model obtains the second statistical feature and / or the third statistical feature according to the historical fault information. That is to say, the computer device statistically obtains the number of faults of each fault type that occurred in the first memory row within a historical time period to obtain the second statistical feature, and / or statistically obtains the number of error corrections that occurred in the first memory row within the historical time period to obtain the third statistical feature. Subsequently, the computer device compares the second statistical feature with the second threshold through the fault analysis model, and / or compares the third statistical feature with the third threshold to determine the fault level.
[0121] It should be noted that the historical fault information further includes the fault type and fault correction information of the memory faults that occurred within the historical time period. Among them, the fault type includes the CE type and the UCE type. Optionally, the CE type includes the patrol CE type, the read CE type, etc. The fault correction information includes information such as the amount of error correction data (also referred to as error correction data, with the unit such as bit) for error correction (such as ECC error correction) of each memory fault sent, and the error correction code.
[0122] In an embodiment of the present application, as can be seen from the foregoing, when the computer device periodically starts the memory fault analysis or starts the memory fault analysis when a memory fault occurs, based on this, there are many implementation manners for the computer device to statistically analyze the fault information of the first memory row in the historical fault information to obtain the second statistical feature and / or the third statistical feature. That is, there are various situations for the computer device to determine the first memory row to be statistically analyzed, which are the same as the various situations for determining the first memory row in the process of statistically obtaining the first statistical feature introduced above. Please refer to the foregoing introduction and details are not repeated here.
[0123] In the case where the first memory row refers to the memory row where the current or most recent memory failure occurred, the computer device statistically obtains the data corresponding to a memory row and directly uses the statistically obtained data as the second statistical feature and / or the third statistical feature. In the case where the first memory row refers to one of the memory rows included in the first bank or the memory, the computer device statistically obtains the data corresponding to multiple memory rows, and the computer device uses the statistically obtained data as the second statistical feature and / or the third statistical feature corresponding to the corresponding memory row.
[0124] In the embodiments of the present application, after obtaining the second statistical feature and / or the third statistical feature, the computer device compares the second statistical feature with a second threshold, and / or compares the third statistical feature with a third threshold to determine the current failure level of the memory.
[0125] It should be noted that since there are many types of failure types, there may be one or more failure types in the historical failure information. Therefore, the computer device needs to statistically count the number of failures of one or more failure types that occur in the first memory row to obtain one or more second statistical features corresponding to the memory row, and each second statistical feature corresponds to one failure type.
[0126] Optionally, one second threshold or multiple second thresholds are stored in the computer device. For example, the failure analysis model includes one second threshold or multiple second thresholds.
[0127] In the case where the computer device stores one second threshold, the computer device compares each of the one or more second statistical features corresponding to each obtained memory row with the second threshold. When all or part of the one or more second statistical features are greater than the second threshold, the failure level is determined to be a high-risk level.
[0128] In the case where the computer device stores multiple second thresholds, each of the multiple second thresholds corresponds to one failure type. For the one or more second statistical features corresponding to each obtained memory row, the computer device compares each second statistical feature with the second threshold corresponding to the same failure type. When all or part of the one or more second statistical features are greater than the corresponding second threshold, the failure level is determined to be a high-risk level.
[0129] Exemplarily, the fault types include inspection CE type, read CE type, and UCE type. The memory faults that occurred on the first memory row within the historical time period include 3 inspection CE types and 1 read CE type. Then, the computer device counts the first memory row and obtains two second statistical features, which are 3 and 1 respectively. 3 corresponds to the inspection CE type, and 1 corresponds to the read CE type. Suppose the computer device stores a second threshold, and the second threshold is 5. Then, the computer device compares both 3 and 1 with 5 and determines that the fault level is a low-risk level. Suppose the computer device stores 3 second thresholds, which are 8, 5, and 2 respectively. Among them, 8 corresponds to the inspection CE type, 5 corresponds to the read CE type, and 2 corresponds to the UCE type. Then, the computer device compares 3 with 8 and 1 with 5 and determines that the fault level is a low-risk level.
[0130] It should be noted that in the case where the first memory row refers to the memory row where the memory fault occurred this time or the most recent time, since only the data corresponding to one memory row is counted, in this way, when the second statistical feature corresponding to this memory row is greater than the second threshold, and / or the third statistical feature is greater than the third threshold, it is determined that the fault level is a high-risk level. If it is determined that the current fault mode of the memory is a memory row fault according to the fault information analysis of this memory row by the foregoing method, then it is determined that this memory row is a faulty row, and it is necessary to start the fault repair of the memory.
[0131] However, in the case where the first memory row refers to one of the memory rows included in the first bank or the memory, since the data corresponding to multiple memory rows are counted respectively, in this way, when the first statistical feature corresponding to the same memory row is greater than the first threshold, and the corresponding second statistical feature is greater than the second threshold and / or the third statistical feature is greater than the third threshold, it is determined that this memory row is a faulty row, and it is necessary to start the fault repair of the memory.
[0132] Optionally, risk mode options are displayed on the interaction interface. The risk mode options include a memory high-risk mode option and a memory low-risk mode option. That is to say, the computer device provides an interaction interface, and the user can select a risk mode through the interaction interface.
[0133] Optionally, the first threshold, the second threshold, and the third threshold are variables set according to the risk mode.
[0134] Optionally, the first threshold of the memory high-risk mode is less than the first threshold of the memory low-risk mode; and / or, the second threshold of the memory high-risk mode is less than the second threshold of the memory low-risk mode; and / or, the third threshold of the memory high-risk mode is less than the third threshold of the memory low-risk mode.
[0135] Optionally, the duration of the historical time period is a set fixed parameter. For example, the historical time period refers to the time period from the start of the computer device's installation and operation to the current analysis of the fault information, or the user configures the duration of the historical time period through the computer device. For example, the duration of the historical time period is configured as one month, and the historical time period refers to the one-month time before the current analysis of the fault information.
[0136] Optionally, the duration of the historical time period is a variable set according to the risk mode, and the duration of the historical time period in the high memory risk mode is less than the duration of the historical time period in the low memory risk mode.
[0137] Optionally, when the computer device analyzes that the fault mode is a memory row fault, or when it analyzes that the fault mode is a memory row fault and the fault level is a high risk level, it prompts the existence of a memory fault risk through the interaction interface.
[0138] Optionally, the user can also modify one or more of the first threshold, the second threshold, the third threshold, and the duration of the historical time period through the interaction interface.
[0139] As can be seen from the above, the user can flexibly select the risk mode according to the needs. For example, if the user's business risk is relatively high, the high risk mode can be selected. In this way, the first threshold and / or the second threshold and / or the third threshold are lower and / or the historical time period is shorter. The computer device analyzes the historical fault information within a shorter time period to obtain the first statistical feature, the second statistical feature, and / or the third statistical feature, and compares the obtained data with the smaller threshold to analyze whether it is a memory row fault and a high risk level. In this way, the computer device can ensure timely identification of less serious memory row faults. If the user's business risk is relatively low, the low risk mode can be selected, which can ensure high recognition, that is, timely identification of more serious memory row faults.
[0140] In the embodiment of the present application, the computer device provides an interaction interface for the user to select the risk mode. The computer device determines the duration of the fault information to be analyzed and / or the threshold size during threshold judgment according to the risk mode selected by the user, statistically analyzes the fault information within the corresponding duration, and performs threshold comparison. When the fault mode is identified as a memory row fault, or when the fault mode is identified as a memory row fault and the fault level is a high risk level, the memory fault is repaired in a timely manner. In this way, by integrating the risk mode selected by the user with the method of threshold comparison, while accurately predicting memory row faults, the computing pressure on the computer device is reduced.
[0141] Optionally, in some other embodiments, for the second statistical feature and the third statistical feature, the computer device statistically analyzes data in a more fine-grained manner. For example, the computer device statistically analyzes at least one of the maximum number of occurrences and the average number of occurrences of memory faults of each fault type on the first memory row within the first time interval to obtain the second statistical feature, and statistically analyzes at least one of the maximum error correction data volume and the average error correction data volume of memory faults of each fault type on the first memory row within the first time interval to obtain the third statistical feature. The historical time period includes multiple time intervals, and the first time interval is one of the multiple time intervals.
[0142] The computer device determines the fault level (risk level or risk grade) based on the maximum number of occurrences and / or the average number of occurrences, and the maximum error correction data volume and / or the average error correction data volume. For example, in the case where the computer device determines the maximum number of occurrences and the maximum error correction data volume, when the maximum number of occurrences is greater than or equal to the second threshold, and / or the maximum error correction data volume is greater than or equal to the third threshold, the fault level is determined to be a high risk level, where the fault level is divided into a low risk level and a high risk level. Alternatively, the computer device determines the fault level based on the threshold. Optionally, the fault level is divided into multiple levels, such as level one, level two, level three, etc. Level one indicates a relatively serious memory risk, and level three indicates a relatively less serious memory risk.
[0143] It should be noted that in this embodiment, the average number of occurrences includes one or more of the arithmetic mean, geometric mean, and harmonic mean. In addition, in addition to statistically analyzing the maximum number of occurrences and / or the average number of occurrences, the maximum error correction data volume and / or the average error correction data volume, other data can also be statistically analyzed, such as the median, standard deviation, etc. of various data. That is, there are many statistical methods, and the embodiments of the present application only take the statistical analysis of the maximum number of occurrences, the average number of occurrences, the maximum error correction data volume, and the average error correction data volume as examples for illustration.
[0144] Optionally, when the computer device can also determine the fault level, the computer device stores a first fault level. When the computer device identifies a memory row fault and the identified fault level is the same as or exceeds the first fault level, the computer device automatically repairs the memory row fault. Alternatively, the computer device first displays a relatively serious memory fault currently through the interaction interface to prompt the user to select whether to repair the memory fault, and the computer device determines whether to repair the memory row fault according to the user's selection operation.
[0145] Optionally, the first fault level stored in the computer device is the default configuration. Alternatively, the first fault level is the fault level selected by the user, that is, the user pre-selects the fault level through the interaction interface provided by the computer device according to the business risk requirements.
[0146] In this embodiment, the computer device statistically obtains fine-grained statistical features each time to identify the fault mode and fault level, and more accurately predicts the memory row fault and the risk level.
[0147] Optionally, in some other embodiments, the way for the computer device to analyze historical fault information and determine the fault mode and fault level can also be: the computer device determines the fault mode by means of threshold judgment based on statistical data, and determines the fault level by means of an intelligent analysis based on a fault analysis model. In this implementation, the computer device statistically analyzes the fault time, fault location, etc. in the historical fault information, and identifies the fault row mode by means of threshold comparison. In addition, the fault analysis model is used to intelligently analyze the fault location, fault time, fault type, and fault correction information in the historical fault information to identify the fault level. Optionally, in this implementation, the computer device provides an interactive interface for the user to select and configure the duration of the historical time period, the first threshold, the first fault level, etc., and the computer device accurately predicts the memory row fault and the fault level according to the configuration selected by the user.
[0148] Step 102: Start the fault repair of the memory according to the current fault analysis result of the memory.
[0149] In the embodiment of the present application, when the fault analysis result includes a fault mode and the current fault mode of the memory is a memory row fault, the computer device starts the fault repair of the memory. Optionally, when the fault analysis result further includes a fault level, and the fault mode is a memory row fault and the fault level is a high-risk level, the fault repair of the memory is started.
[0150] In the embodiment of the present application, the fault repair includes: replacing the faulty row with a redundant row in the memory and repairing the data on the redundant row.
[0151] Wherein, the faulty row refers to the memory row where the memory row fault occurs. For example, when the first memory row refers to the memory row where the current (or the most recent) memory fault occurs, the faulty row is the first memory row. When the first memory row refers to one of the memory rows included in the first bank (or memory), the computer device can determine the faulty row by means of threshold judgment or intelligent analysis, and the faulty row is a memory row on the first bank (or memory).
[0152] In the embodiment of the present application, the redundant row and the faulty row are located on the same bank in the memory, and the computer device replaces the faulty row with the redundant row on the bank where the faulty row is located.
[0153] Optionally, after determining that the fault repair of the memory needs to be started, the computer device also generates a row fault isolation request, and after generating the row fault isolation request, replaces the faulty row with a redundant row in the memory.
[0154] As described above, the user can select a risk mode according to business risk requirements. In this way, after the computer device determines that the failure mode is a memory row failure according to the risk mode selected by the user, a row failure isolation request is generated, indicating that the conditions for processing the memory row failure are currently met, and the computer device performs a memory row replacement. Optionally, the computer device can also prompt the user to select a memory row failure repair. After the computer device receives an instruction from the user to confirm the memory row failure repair, it performs a memory failure row replacement.
[0155] Optionally, the technology for online memory failure row replacement in the embodiments of the present application includes a soft post package repair (sPPR) technology.
[0156] In the embodiments of the present application, the implementation method for the computer device to repair the data on the redundant row is as follows: perform a read operation on the redundant row. If the data read from the redundant row is incorrect data, correct the incorrect data and write the corrected data back to the redundant row to achieve the repair of the data on the redundant row. That is to say, in the embodiments of the present application, the faulty data is repaired through the read operation and data write-back of the redundant row.
[0157] It should be noted that the computer device triggers a read operation on the redundant row to read all the data on the memory chip where the redundant row is located. When the redundant row is read, it determines whether the data on the redundant row is incorrect data according to the other data on the memory chip read, and corrects the incorrect data according to the other data read. In some other embodiments, the computer device triggers a read operation on the redundant row to read the data on the bank where the redundant row is located and some other banks in the memory, and performs data error correction on the redundant row according to the data read. That is to say, which banks or which memory chips the computer device actually reads the data from to perform data error correction on the redundant row is related to the storage algorithm (such as memory interleaving) when the actual memory stores data, which banks the chip select signal of the memory read operation is connected to, etc.
[0158] In the embodiments of the present application, the memory read operation is performed in a segmented read manner. The computer device is default-configured with a read interval for the memory read operation. For example, the read interval is 4 bits, that is, 4 bits of data are read each time, or the read interval is one or two cells, that is, one or two cells of data are read each time. The user can also change the default configuration.
[0159] For example, the read interval is 4 bits. For the data of the redundant row, assuming the data on the redundant row is 100 bits, the computer device reads 4-bit data each time in sequence and performs repair. After the repair, it reads the next 4-bit data for repair until all the data on the redundant row is repaired.
[0160] Optionally, the computer device divides the redundant row into M segments, each segment including one or more storage units, where M is an integer greater than 1. Let i = 1, perform a read operation on the i-th segment of the redundant row. If the data read from the i-th segment of the redundant row is incorrect data, correct the incorrect data and write the corrected data back to the i-th segment; if i is not equal to M, let i = i + 1, and return to perform a read operation on the i-th segment of the redundant row until i is equal to M.
[0161] Exemplarily, 4-bit data is read each time for error correction, and the read 4-bit data is corrected through an error correction algorithm, and the corrected data is written back to the position where the 4-bit data is located.
[0162] It should be noted that in the process of the computer device performing a read operation on the redundant row, the data on the redundant row is corrected through an error correction algorithm (such as ECC, single device data corrction (SDDC), etc.).
[0163] Figure 2 It is a schematic diagram of a method for repairing data on a redundant row through a read operation shown in an embodiment of the present application. Refer to Figure 2 , the method includes the following steps:
[0164] Step 201: The computer device performs row address parsing. That is, the computer device performs row address parsing on the faulty row, replaces the faulty row with the redundant row, that is, maps the address of the memory data of the faulty row to the redundant row. At this time, the data of the redundant row is empty.
[0165] Step 202: The computer device starts a memory area read operation. That is, the computer device reads the data on multiple banks through a memory read operation on the redundant row, including the first bank where the redundant row is located. When the data on the redundant row is read, the computer device determines that the data on the redundant row is incorrect data (shown by the black-filled squares) according to the data on the other banks read.
[0166] Step 203: The computer device performs data error correction. That is, the computer device corrects the incorrect data according to the data on the other banks read.
[0167] Step 204: The computer device performs data write-back. That is, the computer device writes the corrected data back to the redundant row, implementing data repair after the redundant row replaces the faulty row.
[0168] Exemplarily, Figure 2 A small square shown represents 4-bit data, and the computer device reads 4-bit data on the redundant row each time for correction. That is, when the computer device reads the redundant row, it sequentially reads a small square included in the redundant row. Suppose it reads Figure 2 the second small square shown, that is, the 4-bit data at the position of the black-filled square. After correcting the 4-bit data corresponding to the black-filled square according to the data on other banks read, the corrected data is obtained and written back to the position of the black-filled square on the redundant row. Then, it reads the 4-bit data in a small square after the black-filled square on the redundant row, that is, the 4-bit data in the third small square, and performs data error correction and writes the data back to the corresponding position. And so on, the computer device performs the actions of reading, correcting, and writing back in a segmented and successive manner to repair the data on the redundant row.
[0169] In the embodiment of the present application, after the data read from the redundant row is incorrect data, a CE is generated in the computer device, and the computer device suppresses the CE.
[0170] That is, since the computer device detects incorrect data when reading the redundant row, the computer device will consider that a CE is detected. Since this CE is not caused by a memory fault of the computer device, it is necessary to suppress the CE, that is, not process the CE, or in other words, the computer device does not record the CE.
[0171] Optionally, the computer device suppresses the CE during the process from the start of the read operation on the redundant row to the completion of the data repair on the redundant row.
[0172] Optionally, after the data repair on the redundant row is completed, the computer device releases the suppression operation of the CE. That is, the CE generated by the computer device after repairing the redundant row is caused by a real memory fault. Therefore, it is necessary to process the CE, that is, release the suppression operation of the CE and record the CE.
[0173] It should be noted that normally, each time the computer device generates a CE, a CE interrupt will occur, and the fault information of the generated CE will be recorded in the memory fault log. However, in the embodiment of the present application, by suppressing the CE during the read operation, the computer device will not record the fault information of the CE generated during this process in the memory fault log.
[0174] In the embodiment of the present application, the computer device implements the above functions through modules, see Figure 3 The computer device includes an execution module and a fault identification module. The computer device implements the above-mentioned memory fault processing method through the execution module and the fault identification module. The method includes the following steps.
[0175] Step 301: The execution module detects a memory fault and reports the fault information of the memory fault (including the fault location and fault time) to the fault identification module, that is, CE error reporting, to trigger the fault identification module to start fault analysis.
[0176] Step 302: The fault identification module analyzes the memory error, that is, analyzes the historical fault information, such as analyzing the physical address (fault location).
[0177] Step 303: The fault identification module performs memory fault identification prediction, that is, based on historical fault information, analyzes and determines the fault mode, or determines the fault mode and fault level, and when the determined fault mode meets the conditions for memory fault repair, or when the determined fault mode and fault level meet the conditions for memory fault repair, triggers the execution module to execute memory fault repair.
[0178] Step 304: The execution module executes sppr to replace the faulty memory row, that is, replace the faulty row with a redundant row.
[0179] Step 305: the execution module starts a memory row area read operation, data error correction, and data write-back to repair the data on the redundant row, that is, repairs the faulty data through a memory read operation on the redundant row.
[0180] Step 306: The execution module configures CE suppression of the memory row to suppress CE during the read operation of the redundant row.
[0181] Step 307: After the execution module completes the read operation on the redundant row, the CE inhibition is released, that is, the CE inhibition is released after the data is repaired.
[0182] Optionally, the execution module is a memory control module in a memory controller (such as a double data rate dynamic random access memory control (DDRC)) in a processor included in the computer device, and the fault identification module is a newly added module on the chip where the BMC is located. Alternatively, the fault identification module can also be added to any processing device included in the computer device.
[0183] Figure 4 FIG. 1 is a flowchart of another method for processing a memory failure provided by an embodiment of the present application.Figure 3 Based on Figure 4 , this method mainly includes error reporting, fault analysis (recognition), row replacement, and data write-back.
[0184] Among them, the process of error reporting includes: when the execution module detects a memory fault, hardware error correction (such as ECC) is performed, and the fault information of this memory fault (including the fault time and fault location) is reported to the fault recognition module, and the fault information is reported to the module for recording the memory fault log to record the fault information of this memory fault.
[0185] The process of fault analysis includes: the fault recognition module recognizes the fault mode of the memory fault (or recognizes the fault mode and fault level) according to the received fault information and the memory fault log. When it is recognized that the fault mode is a memory row fault (or it is recognized that the fault mode is a memory row fault and the fault level is a high-risk level), the execution module is triggered to perform memory fault row replacement.
[0186] The process of row replacement includes: the execution module triggers memory row replacement, that is, replacing the faulty row with a redundant row.
[0187] The process of data write-back includes: the execution module performs a memory area read operation on the redundant row, corrects the error data on the redundant row through an error correction algorithm, that is, performs data error correction, and writes the corrected data back to the redundant row. Optionally, if the data on the redundant row cannot be repaired through the error correction algorithm, a UCE may be generated, resulting in the computer reporting a crash and restart.
[0188] In summary, in the embodiment of the present application, by analyzing historical fault information to obtain a fault analysis result, and then performing fault repair on the memory according to the fault analysis result, this solution can analyze memory faults more accurately. In addition, this solution can start the fault repair of the memory without cold reset, that is, it can repair memory faults in a timely manner, prevent system crashes, and reduce business impacts.
[0189] After obtaining the fault analysis result by analyzing the fault information of the first memory row in the historical time period as described above, the implementation method for the computer device to start the fault repair of the memory is as follows: when the fault mode is a memory row fault, or when the fault mode is a memory row fault and the fault level is a high-risk level, start the fault repair of the memory. The fault repair is to replace the faulty row with a redundant row and repair the data on the redundant row. In some other embodiments, the computer device analyzes the fault information of the second bank in the historical time period to obtain the fault analysis result. Correspondingly, the implementation method for the computer device to start the fault repair of the memory is as follows: when the fault mode is a memory bank fault, or when the fault mode is a memory bank fault and the fault level is a high-risk level, start the fault repair of the memory. The fault repair is to replace the faulty bank with a redundant bank and repair the data on the redundant bank.
[0190] Among them, in the case where a memory fault occurs this time and the memory fault analysis is started, the second bank refers to the bank where the memory row where the memory fault occurs this time is located, or the second bank refers to a bank on the memory die where the memory row where the memory fault occurs this time is located, or the second bank refers to any bank in the memory. In the case where the memory fault analysis is started periodically, the second bank refers to the bank where the memory row where the most recent memory fault occurred is located, or the second bank refers to a bank on the memory die where the memory row where the most recent memory fault occurred is located, or the second bank refers to any bank in the memory.
[0191] Next, refer to Figure 5 to introduce this embodiment. Figure 5 is a flowchart of a method for processing memory faults provided by an embodiment of the present application. This method is applied to a computer device. Please refer to Figure 5 The method includes the following steps.
[0192] Step 501: Start the fault analysis of the memory at the first moment. The fault analysis includes: obtaining the current fault analysis result of the memory by analyzing the historical fault information.
[0193] In the embodiments of the present application, when a computer device detects a memory failure, it analyzes historical failure information to obtain a failure analysis result. Alternatively, the computer device periodically analyzes historical failure information to obtain a failure analysis result. Alternatively, the computer device periodically analyzes failure information to obtain a failure analysis result, and if a memory failure is detected within a cycle interval, it analyzes historical failure information to obtain a failure analysis result and restarts the cycle analysis based on the time when the memory failure is detected this time. Alternatively, the computer device periodically analyzes historical failure information to determine a failure mode, and if a memory failure is detected within a cycle interval, it analyzes historical failure information to obtain a failure analysis result, but does not restart the cycle analysis based on the time when the memory failure is detected this time, that is, it does not affect the cycle analysis.
[0194] It should be noted that the historical failure information is the failure information of memory failures that occurred within a historical time period, and the duration of the historical time period may be the same as or different from the historical time period in the foregoing embodiments. Since it is necessary to analyze whether there is a relatively serious memory bank failure, therefore, when the duration of the historical time period is longer than the historical time period in the foregoing embodiments, the analysis of the memory bank failure is more accurate to a certain extent.
[0195] Optionally, in the embodiments of the present application, the computer device analyzes historical failure information through a failure analysis model to obtain the current failure analysis result of the memory, that is, the computer device inputs the historical failure information into the failure analysis model to obtain the current failure analysis result of the memory, and the failure analysis model is an intelligent computing analysis model.
[0196] In the embodiments of the present application, the failure analysis result includes a failure mode.
[0197] Optionally, the historical failure information includes the failure location and failure time of memory failures that occurred within a historical time period. The computer device counts the failure location and failure time of historical memory failures to obtain the number of failure bits that appear in the second bank, that is, obtains the fourth statistical feature. When the number of failure bits that appear in the second bank is greater than or equal to the fourth threshold within the historical time period, that is, when the fourth statistical feature is greater than the fourth threshold, it is determined that the failure mode is a memory bank failure. Among them, the fourth threshold represents the number of failure bits that each bank can tolerate.
[0198] Optionally, assuming that the computer device analyzes historical failure information through a failure analysis model, then the failure analysis model includes the fourth threshold.
[0199] Optionally, the fault analysis result further includes a fault level, and the historical fault information further includes the fault type and / or fault correction information of the memory faults that occurred within the historical time period. The computer device obtains a fifth statistical feature and / or a sixth statistical feature based on the historical fault information. The fifth statistical feature represents the number of faults of each fault type that occurred in the second bank within the historical time period, and the sixth statistical feature represents the number of error corrections that occurred in the second bank within the historical time period. When the fifth statistical feature is greater than the fifth threshold, or when the sixth statistical feature is greater than the sixth threshold, or when the fifth statistical feature is greater than the fifth threshold and the sixth statistical feature is greater than the sixth threshold, the fault level is determined to be a high-risk level. Wherein, the fifth threshold represents the number of faults of each fault type that each bank can tolerate, and the sixth threshold represents the number of error corrections that each bank can tolerate.
[0200] Optionally, assuming that the computer device analyzes the historical fault information through a fault analysis model, then the fault analysis model further includes a fifth threshold and / or a sixth threshold.
[0201] Optionally, the duration of the historical time period and / or the fourth threshold and / or the fifth threshold and / or the sixth threshold are variables set according to the risk mode.
[0202] Optionally, the risk mode includes a memory high-risk mode and a memory low-risk mode. The duration of the historical time period in the memory high-risk mode is shorter than the duration of the second time period in the memory low-risk mode; and / or, the fourth threshold in the memory high-risk mode is less than the second threshold in the memory low-risk mode; and / or, the fifth threshold in the memory high-risk mode is less than the sixth threshold in the memory low-risk mode; and / or, the sixth threshold in the memory high-risk mode is less than the sixth threshold in the memory low-risk mode.
[0203] Optionally, the computer device also provides an interaction interface, and risk mode options are displayed on the interaction interface. The risk mode options include a high-risk mode option and a low-risk mode option. The user can select the risk mode through the interaction interface according to the business risk requirements.
[0204] Optionally, the interaction interface is further used to prompt the existence of a memory fault risk when it is confirmed that the fault mode is a memory bank fault.
[0205] It should be noted that in this embodiment, different from the above Figure 1 embodiment, the second bank in this embodiment and Figure 1 the first memory in the embodiment are concepts of the same level. Figure 1 The embodiment analyzes the fault mode of the memory fault at the granularity of the memory row. Figure 5 This embodiment analyzes the fault mode of the memory fault at the granularity of the bank. For Figure 5 the implementation manner of the computer device to determine the fault mode, refer to the foregoingFigure 1 The relevant content in the embodiments will not be elaborated here.
[0206] Step 502: When the failure mode is a memory bank failure, start the failure repair of the memory. Among them, the failure repair includes: replacing the faulty bank with a redundant bank and repairing the data on the redundant bank.
[0207] In the embodiments of the present application, if the computer device determines that the failure mode is a memory bank failure, it replaces the faulty bank with a redundant bank in the memory and repairs the faulty data. The faulty bank refers to the bank where the memory failure occurs.
[0208] Optionally, the redundant bank and the faulty bank are located on the same channel in the memory.
[0209] Figure 5 What is different from the Figure 1 embodiment in the shown embodiment is that Figure 1 in the embodiment, the faulty row is replaced with a redundant row, and the redundant row and the faulty row are on the same bank. Figure 5 in the embodiment, the faulty bank is replaced with a redundant bank, and the redundant bank and the faulty bank are located on the same channel in the memory.
[0210] It should be noted that the memory includes multiple channels, each channel includes multiple dual inline memory modules (DIMMs), one DIMM includes multiple ranks, one rank includes multiple chips (memory dies), and one chip includes multiple banks.
[0211] To sum up, in the embodiments of the present application, by analyzing historical failure information, the current failure mode of the memory is determined. When the failure mode is a memory bank failure, the faulty bank is replaced with a redundant bank and data repair is performed. This solution can more accurately identify the failure mode, and the memory bank replacement can be performed without a cold reset, enabling the memory failure to be repaired in a timely manner, preventing system downtime, and reducing the impact on services.
[0212] Figure 6 It is a schematic structural diagram of a memory failure processing device 600 provided by the embodiments of the present application. The memory failure processing device 600 can be implemented as part or all of a computer device by software, hardware, or a combination of both. The computer device can be the Figure 9 computer device shown below. Refer to Figure 6 . The device 600 includes: an analysis module 601 and a processing module 602.
[0213] An analysis module 601 is configured to start a fault analysis of the memory at a first moment; the fault analysis includes: obtaining a current fault analysis result of the memory by analyzing historical fault information, where the historical fault information is the fault information accumulated by the memory within a historical time period, and the historical time period is a time period before the first moment or a time period before and including the first moment; for the specific implementation method, refer to the detailed introduction of step 201 in the foregoing Figure 1 embodiment, which will not be elaborated here.
[0214] A processing module 602 is configured to start a fault repair of the memory according to the current fault analysis result of the memory. For the specific implementation method, refer to the detailed introduction of step 102 in the foregoing Figure 1 embodiment, which will not be elaborated here.
[0215] Optionally, the first moment is a moment before a computer system has an uncorrectable error (UCE) fault.
[0216] Optionally, the first moment includes:
[0217] A moment periodically started according to a preset condition; and / or, a moment when it is determined that a memory fault occurs in the memory after the computer system runs.
[0218] Optionally, the analysis module 601 includes:
[0219] An analysis sub-module is configured to input the historical fault information into a fault analysis model to obtain a current fault analysis result of the memory, and the fault analysis model is an intelligent computing analysis model.
[0220] Optionally, if the fault analysis result includes a fault mode, the processing module 602 includes:
[0221] A first repair sub-module is configured to start a fault repair of the memory when the fault mode is a memory row fault, where the fault repair includes: replacing the faulty row with a redundant row and repairing the data on the redundant row. For the specific implementation method, refer to the detailed introduction of step 102 in the foregoing Figure 1 embodiment, which will not be elaborated here.
[0222] Optionally, the analysis module 601 is specifically configured to:
[0223] Obtain a first statistical feature according to the historical fault information, where the first statistical feature represents the number of faulty bits that appear in a first memory row within the historical time period, the first memory row is any memory row, and the first threshold represents the number of faulty bits that each memory row can tolerate; for the specific implementation method, refer to the detailed introduction of step 101 in the foregoing Figure 1 embodiment, which will not be elaborated here.
[0224] When the first statistical feature is greater than the first threshold, determine that the failure mode is a memory row failure.
[0225] Optionally, the failure analysis result further includes a failure level, and the processing module 602 includes:
[0226] A second repair sub-module, configured to start a failure repair of the memory when the failure mode is a memory row failure and the failure level is a high-risk level.
[0227] Optionally, the analysis module 601 is further specifically configured to:
[0228] Obtain a second statistical feature and / or a third statistical feature according to historical failure information, where the second statistical feature represents the number of failures of each failure type that occur in the first memory row within a historical time period, and the third statistical feature represents the number of error corrections that occur in the first memory row within a historical time period; the specific implementation manner refers to the detailed introduction of step 101 in the foregoing Figure 1 Embodiment, which will not be elaborated here.
[0229] When the second statistical feature is greater than the second threshold, or when the third statistical feature is greater than the third threshold, or when the second statistical feature is greater than the second threshold and the third statistical feature is greater than the third threshold, determine that the failure level is a high-risk level, where the second threshold represents the number of failures of each failure type that each memory row can tolerate, and the third threshold represents the number of error corrections that each memory row can tolerate.
[0230] Optionally, referring to Figure 7 , the device 600 further includes:
[0231] An interaction module 603, configured to display risk mode options on an interaction interface, where the risk mode options include a memory high-risk mode option and a memory low-risk mode option.
[0232] Optionally, the first threshold, the second threshold, and the third threshold are variables set according to the risk mode.
[0233] Optionally, the first repair sub-module is specifically configured to:
[0234] Perform a read operation on the redundant row;
[0235] If the data read from the redundant row is incorrect data, correct the incorrect data and write the corrected data back to the redundant row to implement the repair of the data on the redundant row. The specific implementation manner refers to the detailed introduction of step 102 in the foregoing Figure 1 Embodiment, which will not be elaborated here.
[0236] Optionally, referring to Figure 8 , the device 600 further includes:
[0237] A generation module 604, configured to generate a correctable error (CE) after determining that the data read from a redundant row is incorrect data.
[0238] A suppression module 605, configured to suppress the CE. For the specific implementation method, refer to the detailed description of step 102 in the foregoing Figure 1 embodiment, which will not be elaborated here.
[0239] Optionally, referring to Figure 8 , the apparatus 600 further includes:
[0240] A release module 606, configured to release the suppression operation of the CE after the data repair on the redundant row is completed. For the specific implementation method, refer to the detailed description of step 102 in the foregoing Figure 1 embodiment, which will not be elaborated here.
[0241] Optionally, the apparatus 600 further includes:
[0242] A generation module, configured to generate a row fault isolation request after determining that the fault mode is a memory row fault.
[0243] Optionally, the redundant row and the faulty row are located on the same bank in the memory.
[0244] Optionally, if the fault analysis result includes a fault mode, the processing module 602 includes:
[0245] A third repair sub-module, configured to start the fault repair of the memory when the fault mode is a memory bank fault, where the fault repair includes: replacing the faulty bank with a redundant bank and repairing the data on the redundant bank.
[0246] Optionally, the redundant bank and the faulty bank are located on the same channel in the memory.
[0247] In the embodiment of the present application, the fault analysis result is obtained by analyzing historical fault information, and then the memory is fault-repaired according to the fault analysis result. This solution can analyze memory faults more accurately, and can start the fault repair of the memory without cold reset, preventing system downtime and reducing business impact.
[0248] It should be noted that: when the memory fault processing apparatus provided in the above embodiment processes a memory fault, only the above-mentioned functional module division is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the apparatus is divided into different functional modules to complete all or part of the functions described above. In addition, the memory fault processing apparatus provided in the above embodiment and Figures 1 to 5The embodiments of the method for handling memory faults shown belong to the same inventive concept. For the specific implementation process, refer to the method embodiments, which will not be elaborated here.
[0249] An embodiment of the present application provides a computer device. A computer program is stored in the computer device. When the computer program runs on the computer device, it implements the method for handling memory faults in the above Figures 1 to 4 embodiment, or implements the method for handling memory faults in the Figure 5 embodiment. For the specific implementation manner, refer to the detailed introduction in the foregoing Figures 1 to 5 method embodiment shown, which will not be elaborated here.
[0250] Optionally, the computer device includes a processor and a chip where the BMC is located. The processor includes a memory controller, and the memory controller includes an execution module. The BMC in the chip where the BMC is located includes a fault identification module. The memory controller runs the execution module to implement the corresponding functions of the execution module in the above Figure 3 embodiment, and the BMC runs the fault identification module to implement the corresponding functions of the fault identification module in the above Figure 3 embodiment.
[0251] Optionally, in addition to being set in the BMC, the fault identification module can also be added to other processing devices included in the computer device to implement the corresponding functions.
[0252] In the embodiment of the present application, the computer device obtains a fault analysis result by analyzing historical fault information, and then repairs the memory fault according to the fault analysis result. This solution can analyze memory faults more accurately, and can start repairing memory faults without cold reset, that is, it can repair memory faults in time, prevent system downtime, and reduce business impact.
[0253] It should be noted that: when the computer device provided in the above embodiment processes memory faults, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the computer device provided in the above embodiment and the Figure 1 or Figure 5 embodiments of the method for handling memory faults shown belong to the same inventive concept. For the specific implementation process, refer to the method embodiments, which will not be elaborated here.
[0254] Please refer to Figure 9 , Figure 9 which is a schematic structural diagram of a computer device shown according to an embodiment of the present application. The computer device includes one or more processors 901, a communication bus 902, a memory 903, and one or more communication interfaces 904.
[0255] The processor 901 is a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or one or more integrated circuits for implementing the solution of this application. For example, an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. Optionally, the above PLD is a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0256] The communication bus 902 is used to transfer information between the above components. Optionally, the communication bus 902 is divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus.
[0257] Optionally, the memory 903 is a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compressed optical disc, a laser disc, a digital versatile disc, a Blu-ray disc, etc.), a magnetic disk storage medium, or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 903 exists independently and is connected to the processor 901 through the communication bus 902, or the memory 903 is integrated with the processor 901.
[0258] The communication interface 904 uses any device such as a transceiver for communicating with other devices or communication networks. The communication interface 104 includes a wired communication interface and optionally also includes a wireless communication interface. Among them, the wired communication interface is, for example, an Ethernet interface, etc. Optionally, the Ethernet interface is an optical interface, an electrical interface, or a combination thereof. The wireless communication interface is a wireless local area networks (WLAN) interface, a cellular network communication interface, or a combination thereof, etc.
[0259] Optionally, in some embodiments, the computer device includes multiple processors, and each of these processors is a single-core processor or a multi-core processor. Optionally, the processor here refers to one or more devices, circuits, and / or processing cores for processing data (such as computer program instructions).
[0260] In a specific implementation, as an embodiment, the computer device further includes an output device 906 and an input device 907. The output device 906 communicates with the processor 901 and can display information in various ways. For example, the output device 906 is a liquid crystal display (LCD), a light emitting diode (LED) display device, a cathode ray tube (CRT) display device, or a projector, etc. The input device 907 communicates with the processor 901 and can receive user input in various ways. For example, the input device 907 is a mouse, a keyboard, a touch screen device, or a sensing device, etc.
[0261] In some embodiments, the memory 903 is used to store the program code 910 for executing the solution of this application, and the processor 901 can execute the program code 910 stored in the memory 903. The program code includes one or more software modules, and the computer device can, through the processor 901 and the program code 910 in the memory 903, Figure 1 or Figure 5 implement the memory fault handling method provided by the above
[0262] In other embodiments, the program code for executing the solution of this application is stored in the processor 901, and the processor 901 is used to execute the program code to implement the above Figure 1 or Figure 5 memory fault handling method provided by the embodiment. The program code includes one or more software modules. For example, the processor 901 includes a memory controller, and the program code is stored in the memory controller. The memory controller includes Figure 3 the execution module and the fault identification module shown, and the above Figure 1 orFigure 5 The memory fault handling method provided by the embodiment.
[0263] In some other embodiments, part of the program code for executing the solution of this application is stored in the processor 901. For example, the processor 901 includes a memory controller, and the memory controller includes Figure 3 The execution module shown. The computer device also includes other processing devices other than the processor 901, and another part of the program code for executing the solution of this application is stored in the other processing devices. The processor 901 and the other processing devices jointly implement the above Figure 1 or Figure 5 The memory fault handling method provided by the embodiment. For example, the other processing device is the chip where the baseboard management controller (BMC) is located, and the BMC includes Figure 3 The fault identification module shown. By running the fault identification module through the BMC, it jointly implements the above Figure 1 or Figure 5 The memory fault handling method provided by the embodiment.
[0264] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital versatile disc (DVD)), or a semiconductor medium (such as a solid state disk (SSD)), etc. It should be noted that the computer-readable storage medium mentioned in the embodiments of the present application can be a non-volatile storage medium, in other words, it can be a non-transitory storage medium.
[0265] It should be understood that the "at least one" mentioned herein refers to one or more, and "a plurality" refers to two or more. In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B can mean A or B; the "and / or" herein is only a description of the association relationship of the associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish the same items or similar items with basically the same functions and roles. Those skilled in the art can understand that the terms "first" and "second" do not limit the quantity and execution order, and the terms "first" and "second" do not necessarily mean different.
[0266] The above are the embodiments provided by the present application, which are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for handling memory faults, characterized in that, The method includes: Starting a fault analysis of the memory at a first moment; the fault analysis includes: obtaining the current fault analysis result of the memory by analyzing historical fault information, where the historical fault information is the fault information accumulated by the memory within a historical time period, and the historical time period is a time period before the first moment or a time period before and including the first moment; Starting a fault repair of the memory according to the current fault analysis result of the memory; Wherein, the fault analysis result includes a fault mode, and the fault mode includes a memory row fault. Obtaining the current fault analysis result of the memory includes: Obtaining a first statistical feature according to the historical fault information, where the first statistical feature represents the number of fault bits that have occurred in a first memory row within the historical time period, and the first memory row is any memory row; When the first statistical feature is greater than a first threshold, determining that the fault mode is a memory row fault, and the first threshold represents the number of fault bits that each memory row can tolerate.
2. The method according to claim 1, characterized in that, The first moment is the moment before an uncorrectable error (UCE) fault occurs in the computer system.
3. The method according to claim 1, wherein The first moment includes: The moment periodically started according to preset conditions; and / or, the moment when it is determined that the memory has a memory fault after the computer system runs.
4. The method according to claim 1, characterized in that, The obtaining the current fault analysis result of the memory by analyzing historical fault information includes: Inputting the historical fault information into a fault analysis model to obtain the current fault analysis result of the memory, and the fault analysis model is an intelligent computing analysis model.
5. The method according to any one of claims 1-4, characterized in that The starting a fault repair of the memory according to the current fault analysis result of the memory includes: When the fault mode is the memory row fault, starting a fault repair of the memory, where the fault repair includes: replacing the faulty row with a redundant row and repairing the data on the redundant row.
6. The method according to any one of claims 1-4, characterized in that, If the fault analysis result further includes a fault level, then the starting a fault repair of the memory according to the current fault analysis result of the memory includes: When the fault mode is the memory row fault and the fault level is a high-risk level, starting a fault repair of the memory.
7. The method according to claim 6, wherein The obtaining the current fault analysis result of the memory further includes: Obtaining a second statistical feature and / or a third statistical feature according to the historical fault information, where the second statistical feature represents the number of faults of each fault type that have occurred in the first memory row within the historical time period, and the third statistical feature represents the number of error corrections that have occurred in the first memory row within the historical time period; When the second statistical feature is greater than a second threshold, or when the third statistical feature is greater than a third threshold, or when the second statistical feature is greater than the second threshold and the third statistical feature is greater than the third threshold, determining that the fault level is a high-risk level, the second threshold represents the number of faults of each fault type that each memory row can tolerate, and the third threshold represents the number of error corrections that each memory row can tolerate.
8. The method according to claim 7, wherein The method further includes: Display risk mode options on the interaction interface, where the risk mode options include a high memory risk mode option and a low memory risk mode option.
9. The method according to claim 8, wherein The first threshold, the second threshold, and the third threshold are variables set according to the risk mode.
10. The method according to claim 5, characterized in that, The repairing of the data on the redundant row includes: Performing a read operation on the redundant row; If the data read from the redundant row is incorrect data, correcting the incorrect data and writing the corrected data back to the redundant row to repair the data on the redundant row.
11. The method according to claim 10, characterized in that, After the data read from the redundant row is incorrect data, the method further includes: Generating a correctable error CE; Suppressing the CE.
12. The method according to claim 11, wherein After the repair of the data on the redundant row is completed, the method further includes: Removing the suppression operation of the CE.
13. The method according to any one of claims 1-4, characterized in that, If the failure mode further includes a memory bank failure, then the starting of the failure repair of the memory according to the current failure analysis result of the memory includes: When the failure mode is the memory bank failure, starting the failure repair of the memory, where the failure repair includes: replacing the failed bank with a redundant bank and repairing the data on the redundant bank.
14. A processing device for memory faults, characterized in that, The device includes: An analysis module, configured to start a failure analysis of the memory at a first moment; the failure analysis includes: obtaining the current failure analysis result of the memory by analyzing historical failure information, where the historical failure information is the failure information accumulated by the memory within a historical time period, and the historical time period is a time period before the first moment or a time period before and including the first moment; A processing module, configured to start a failure repair of the memory according to the current failure analysis result of the memory; Wherein the failure analysis result includes a failure mode, the failure mode includes a memory row failure, and the analysis module is specifically configured to: obtain a first statistical feature according to the historical failure information, the first statistical feature representing the number of failure bits that occur in a first memory row within the historical time period, and the first memory row is any memory row; when the first statistical feature is greater than a first threshold, determining that the failure mode is a memory row failure, and the first threshold represents the number of failure bits that each memory row can tolerate.
15. The device according to claim 14, characterized in that, The first moment is the moment before an uncorrectable error UCE failure occurs in the computer system.
16. The device according to claim 14, characterized in that, The first moment includes: A moment periodically started according to preset conditions; and / or, a moment when it is determined that the memory has a memory failure after the computer system runs.
17. The device according to claim 14, wherein The analysis module includes: An analysis sub-module, configured to input historical failure information into a failure analysis model to obtain the current failure analysis result of the memory, and the failure analysis model is an intelligent computing analysis model.
18. The device according to any one of claims 14-17, characterized in that, The processing module includes: A first repair sub-module, configured to start a failure repair of the memory when the failure mode is the memory row failure, where the failure repair includes: replacing the failed row with a redundant row and repairing the data on the redundant row.
19. The device according to any one of claims 14-17, characterized in that, The fault analysis result also includes a fault level, and the processing module includes: The second repair submodule is used to start fault repair of the memory when the fault mode is the memory row fault and the fault level is a high risk level.
20. The device according to claim 19, characterized in that, The analysis module is also specifically used for: Obtaining a second statistical feature and / or a third statistical feature according to the historical fault information, wherein the second statistical feature indicates the number of faults of each fault type occurring in the first memory row in the historical time period, and the third statistical feature indicates the number of error corrections occurring in the first memory row in the historical time period; When the second statistical feature is greater than a second threshold, or when the third statistical feature is greater than a third threshold, or when the second statistical feature is greater than the second threshold and the third statistical feature is greater than the third threshold, the fault level is determined to be a high risk level, the second threshold represents the number of faults of each fault type that each memory row can tolerate, and the third threshold represents the number of error corrections that each memory row can tolerate.
21. The device according to claim 20, characterized in that, The device also includes: The interactive module is used to display risk mode options on the interactive interface, wherein the risk mode options include a high-risk memory mode option and a low-risk memory mode option.
22. The device according to claim 21, characterized in that, The first threshold, the second threshold, and the third threshold are variables set according to the risk model.
23. The device according to claim 18, characterized in that, The first repair submodule is specifically used for: performing a read operation on the redundant row; If the data read from the redundant row is erroneous data, the erroneous data is corrected and the corrected data is written back to the redundant row to repair the data on the redundant row.
24. The device according to claim 23, characterized in that, The device also includes: A generating module, configured to generate a correctable error CE after the data read from the redundant row is erroneous data; A suppression module is used to suppress the CE.
25. The device according to claim 24, characterized in that, The device also includes: The release module is used to release the suppression operation of the CE after the repair of the data on the redundant row is completed.
26. The device according to any one of claims 14-17, characterized in that, If the failure mode also includes a memory bank failure, the processing module includes: The third repair submodule is used to start the fault repair of the memory when the fault mode is the memory bank fault, wherein the fault repair includes: replacing the faulty bank with a redundant bank and repairing the data on the redundant bank.
27. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program to implement the method according to any one of claims 1 to 13.
28. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Memory detection model training method and memory detection method and device
CN110598802A
(build-in self repair method and device for embedded SRAM)
KR1020060094592A
Hybrid memory system with configurable error thresholds and failure analysis capability
US20190340080A1