Memory fault prediction method, device and equipment

By aggregating the ECC verification features that can be corrected by memory errors, aggregating error features are generated, the problem of low memory failure prediction accuracy in the prior art is solved, the prediction accuracy is improved, and the risk of equipment downtime is reduced.

CN114996065BActive Publication Date: 2025-08-08ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210604963.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-30
Publication Date
2025-08-08
Estimated Expiration
2042-05-30

AI Technical Summary

Technical Problem

In the prior art, the accuracy of memory failure prediction is poor, resulting in a high risk of equipment downtime.

Method used

By obtaining the ECC verification error characteristics that can correct errors in the device to be predicted multiple times in the current time window, including the error position and error form characteristics, perform feature aggregation, and generate aggregated error characteristics to predict whether memory uncorrectable errors will occur.

Benefits of technology

Improves the accuracy of memory failure prediction and reduces the risk of device downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114996065B_ABST
    Figure CN114996065B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a memory fault prediction method, apparatus, and device. The method comprises: obtaining multiple ECC check error signatures indicating that a device to be predicted has multiple correctable memory errors within a current time window, wherein the ECC check error signatures include error location signatures and error form signatures; based on the error location signatures in the ECC check error signatures, performing feature aggregation on the multiple ECC check error signatures to obtain aggregated error signatures; and predicting, based on the aggregated error signatures, whether the device to be predicted will have uncorrectable memory errors. The present application can improve the accuracy of predicting whether uncorrectable memory errors will occur.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a memory fault prediction method, apparatus, and device. Background Art

[0002] Memory failure is the most common failure in hardware systems, greatly affecting the reliability, availability, and serviceability (RAS) of the system.

[0003] Typically, after reading data from memory, the memory controller performs error checking. If a correctable error (CE) occurs, the error is corrected. If an uncorrectable error (UCE) occurs, the location of the error is re-accessed. If an uncorrectable error occurs after multiple accesses, the hardware system issues an UCE signal, causing the device to crash. To reduce downtime, the current approach is to predict whether a device will experience uncorrectable errors in the future based on the number of correctable errors that have occurred over a period of time.

[0004] However, this prediction method has the problem of poor accuracy. Summary of the Invention

[0005] Embodiments of the present application provide a memory fault prediction method, apparatus, and device to address the problem of poor accuracy in predicting whether an uncorrectable error will occur in the prior art.

[0006] In a first aspect, an embodiment of the present application provides a memory fault prediction method, comprising:

[0007] Acquire multiple ECC check error features of memory correctable errors that occur multiple times in the device to be predicted within the current time window, wherein the ECC check error features include error position features and error form features;

[0008] Based on the error position feature in the ECC error feature, performing feature aggregation on the multiple ECC error features to obtain an aggregated error feature;

[0009] Predict whether an uncorrectable memory error will occur in the device to be predicted based on the aggregated error feature.

[0010] In a second aspect, an embodiment of the present application provides a memory fault prediction device, comprising:

[0011] An acquisition module is configured to acquire multiple ECC error features of memory correctable errors that occur multiple times in the device to be predicted within a current time window; the ECC error features include error location features and error form features;

[0012] an aggregation module, configured to aggregate the plurality of ECC error features based on an error position feature in the ECC error feature to obtain an aggregated error feature;

[0013] A prediction module is used to predict whether the device to be predicted will have an uncorrectable memory error based on the aggregated error characteristics.

[0014] In a third aspect, an embodiment of the present application provides an electronic device comprising: a memory and a processor; wherein the memory is used to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement a method as described in any one of the first aspects.

[0015] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed, the method as described in any one of the first aspects is implemented.

[0016] An embodiment of the present application further provides a computer program, which, when executed by a computer, is used to implement the method as described in any one of the first aspects.

[0017] In an embodiment of the present application, multiple ECC check error features of the device to be predicted that have multiple correctable memory errors in the current time window can be obtained. The ECC check error features include error position features and error form features. Based on the error position features in the ECC check error features, multiple ECC check error features are feature aggregated to obtain aggregated error features. Based on the aggregated error features, it is predicted whether the device to be predicted will have uncorrectable memory errors. This enables prediction of whether uncorrectable memory errors will occur based on the specific ECC check error features of multiple correctable memory errors. In particular, it predicts whether uncorrectable memory errors will occur in the device to be predicted based on the aggregated error features obtained by feature aggregation of multiple ECC check error features. Therefore, the prediction of memory failures can be based on the microscopic ECC check error features, and the historical ECC check error situations of the device to be predicted can be considered from a macroscopic perspective, thereby improving the accuracy of the prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0019] Figure 1 A schematic diagram of an application scenario of the memory fault prediction method according to an embodiment of the present application;

[0020] Figure 2 A flowchart of a memory fault prediction method provided in one embodiment of the present application;

[0021] Figure 3 A schematic diagram of a single memory chip single read data error situation provided by an embodiment of the present application;

[0022] Figure 4 A schematic diagram of a training model and prediction using the model provided in one embodiment of the present application;

[0023] Figure 5 A schematic diagram of the structure of a memory fault prediction device provided in one embodiment of the present application;

[0024] Figure 6 A schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0025] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] The terms used in the examples of this application are for the purpose of describing specific embodiments only and are not intended to limit this application. The singular forms "a," "the," and "the" used in the examples of this application and the appended claims are also intended to include plural forms, and unless the context clearly indicates otherwise, "a plurality" generally includes at least two, but does not exclude the inclusion of at least one.

[0027] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0028] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0029] It should also be noted that the terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such product or system. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the product or system comprising the element.

[0030] In addition, the step sequence in the following method embodiments is only an example and not a strict limitation.

[0031] Figure 1 A schematic diagram of an application scenario of the memory fault prediction method provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, this application scenario may include a first device 11 and at least one second device 12. The first device 11 can predict memory failures in the second device 12, specifically, whether the second device will experience an uncorrectable memory error. The second device 12 can be any type of electronic device for which prediction of an uncorrectable memory error is required. The second device 12 can be referred to as the device to be predicted. For example, the second device 12 can be a physical machine, which refers to a dedicated physical server with a virtualized computer system deployed.

[0032] The second device 12 includes a memory controller, which can be integrated into the CPU (Central Processing Unit) of the second device 12. In response to the CPU's write instruction, the memory controller can generate a corresponding error correction code (ECC) based on the data to be written to the memory, and write the data and the error correction code into the memory. In response to the CPU's read instruction, the memory controller can read the data and error correction code from the memory, and perform ECC verification on the data based on the error correction code. If a correctable error occurs, the error will be corrected according to the error correction code. If an uncorrectable error occurs, the error location will be re-accessed. If an uncorrectable memory error occurs during multiple accesses, the hardware system will issue a UCE signal and cause the second device 12 to shut down.

[0033] Typically, to reduce downtime, memory failure prediction for the second device 12 is performed using the following method: based on the number of correctable memory errors that occurred in the second device 12 over a period of time, it is predicted whether the second device 12 will experience uncorrectable memory errors in the future. However, this method of predicting based on the number of correctable memory errors over a period of time does not consider the specific ECC error characteristics of correctable memory errors, resulting in a technical problem of poor prediction accuracy.

[0034] For example, if many errors at a certain error location are single-bit errors, since the number of correctable memory errors is large, the prediction result obtained based on the number of correctable memory errors is that uncorrectable errors will occur. However, since single-bit errors will be corrected, no matter how many single-bit errors occur, they will not lead to uncorrectable memory errors. Therefore, such a prediction result is inaccurate, resulting in poor prediction accuracy.

[0035] In order to solve the technical problem of poor prediction accuracy in the prior art, in an embodiment of the present application, multiple ECC check error features of the device to be predicted that have multiple correctable memory errors in the current time window can be obtained. The ECC check error features include error position features and error form features. Based on the error position features in the ECC check error features, multiple ECC check error features are feature aggregated to obtain aggregated error features. Based on the aggregated error features, it is predicted whether the device to be predicted will have uncorrectable memory errors. This achieves the prediction of whether uncorrectable memory errors will occur based on the specific ECC check error features of multiple correctable memory errors. In addition, it specifically predicts whether uncorrectable memory errors will occur in the device to be predicted based on the aggregated error features obtained by feature aggregation of multiple ECC check error features. Therefore, the prediction of memory failures can be based on the microscopic ECC check error features, and the historical ECC check error situations of the device to be predicted can be considered from a macro perspective, thereby improving the accuracy of the prediction.

[0036] For example, in the case where many errors at a certain error location are single-bit errors, since the memory failure is predicted based on the error location characteristics and error form characteristics of the correctable memory error, it is possible to obtain a prediction result that no uncorrectable error will occur based on the knowledge that single-bit errors will be corrected and no matter how many single-bit errors occur, no uncorrectable memory error will occur. This can reduce the occurrence of inaccurate predictions and improve the accuracy of the predictions.

[0037] It should be noted that for the specific description of correctable memory errors and uncorrectable memory errors, please refer to the specific description in the relevant technology, which will not be repeated here.

[0038] It should be noted that Figure 1 In the example above, a device other than the device to be predicted performs memory fault prediction on the device to be predicted. It is understandable that in other embodiments, the device to be predicted may also perform memory fault prediction itself.

[0039] The following describes some embodiments of the present application in detail with reference to the accompanying drawings. Unless there is a conflict, the following embodiments and features in the embodiments may be combined with each other.

[0040] Figure 2 This is a flow chart of a memory fault prediction method provided by an embodiment of the present application. The execution subject of this embodiment can be the device to be predicted itself, or can be other devices other than the device to be predicted. Figure 2 As shown, the method of this embodiment may include:

[0041] Step 21: obtaining multiple ECC error features of memory correctable errors that occur multiple times in the predicted device within the current time window, where the ECC error features include error location features and error form features;

[0042] Step 22: Based on the error position feature in the ECC error feature, perform feature aggregation on multiple ECC error features to obtain an aggregated error feature;

[0043] Step 23: predict whether the device to be predicted will have an uncorrectable memory error based on the aggregated error characteristics.

[0044] In the embodiment of the present application, the size of the current time window can be flexibly implemented, and the size of the current time window can be, for example, 3 days. If the device to be predicted experiences multiple correctable memory errors within the current time window, each occurrence of a correctable memory error can have a corresponding ECC error signature. The ECC error signature can include an error location signature and an error form signature.

[0045] It should be noted that if a memory access reads data from a single memory chip with an error and the error is corrected using the ECC algorithm, then this memory access is considered to have a single correctable error. If a memory access reads data from multiple memory chips with errors and the error is corrected using the ECC algorithm, then this memory access is considered to have multiple correctable errors. A single memory access refers to the prefetch of a cache line block (64 bytes).

[0046] Among them, the error location feature refers to a feature used to describe the location of the memory where a correctable memory error occurs. The error location feature can be used to describe the location of the memory where the correctable error occurs. For example, the location described by the error location feature can be accurate to the CELL of the memory. CELL is the basic unit of the memory and can be used to store 1 bit of data.

[0047] Based on this, in one embodiment, when the CPU of the device to be predicted is a multi-core CPU and the number of memory controllers integrated in a single single-core CPU can be multiple, the error location characteristics may include processor (Socket), memory controller (Integrated Memory Controller, IMC), memory channel (Channel), slot (Slot), Rank, BankGroup, Bank, Row and Column, which can be used to describe which single-core CPU, which memory controller, which channel, which slot, which Rank, which BankGroup, which Bank, which row and which column has a memory correctable error. It should be noted that for the specific description of Channel, Rank, Bank Group and Bank, please refer to the specific description in the relevant technology and will not be repeated here.

[0048] Since the memory controller can access multiple memory chips of a Rank and is limited by the error correction capability of the memory controller, if two different memory chips (also called memory particles) have errors at the same time in the memory data corresponding to the same cache line block (64 bytes), an uncorrectable error will definitely occur. Therefore, knowing the historical errors of the memory chip is helpful to further improve the accuracy of fault prediction. In one embodiment, the memory location feature can include the memory chip to describe which specific memory chip has the memory correctable error.

[0049] Error form characteristics refer to characteristics used to describe the error form of memory correctable errors. In actual applications, after the memory controller issues a read command, each of the multiple memory chips on a rank can return data through multiple bursts. The number of bits transmitted in a burst is the memory bit width, usually 4 or 8 bits. Due to the influence of noise, multiple bits returned by the same memory chip through the same burst or different bursts may be erroneous at the same time. Therefore, the error form of memory correctable errors can be described from the perspective of the burst. Based on this, in one embodiment, the error form characteristics may include error form characteristics described from the perspective of the burst, wherein the error form characteristics described from the perspective of the burst may include: error form characteristics of error bits within the same burst, and / or error form characteristics of error bits between different bursts. Exemplarily, the error form characteristics described from the perspective of the burst may include one or more of the following: the number of error bits in the same burst, the position of the error bits in the same burst, whether the error bits in the same burst are continuous, the number of bursts in which error bits appear, the position of the burst in which error bits appear, or whether the bursts in which error bits appear are continuous.

[0050] In practical applications, multiple bits share the same set of data I / O channels (Data Queues, DQs). When a DQ has a problem, multiple bits will typically be erroneous. Therefore, the error form of a correctable memory error can be described from the DQ perspective. Based on this, in one embodiment, the error form characteristics may include error form characteristics described from the DQ perspective, wherein the error form characteristics described from the DQ perspective may include: error form characteristics of erroneous bits within the same DQ, and / or error form characteristics of errors between different DQs. Exemplarily, the error form characteristics described from the DQ perspective may include one or more of the following.

[0051] Assuming that the memory chip is an X4 chip (that is, it provides 4 DQs and one burst can return 4 bits stored in the memory chip), then one memory access can prefetch the data required for a cacheline (64 bytes). One memory chip can contribute 8 bursts of data, a total of 32 bits, and the data provided by 16 memory chips can together form a cacheline. If a memory correctable error occurs when reading a memory chip, and the error condition is as follows Figure 3As shown, a circle can represent 1 bit of data, a circle filled with white can represent a non-error bit, and a circle filled with black can represent an error bit. Therefore, a memory correctable error occurs in the reading of the memory chip this time, and the ECC check error characteristics of the memory correctable error this time may include, for example, an error position feature for describing that the memory chip where the correctable error occurs is the memory chip, and an error form feature for describing that 3 DQs, specifically DQ0, DQ1, and DQ2, have errors, and 4 Bursts, specifically the 1st Burst, 3rd Burst, 5th Burst, and 8th Burst have errors.

[0052] In actual applications, when the memory controller verifies that a correctable error has occurred in the memory, it can record the relevant error information in its own register, and the error location characteristics and error form characteristics can be generated based on the data recorded in the register. Taking the processor of the device to be predicted as an Intel processor as an example, the ECC check error characteristics can be generated based on the data recorded in the register of the memory controller for recording the retry read error log (retry read error log). Of course, in other embodiments, the relevant error information can also be recorded in other registers of the memory controller, and this application does not limit this. It should be noted that the specific content of the register of the memory controller for recording the read error log (retry read error log) can be found in the specific description in the relevant technology, which will not be repeated here.

[0053] Exemplarily, the reading of the register can be event-triggered, and the operating system of the device to be predicted can capture the event indicating the occurrence of a correctable memory error. In response to the event, the register can be read through the driver. Taking the operating system as the Linux operating system and the processor as the Intel processor as an example, the data in the register can be read through the EDAC driver (Linux error detection and correction driver).

[0054] In one embodiment, multiple ECC error signatures may be received when multiple correctable memory errors occur on the device to be predicted within a current time window. In another embodiment, data in a register when multiple correctable memory errors occur on the device to be predicted within the current time window may be obtained, and an ECC error signature for each correctable memory error may be generated based on the data in the register when each correctable memory error occurs.

[0055] Optionally, considering that the characteristics of the memory error of the device to be predicted and the memory error correction capability are related to its own static characteristics, that is, the static characteristics of the device to be preset can affect the characteristics of the memory error of the device, and can also affect the memory error correction capability of the device, so in order to further improve the accuracy of the prediction, the memory fault prediction can also be based on the static characteristics of the device to be predicted. Based on this, in one embodiment, the method provided in this embodiment can also include obtaining the target static characteristics of the device to be predicted. Among them, the target static characteristics can specifically be one or more characteristics of the device to be predicted that can affect the characteristics of the memory error or the memory error correction capability. Exemplarily, the target static characteristics can include one or more of the following: CPU model, memory batch, number of memory sticks, memory stick insertion method, operating system (OS), basic input output system (BIOS) model.

[0056] Among them, different CPU models may have different ECC algorithms. The ECC algorithm can determine the memory error correction capability, so the target static feature can include the CPU model. The memory batch can determine the characteristics of the memory error, so the target static feature can include the memory batch. The number of memory sticks and the memory stick insertion method determine the interleaving method of memory access, which can determine the characteristics of the memory error, so the target static feature can include the number of memory sticks and / or the memory stick insertion method. The operating system and BIOS model can determine the behavior of the hardware system and the possible load conditions, thereby determining the characteristics of the memory error, so the target static feature can include the operating system and / or BIOS model.

[0057] Taking into account that the time taken for devices with different static characteristics to go from the occurrence of the first correctable error to the occurrence of an uncorrectable error may vary greatly, for example, the inventor observed column errors in some batches of Micron and row errors in Samsung C generation in a production environment. The time taken from the occurrence of the first correctable error to the occurrence of an uncorrectable error was very short. However, in some batches of Hynix, the failure gradually worsened, and the time taken from the occurrence of the first correctable error to the occurrence of an uncorrectable error was longer. Therefore, in order to further improve the accuracy of the prediction, in one embodiment, the number of current time windows can be multiple, and the sizes of the multiple current time windows are different to adapt to various situations where the time taken from the occurrence of the first correctable error to the occurrence of an uncorrectable error varies greatly. The sizes of the multiple current time windows can be determined based on experience, and the multiple ECC check error features can include multiple ECC check error features in which the device to be predicted has multiple memory correctable errors in each current time window.

[0058] In an embodiment of the present application, after obtaining multiple ECC error signatures indicating multiple occurrences of correctable memory errors in the device to be predicted within the current time window, the multiple ECC error signatures can be aggregated to obtain an aggregated error signature. The feature aggregation can be performed based on the error location features within the ECC error signatures, and the aggregation method can be a statistical method such as summation, averaging, or magnitude. The aggregation granularity can be fixed or variable.

[0059] In one embodiment, step 22 may specifically include: determining a target granularity for feature aggregation, and aggregating multiple ECC error features into an aggregated error feature of the target granularity based on the error location feature in the ECC error feature. The target granularity may be a granularity of a larger range of memory locations where correctable memory errors occur than that described by the error location feature, so as to achieve aggregation of features to a larger fault range. Exemplarily, the target granularity may include any one of the following: BANK row granularity, BANK column granularity, BANK granularity, RANK granularity, memory bar granularity, channel granularity, memory controller granularity, CPU granularity, or device granularity. The target granularity may be determined based on experience. For example, the inventor observed in a production environment that the column error features of some batches of Micron were very obvious and prone to downtime. The ECC error features of a single correctable error may be aggregated to the BANK column granularity for prediction.

[0060] It should be understood that in the case where the current time window includes multiple current time windows with different window sizes, the aggregated error feature may include the aggregated error feature corresponding to each current time window in the multiple current time windows, wherein the aggregated error feature corresponding to each current time window is obtained by performing feature aggregation on multiple ECC check error features of memory correctable errors that occur multiple times in the device to be predicted in each current time window.

[0061] In the embodiment of the present application, after obtaining the aggregate error feature, it is possible to predict whether the device to be predicted will have an uncorrectable memory error based on the obtained aggregate error feature.

[0062] In one embodiment, the aggregated error features can be used to predict whether the device being predicted will experience uncorrectable memory errors. In this case, when there is only one current time window, the aggregated error features can be used as the feature based on which the prediction is based. When there are multiple current time windows, the concatenation of the aggregated error features corresponding to the multiple current time windows can be used as the feature based on which the prediction is based. In another embodiment, the aggregated error features and the target static features can be used to predict whether the device being predicted will experience uncorrectable memory errors.

[0063] Since there is a large difference between the ECC check error characteristics that will cause a crash and the ECC check error characteristics that will not cause a crash, it is possible to predict whether the device to be predicted will have an uncorrectable memory error based on which characteristic the aggregated error characteristics are more similar to. Specifically, when the characteristics on which the prediction is based are more similar to the characteristics that will cause a crash, it can be predicted that the device to be predicted will have an uncorrectable memory error. When the characteristics on which the prediction is based are more similar to the characteristics that will not cause a crash, it can be predicted that the device to be predicted will not have an uncorrectable memory error.

[0064] Exemplarily, a machine learning approach can be used to perform the prediction, i.e., the features on which the prediction is based can be input into a prediction model to obtain a prediction result of whether the device to be predicted will experience an uncorrectable error. Based on this, in one embodiment, step 23 can specifically include: inputting the aggregated error features into the prediction model to obtain a prediction result of whether the device to be predicted will experience an uncorrectable memory error. In another embodiment, step 23 can specifically include: inputting the concatenation result of the aggregated error features corresponding to multiple current time windows into the prediction model to obtain a prediction result of whether the device to be predicted will experience an uncorrectable memory error.

[0065] Taking the aggregated error feature as the feature based on which prediction is based as an example, the prediction model can be trained in the following manner: construct a prediction model with training parameters set in the prediction model; input multiple sample aggregated error features into the prediction model respectively to generate prediction results; based on the difference between the prediction results and the expected results corresponding to the sample labels of the sample aggregated error features, iteratively adjust the training parameters until the difference meets the preset requirements.

[0066] The sample labels of the sample aggregated error features can be positive samples or negative samples. Positive samples can refer to samples with uncorrectable errors, and the corresponding expected result can be 1. Negative samples can refer to samples without uncorrectable errors, and the corresponding expected result can be 0. Therefore, after the aggregated error features are input into the prediction model, if the output result is 1, it can be said that the probability of an uncorrectable error is 1, that is, it is predicted that an uncorrectable error will occur. If the output result is 0, it can be said that the probability of an uncorrectable error is 0, that is, it is predicted that an uncorrectable error will not occur.

[0067] For example, if the number of time windows is multiple, the operating system is Linux, and the processor is Intel, Figure 4 As shown in Figure 2, the prediction model can be divided into two stages: offline learning and online prediction + learning.

[0068] During the offline learning phase, register data related to correctable memory errors can be collected through the EDAC driver, and the register data can be written in batches to the offline data warehouse. For the register data written to the offline data warehouse, the feature generation module can generate ECC error features corresponding to multiple time windows. For the generated ECC error features, the feature aggregation module can obtain aggregated error features corresponding to multiple time windows. For the obtained aggregated error features, the time aggregation module can obtain spliced aggregated error features. Based on the spliced aggregated error features and the corresponding static features, a prediction model for online memory fault prediction can be obtained through an offline training process.

[0069] During the online prediction and learning phase, register data related to correctable memory errors collected online is processed by the feature generation module, feature aggregation module, and time aggregation module. It is then input into the prediction model along with the corresponding static features to generate prediction results. Furthermore, the online memory fault prediction results can be fed back into the offline data warehouse for model optimization.

[0070] In an embodiment of the present application, if it is predicted that the device to be predicted will not have an uncorrectable memory error, it can be said that the probability that the ECC can fully cover the memory check error after the device to be predicted has a high probability; if it is predicted that the device to be predicted will have a correctable memory error, it can be said that the probability that the ECC can fully cover the memory check error after the device to be predicted has a low probability. By predicting whether the device to be predicted will have an uncorrectable memory error, it is possible to find a device with a low probability of being fully covered by the ECC after the memory check error occurs, so as to perform further processing on the device. For example, if the device to be predicted is a physical machine, and it is predicted that the device to be predicted will have an uncorrectable memory error, the virtual machine on the device to be predicted can be migrated to other devices in advance. Of course, in other embodiments, other types of further processing can also be performed on the device to be predicted, and this application does not limit this.

[0071] The memory fault prediction method provided in this embodiment obtains multiple ECC check error features of the device to be predicted that has multiple correctable memory errors in the current time window, the ECC check error features including error location features and error form features, and based on the error location features in the ECC check error features, performs feature aggregation on multiple ECC check error features to obtain aggregated error features, and predicts whether the device to be predicted will have uncorrectable memory errors based on the aggregated error features. This achieves the prediction of whether uncorrectable memory errors will occur based on the specific ECC check error features of multiple correctable memory errors, and specifically predicts whether the device to be predicted will have uncorrectable memory errors based on the aggregated error features obtained by aggregating multiple ECC check error features. Therefore, the prediction of memory faults can be based on the microscopic ECC check error features, and the historical ECC check error situations of the device to be predicted can be considered from a macroscopic perspective, thereby improving the accuracy of the prediction.

[0072] Figure 5 This is a schematic diagram of the structure of a memory fault prediction device provided by an embodiment of the present application; Figure 5 As shown, this embodiment provides a memory fault prediction device, which can execute the memory fault prediction method described in the above embodiment. Specifically, the device may include:

[0073] An acquisition module 51 is configured to acquire multiple ECC error features of a device to be predicted that has multiple memory correctable errors within a current time window; the ECC error features include error location features and error form features;

[0074] An aggregation module 52 is configured to aggregate the plurality of ECC error features based on error position features in the ECC error features to obtain an aggregated error feature;

[0075] The prediction module 53 is configured to predict whether an uncorrectable memory error will occur in the device to be predicted based on the aggregated error characteristics.

[0076] In one embodiment, the error location characteristics include processor, memory controller, memory channel, slot, rank, bank group, bank, row and column.

[0077] In one embodiment, the error location feature also includes a memory chip.

[0078] In one embodiment, the error form characteristics include: error form characteristics described from a Burst perspective, and / or error form characteristics described from a DQ perspective.

[0079] In one embodiment, the error form characteristics described from the perspective of Burst include one or more of the following: the number of erroneous bits in the same Burst, the position of the erroneous bits in the same Burst, whether the erroneous bits in the same Burst are continuous, the number of bursts with erroneous bits, the position of the bursts with erroneous bits, or whether the bursts with erroneous bits are continuous.

[0080] In one embodiment, the error form characteristics described from the DQ perspective include one or more of the following: the number of erroneous bits in the same DQ, the position of the erroneous bits in the same DQ, whether the erroneous bits in the same DQ are continuous, the number of DQs where erroneous bits occur, the position of the DQs where erroneous bits occur, or whether the DQs where erroneous bits occur are continuous.

[0081] In one embodiment, the aggregation module 52 is specifically configured to determine a target granularity for feature aggregation; and aggregate the multiple ECC error features into an aggregated error feature of the target granularity based on error position features in the ECC error features.

[0082] In one embodiment, the target granularity includes any one of the following: BANK row granularity, BANK column granularity, BANK granularity, RANK granularity, memory bar granularity, channel granularity, memory controller granularity, CPU granularity or device granularity.

[0083] In one embodiment, the acquisition module 51 is further configured to acquire target static features of the device to be predicted;

[0084] The prediction module 53 is specifically configured to predict whether an uncorrectable memory error will occur in the device to be predicted based on the aggregated error feature and the target static feature.

[0085] In one embodiment, the target static features include one or more of the following: CPU model, memory batch, number of memory sticks, memory stick insertion method, operating system or BIOS model.

[0086] In one embodiment, the current time window includes multiple current time windows with different window sizes; and the aggregated error feature includes aggregated error features corresponding to the multiple current time windows respectively.

[0087] In one embodiment, the prediction module 53 is specifically configured to input the aggregated error features into a prediction model to obtain a prediction result of whether the device to be predicted will have an uncorrectable memory error.

[0088] In one embodiment, the prediction model is trained in the following manner: constructing a prediction model, in which training parameters are set; inputting multiple sample aggregation error features into the prediction model respectively to generate prediction results; based on the difference between the prediction results and the expected results corresponding to the sample labels of the sample aggregation error features, iteratively adjusting the training parameters until the difference meets the preset requirements.

[0089] Figure 5 The device shown can perform Figure 2 For the method of the embodiment shown in FIG. 1 , reference may be made to the description of the part not described in detail in the embodiment. Figure 2 The implementation process and technical effects of this technical solution can be found in Figure 2 The description in the illustrated embodiment will not be repeated here.

[0090] In one possible implementation, Figure 5 The structure of the device shown can be realized as an electronic device. Figure 6 As shown, the electronic device may include: a processor 61 and a memory 62. The memory 62 is used to store the data that supports the electronic device to execute the above Figure 2 The processor 61 is configured to execute the program stored in the memory 62 according to the method provided in the illustrated embodiment.

[0091] The program includes one or more computer instructions, wherein when the one or more computer instructions are executed by the processor 61, the following steps can be implemented:

[0092] Acquire multiple ECC check error features of memory correctable errors that occur multiple times in the predicted device within the current time window; the ECC check error features include error position features and error form features;

[0093] Aggregating the multiple ECC error features based on error position features in the ECC error features to obtain an aggregated error feature;

[0094] Predict whether an uncorrectable memory error will occur in the device to be predicted based on the aggregated error feature.

[0095] Optionally, the processor 61 is further configured to execute the aforementioned Figure 2 All or part of the steps on the electronic device side in the illustrated embodiment.

[0096] The structure of the electronic device may further include a communication interface 63 for the electronic device to communicate with other devices or a communication network.

[0097] In addition, an embodiment of the present application provides a computer storage medium having a computer program stored thereon, which, when executed, implements the following Figure 2 The method of any one of the embodiments shown.

[0098] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present embodiment without inventive effort.

[0099] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by adding a necessary general hardware platform, and of course can also be implemented by combining hardware and software. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a computer product. This application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0100] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable device to produce a machine, so that the instructions executed by the processor of the computer or other programmable device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0101] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0102] These computer program instructions can also be loaded onto a computer or other programmable device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0103] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0104] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0105] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, linked lists, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A memory fault prediction method, characterized in that: include: Acquire multiple ECC error features of multiple memory correctable errors occurring in the device to be predicted within the current time window. The ECC error features include error location features and error form features. A memory correctable error occurring during a single read of a single memory chip is considered a single memory correctable error. The error form features included in the ECC error features of a single memory correctable error are features used to describe the situation of the memory correctable error from a burst perspective and / or a DQ perspective. Based on the error position feature in the ECC error feature, performing feature aggregation on the multiple ECC error features to obtain an aggregated error feature; Predict whether an uncorrectable memory error will occur in the device to be predicted based on the aggregated error feature.

2. The method according to claim 1, characterized in that The error location features include: processor, memory controller, memory channel, slot, rank, bank group, bank, row and column.

3. The method according to claim 2, characterized in that The error location feature also includes: memory chip.

4. The method according to claim 1, wherein The features used to describe the correctable memory error from a burst perspective include one or more of the following: the number of erroneous bits in the same burst, the position of the erroneous bits in the same burst, whether the erroneous bits in the same burst are continuous, the number of bursts with erroneous bits, the position of the bursts with erroneous bits, or whether the bursts with erroneous bits are continuous.

5. The method according to claim 1, characterized in that The features used to describe the correctable memory error from the DQ perspective include one or more of the following: the number of erroneous bits in the same DQ, the position of the erroneous bits in the same DQ, whether the erroneous bits in the same DQ are continuous, the number of DQs where erroneous bits occur, the position of the DQs where erroneous bits occur, or whether the DQs where erroneous bits occur are continuous.

6. The method according to any one of claims 1 to 5, characterized in that The step of aggregating the plurality of ECC error features based on the error position feature in the ECC error feature to obtain an aggregated error feature includes: Determine the target granularity for feature aggregation; Based on the error position features in the ECC error features, the multiple ECC error features are aggregated into an aggregated error feature of a target granularity.

7. The method according to claim 6, characterized in that The target granularity includes any one of the following: BANK row granularity, BANK column granularity, BANK granularity, RANK granularity, memory bar granularity, channel granularity, memory controller granularity, CPU granularity or device granularity.

8. The method according to any one of claims 1 to 5, characterized in that The method comprises: obtaining target static features of the device to be predicted; The predicting, based on the aggregated error characteristics, whether the device to be predicted will have an uncorrectable memory error includes: Predict whether an uncorrectable memory error will occur in the device to be predicted based on the aggregated error feature and the target static feature.

9. The method according to claim 8, characterized in that The current time window includes multiple current time windows with different window sizes; the aggregated error feature includes aggregated error features corresponding to the multiple current time windows respectively.

10. The method according to any one of claims 1 to 5, characterized in that The predicting, based on the aggregated error characteristics, whether the device to be predicted will have an uncorrectable memory error includes: The aggregated error features are input into a prediction model to obtain a prediction result of whether the device to be predicted will have an uncorrectable memory error.

11. A memory fault prediction device, characterized in that: include: An acquisition module is used to obtain multiple ECC check error characteristics of memory correctable errors that occur multiple times in the device to be predicted within the current time window; The ECC error feature includes an error location feature and an error form feature. A memory correctable error occurring during a single read of a single memory chip is considered a memory correctable error. The error form feature included in the ECC error feature of a memory correctable error is a feature used to describe the situation of the memory correctable error from a burst perspective and / or a DQ perspective. an aggregation module, configured to aggregate the plurality of ECC error features based on an error position feature in the ECC error feature to obtain an aggregated error feature; A prediction module is used to predict whether the device to be predicted will have an uncorrectable memory error based on the aggregated error characteristics.

12. An electronic device, characterized in that: include: A memory, a processor; wherein the memory is used to store one or more computer instructions, wherein when the one or more computer instructions are executed by the processor, the method according to any one of claims 1 to 10 is implemented.

13. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed, the method according to any one of claims 1 to 10 is implemented.

Citation Information

Patent Citations

  • Page offline based on fault aware prediction of impending memory errors

    CN115910177A

  • Page offlining based on fault-aware prediction of imminent memory error

    US20220050603A1