Memory fault prediction method and related device
By filtering out the fault information of the bank granularity and training the machine learning model, the problem of low memory fault prediction accuracy in the existing technology is solved, and higher prediction accuracy and accuracy are achieved.
Patent Information
- Application Number
- CN202311641341.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-05-30
AI Technical Summary
The prediction accuracy of existing memory fault prediction methods is low, making it difficult to effectively predict the occurrence of uncorrectable errors (UCEs).
By filtering out bank granularity fault information from the fault log of the memory module and using this fault information to train the machine learning model, it can achieve accurate prediction of potential UCE.
It greatly improves the accuracy and accuracy of memory failure prediction, and can achieve high accuracy while achieving high recall, reducing accidental damage to normal memory.
Smart Images

Figure CN120067922A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of memory faults, and particularly to a method for predicting memory faults and related devices. Background Art
[0002] In recent years, with the rapid development of semiconductor technology, the memory capacity has been continuously increased, and the service life of the memory has also been continuously extended. However, this has also led to a higher probability of memory failures. Memory hardware failures have become the main cause of downtime in high-performance computing (HPC) systems and cloud computing center servers, seriously affecting system performance. Therefore, during daily operation, it is necessary to monitor and analyze the memory operation in the system to predict possible memory faults or errors and take targeted measures in a timely manner, thereby improving the stability and reliability of the system.
[0003] For memory fault prediction, the commonly used method in the industry is to predict un-correctable errors (UCEs) based on the information of correctable errors (CEs) that occur in the memory. Specifically, a model is trained using the address information of the occurred CEs and UCEs. After the model is trained, it can predict the possible UCEs based on the address information of the CEs. However, the prediction accuracy of this method is currently low, and there is an urgent need for a method that can improve the prediction accuracy. Summary of the Invention
[0004] This application provides a method for predicting memory faults and related devices, which can improve the accuracy of memory fault prediction.
[0005] In a first aspect of this application, a method for predicting memory faults is provided, which can be applied to a memory fault prediction device. The method includes: obtaining fault information, where the fault information includes the physical address of a first correctable error (CE) that occurs in the first memory bank (bank) of the memory module, and the physical address of a first un-correctable error (UCE) that occurs in the memory module. The physical address of the first CE includes the row address, column address, and bank address of the first CE in the memory module, and the physical address of the first UCE includes the row address, column address, and bank address of the first UCE in the memory module; training a model based on the physical address of the first CE and the physical address of the first UCE, where the model is used to predict the physical address of a potential UCE based on the physical address of the CE.
[0006] First, read the failure logs of the memory modules in the cluster. The memory modules can be dual-inline-memory-modules (DIMMs), single-inline-memory-module (SIMMs), or small outline DIMMs (SO-DIMMs), etc. The failure logs read can be the failure logs within three months, within one year, or within three years. Specifically, no limitation is made here. The failure logs usually record the location information, timestamp, error code, and failure type information of the occurred failures. Among them, the location information includes information such as the cell, row, column, bank, bank group, and rank where the failure is located in the memory module.
[0007] After reading the failure logs, clean the data in the failure logs, remove invalid or incorrect data, and handle missing values to ensure the accuracy of subsequent analysis and modeling. Exemplarily, for a burst UCE (i.e., no CE occurred before this UCE), since this UCE is not caused by a CE, it does not belong to the data required for training, and the burst UCE is deleted. Another example is that since the current memory module with a UCE occurrence will be replaced, only the first UCE in a certain memory module is used as the prediction template, and the failure records after the first UCE are deleted.
[0008] Next, filter out the corresponding failure information from the failure logs of the memory modules. The failure information includes the physical address of the first CE that occurred in the first bank of the memory module, and the physical address of the first UCE that occurred in the memory module. The filtering method for the first CE can be to first filter out the failure information of the rank from the failure logs of the memory module, then filter out the failure information of the bank group from this rank, and finally filter out the failure information of the first bank from this bank group, so as to finally obtain the first CE that occurred in the first bank. That is, for the CEs in the failure logs, they are sliced layer by layer at the granularity of rank, bank group, and bank, so as to obtain the CE information of each bank. The first UCE is a failure that occurred in the memory module and can be all the UCEs in the failure logs.
[0009] After obtaining the fault information, the physical addresses of the first CE and the first UCE are used as training data to train a machine learning model. Specifically, the row address, column address, and bank address of the first CE in the memory module are used as training samples, and the row address, column address, and bank address of the first UCE in the memory module are used as labels to train the machine learning model. The machine learning model can be a Convolutional Neural Network (CNN) or a Recurrent Neural Network (RNN), etc.
[0010] Before training, the training data needs to be converted into a format that the machine learning model can handle. For example, the training data is converted into numerical data and then normalized or standardized to ensure the stability and effectiveness of the model. After converting the training data, the training data can be divided into a training set and a test set. The training set is used for model training, and the test set is used to evaluate the performance and generalization ability of the model.
[0011] After training the model, for newly detected CEs, their physical addresses are used as model inputs, and the model can predict the physical addresses of potential UCEs (i.e., UCEs that may occur).
[0012] In the first aspect of this application, by screening out fault information at the bank granularity to train the model, the prediction accuracy at the bank level can be achieved, that is, it can predict whether a UCE will occur in a certain area on the bank, thereby greatly improving the prediction accuracy and precision. Moreover, with the improvement of prediction accuracy and precision, high recall rate and high accuracy can be achieved simultaneously, reducing the misjudgment of normal memory.
[0013] In a possible implementation manner of the first aspect, the method further includes: determining a fault mode based on the physical address of the potential UCE and the distribution of the physical address of the potential UCE; determining an operation and maintenance strategy based on the fault mode, and the operation and maintenance strategy is used to process the potential UCE.
[0014] In this possible implementation, the fault mode can be determined based on the physical addresses predicted by the model and the distribution of the physical addresses. Exemplarily, when only one physical address where a UCE may occur is predicted, the corresponding fault mode is the single-point error mode. When multiple physical addresses where a UCE may occur are predicted and the multiple physical addresses are in the same row in the bank, the fault mode is the row error mode. When multiple addresses are in the same column in the bank, the fault mode is the column error mode. When it is predicted that there are multiple physical addresses where a UCE may occur in a bank, but the distribution of the multiple physical addresses has no obvious pattern, the fault mode is the disordered error mode. When the predicted multiple physical addresses are distributed in the same column in different banks, the fault mode is the bank-level column error mode. By identifying the fault mode, the fault type of the memory module can be judged more accurately, so as to achieve refined fault tolerance processing, avoid unnecessary replacement of the memory module, and save operation costs.
[0015] In a possible implementation of the first aspect, the method further includes: converting the physical address of the potential UCE into the system address of the operating system; determining the operation and maintenance strategy based on the area where the system address is located in the operating system.
[0016] In this possible implementation, the operation and maintenance strategy is determined based on the area where the potential UCE is located in the operating system. Specifically, the physical address output by the model (i.e., the address of the potential UCE) is converted into the system address of the operating system. For example, the physical address of the memory module is parsed and calculated through the application programming interface (API) of the Basic Input Output System (BIOS) to obtain the system address of the operating system. The operating system can be divided into three areas, namely the kernel area, the infrastructure area, and the service area. If the potential UCE is in the kernel area, it can be processed through page isolation or memory mirroring. If the possible UCE is in the infrastructure area, fault tolerance can be achieved by performing a re-pulling operation on the management process in advance. Or when the memory error type is not serious, the system service can be allowed to self-heal. If the possible UCE is in the service area, risk address pass-through or early hot migration of the virtual machine can be performed to cope with the possible service suspension. Through the above various refined fault tolerance strategies, the accuracy of fault tolerance processing can be improved, while avoiding the occurrence of UCE as much as possible, reducing the impact on the operation of the operating system, and reducing unnecessary replacement of the memory module.
[0017] In a possible implementation of the first aspect, the above steps: training a model based on the physical addresses of the first CE and the first UCE include: dividing the first bank into multiple regions, where the multiple regions include a first region; determining a second CE in the first region in the first CE based on the physical address of the first CE; determining a second UCE in the first region in the first UCE based on the physical address of the first UCE; using the physical address of the second CE as a training sample and the physical address of the second UCE as a label to train the model.
[0018] Divide the first bank into multiple small regions, and use the address information of the CE and UCE in each small region to train the model, so as to improve the accuracy and precision of the fault. Specifically, for the first region among the multiple small regions, screen out the second CE belonging to the first region from the first CE, and screen out the second UCE belonging to the first region from the first UCE. Use the physical address of the second CE as a training sample and the physical address of the second UCE as a label to train the model. The accuracy of the trained model is approximately the range of the small region.
[0019] The range of each divided region needs to comprehensively consider the distribution of the first CE and the possible operation and maintenance fault tolerance means. The range of each region can be a cell range of 200*100, or a cell range of 150*50, etc., and specific details are not limited here. It can be understood that the range of each region cannot be too small, because if the sampling range is too small, the number of samples may be insufficient. When dividing the first bank, certain first CEs can be selected first, and with the selected first CEs as the center, determine the range of the divided region.
[0020] In this possible implementation, by dividing the fault information in the bank into multiple regions in the spatial dimension and using the fault information in each region to train the model respectively, the accuracy and precision of the fault prediction can be further improved.
[0021] In a possible implementation of the first aspect, the fault information further includes the timestamps of the first CE and the first UCE. The above steps: training a model based on the physical addresses of the first CE and the first UCE include: determining a third CE in the first CE that occurred in the first time period based on the timestamp of the first CE; determining a third UCE in the first UCE that occurred in the first time period based on the timestamp of the first UCE; using the physical address of the third CE as a training sample and the physical address of the third UCE as a label to train the model.
[0022] The fault information includes the timestamp of the occurred fault. The training data is divided at certain time intervals based on the timestamp of the fault, and then the training data belonging to a certain time period is used to train the model, thereby further improving the accuracy of the model. Specifically, based on the timestamp of the first CE, the third CE that occurred in the first time period is filtered out from the first CE, and then based on the timestamp of the first UCE, the third UCE that occurred in the first time period is filtered out from the first UCE. The physical address of the third CE is used as the training sample, and the physical address of the third UCE is used as the label to train the model.
[0023] In this possible implementation, dividing the training data in the time dimension can further improve the accuracy and precision of fault prediction.
[0024] In a possible implementation of the first aspect, the duration of the first time period is not greater than one hour.
[0025] In this possible implementation, it is specified that the duration of the first time period is not greater than one hour. For example, the duration of the first time period is 10 seconds, 30 seconds, 15 minutes, 30 minutes, 50 minutes, or 1 hour, etc. Taking a smaller granularity for the first time period can make the prediction accuracy of the model higher, the prediction result more time-sensitive, and it also has more practical application significance for fault tolerance processing. Of course, the duration of the first time period can also be two hours or three hours, etc., but it cannot be a too large time range, such as one day, 5 days, or 10 days. Generally speaking, the duration of the first time period needs to comprehensively consider the business tide characteristics and the lead of fault tolerance measures, etc., but in the case of a granularity at the minute level or hour level, fault tolerance processing can be better carried out.
[0026] In a possible implementation of the first aspect, the above step: training the model based on the physical address of the first CE and the physical address of the first UCE includes: performing image processing on the physical address of the first CE to obtain the first distribution map sequence; performing image processing on the physical address of the first UCE to obtain the second distribution map sequence; using the first distribution map sequence as the training sample and the second distribution map sequence as the label to train the model.
[0027] In this possible implementation, image processing is performed on the training data to convert the training data into a distribution map sequence. Specifically, image processing is performed on the physical address of the first CE to obtain the first distribution map sequence. Then image processing is performed on the physical address of the first UCE to obtain the second distribution map sequence. Using the first distribution map sequence as the training sample and the second distribution map sequence as the label to train the model. By performing image processing on the training data, the training efficiency of the model can be improved.
[0028] In a possible implementation of the first aspect, the bank address of the first UCE belongs to the first bank.
[0029] In this possible implementation, the first UCE occurs in the first bank. Since CE is more likely to cause a nearby UCE to occur, restricting the first UCE to the first bank can further improve the accuracy and precision of fault prediction.
[0030] A memory fault prediction device is provided in the second aspect of the present application, including an acquisition unit and a training unit. The acquisition unit is configured to acquire fault information, where the fault information includes the physical address of the first correctable error (CE) that occurs in the first memory bank of the memory module, and the physical address of the first uncorrectable error (UCE) that occurs in the memory module. The physical address of the first CE includes the row address, column address, and bank address of the first CE in the memory module, and the physical address of the first UCE includes the row address, column address, and bank address of the first UCE in the memory module. The training unit is configured to train a model based on the physical address of the first CE and the physical address of the first UCE, and the model is used to predict the physical address of a potential UCE based on the physical address of the CE.
[0031] In a possible implementation of the second aspect, the device further includes a determination unit configured to determine a fault mode based on the physical address of the potential UCE and the distribution of the physical address of the potential UCE. The determination unit is further configured to determine an operation and maintenance strategy based on the fault mode, and the operation and maintenance strategy is used to process the potential UCE.
[0032] In a possible implementation of the second aspect, the device further includes a conversion unit configured to convert the physical address of the potential UCE into a system address of the operating system. The determination unit is further configured to determine an operation and maintenance strategy based on the area where the system address is located in the operating system.
[0033] In a possible implementation of the second aspect, the training unit is specifically configured to divide the first bank into multiple regions, where the multiple regions include a first region; determine a second CE in the first CE that is in the first region based on the physical address of the first CE; determine a second UCE in the first UCE that is in the first region based on the physical address of the first UCE; use the physical address of the second CE as a training sample and the physical address of the second UCE as a label to train the model.
[0034] In a possible implementation of the second aspect, the fault information further includes the timestamps of the first CE and the first UCE. Specifically, the training unit is configured to determine a third CE that occurred in the first time period in the first CE based on the timestamp of the first CE; determine a third UCE that occurred in the first time period in the first UCE based on the timestamp of the first UCE; use the physical address of the third CE as a training sample and the physical address of the third UCE as a label to train the model.
[0035] In a possible implementation of the second aspect, the duration of the first time period is not greater than one hour.
[0036] In a possible implementation of the second aspect, the training unit is specifically configured to perform image processing on the physical address of the first CE to obtain a first distribution map sequence; perform image processing on the physical address of the first UCE to obtain a second distribution map sequence; use the first distribution map sequence as a training sample and the second distribution map sequence as a label to train the model.
[0037] In a possible implementation of the second aspect, the bank address of the first UCE belongs to the first bank.
[0038] The memory fault prediction device provided in the second aspect of the present application is used to execute the method described in the first aspect or any possible implementation of the first aspect.
[0039] A memory fault prediction device is provided in the third aspect of the present application, including a processor and a memory. The memory is used to store instructions, and the processor is used to obtain the instructions stored in the memory to execute the method described in the first aspect or any possible implementation of the first aspect.
[0040] A computer-readable storage medium is provided in the fourth aspect of the present application. The computer-readable storage medium includes instructions that, when run on a computer, cause the computer to execute the method described in the first aspect or any possible implementation of the first aspect.
[0041] A computer program product containing instructions is provided in the fifth aspect of the present application. When the computer program product runs on a computer, it causes the computer to execute the method described in the first aspect or any possible implementation of the first aspect.
[0042] A chip system is provided in the sixth aspect of the present application. The chip system includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is used to run a computer program or instructions to execute the method described in the first aspect or any possible implementation of the first aspect. Description of the Drawings
[0043] Figure 1 A schematic diagram of a system architecture for the memory fault prediction method provided by an embodiment of this application;
[0044] Figure 2 A schematic diagram of an embodiment of the memory fault prediction method provided by an embodiment of this application;
[0045] Figure 3a A schematic diagram of a fault mode in an embodiment of this application;
[0046] Figure 3b Another schematic diagram of a fault mode in an embodiment of this application;
[0047] Figure 3c Another schematic diagram of a fault mode in an embodiment of this application;
[0048] Figure 3d Another schematic diagram of a fault mode in an embodiment of this application;
[0049] Figure 3e Another schematic diagram of a fault mode in an embodiment of this application;
[0050] Figure 4 A schematic diagram of determining an operation and maintenance strategy in an embodiment of this application;
[0051] Figure 5 Another schematic diagram of an embodiment of the memory fault prediction method provided by an embodiment of this application;
[0052] Figure 6 A schematic diagram of dividing training data in space in an embodiment of this application;
[0053] Figure 7 A schematic diagram of dividing training data in the time dimension in an embodiment of this application;
[0054] Figure 8 A schematic diagram of the model training process in an embodiment of this application;
[0055] Figure 9 A schematic diagram of predicting potential UCE through a model in an embodiment of this application;
[0056] Figure 10 Another schematic diagram of an embodiment of the memory fault prediction method provided by an embodiment of this application;
[0057] Figure 11 A schematic diagram of a structure of the memory fault prediction device provided by an embodiment of this application;
[0058] Figure 12 Another schematic diagram of a structure of the memory fault prediction device provided by an embodiment of this application. Detailed implementation manners
[0059] The embodiments of the present application provide a method for predicting memory faults, which can improve the accuracy of memory fault prediction. The embodiments of the present application also provide corresponding devices, computer-readable storage media, computer program products, etc. The following will be described separately.
[0060] The embodiments of the present application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Those of ordinary skill in the art can know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0061] Terms such as "system" and "network" in the specification, claims and the above-mentioned drawings of the present application can be used interchangeably. Unless otherwise specified, ordinal numbers such as "first" and "second" are used to distinguish multiple objects and are not used to limit the order, time sequence, priority or importance of multiple objects. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0062] Please refer to the following Figure 1 , Figure 1 which is a schematic diagram of a system architecture for the memory fault prediction method provided by the embodiments of the present application.
[0063] For memory fault prediction, a commonly used method in the industry is to predict the information of UCE through the information of CE. As Figure 1 shown, first obtain the information of CE and UCE in the fault log, and then use the information of CE as the training sample and the information of UCE as the label to train an Artificial Intelligence model. After the model is trained, the next reported CE can be used as the input to predict the potential high-risk address, that is, the address information where UCE may occur.
[0064] However, for Figure 1In the current prediction methods shown, most of them are based on the system node level or the memory module level within the node as the granularity for prediction. In other words, it can only predict that a certain node or a certain memory module may have a UCE. The prediction range is too large and the accuracy is low. Correspondingly, the processing method can only be to replace the memory module that is predicted to have a UCE. However, in fact, only a certain memory bank in the memory module may fail, or a certain memory cell in the memory module may fail. Replacing the memory module easily causes waste of resources.
[0065] Moreover, the time granularity of applying this method for prediction is usually large, often in days. For example, it predicts whether a UCE will occur in the memory within the next 1 day or within 15 days. This kind of prediction with a long time granularity does not have much application prospect.
[0066] In view of this, the embodiments of the present application provide a memory fault prediction method. By screening out the fault information at the bank granularity from the fault logs of the memory module and using this fault information to train the model, the accuracy of fault prediction can be greatly improved.
[0067] Please refer to the following Figure 2 , Figure 2 which is a schematic diagram of an embodiment of the memory fault prediction method provided by the embodiments of the present application. As Figure 2 shown, this embodiment includes steps 201 to 204.
[0068] 201. Obtain fault information, where the fault information includes the physical address of the first CE that occurs in the first bank of the memory module, and the physical address of the first UCE that occurs in the memory module.
[0069] First, read the fault logs of the memory modules in the cluster. The memory modules can be dual-inline-memory-modules (DIMMs), single-inline-memory-modules (SIMMs), or small outline DIMMs (SO-DIMMs), etc. The read fault logs can be the fault logs within three months, within one year, or within three years. There is no specific limitation here. The fault logs usually record the location information, timestamp, error code, and fault type information of the occurred faults. Among them, the location information includes information such as the cell, row, column, bank, bank group, and rank where the fault is located in the memory module.
[0070] After reading the fault log, clean the data in the fault log, remove invalid or incorrect data, and handle missing values to ensure the accuracy of subsequent analysis and modeling. Exemplarily, for a sudden UCE (i.e., no CE occurred before this UCE), since this UCE is not caused by a CE, it does not belong to the data required for training, and the sudden UCE is deleted. For another example, since the memory module with a UCE occurrence is currently replaced, only the first UCE that occurs in a certain memory module is used as a prediction template, and the fault records after the first UCE are deleted.
[0071] Next, screen out the corresponding fault information from the fault log of the memory module. The fault information includes the physical address of the first CE that occurs in the first bank of the memory module and the physical address of the first UCE that occurs in the memory module. The screening method for the first CE can be to first screen out the fault information of the rank from the fault log of the memory module, then screen out the fault information of the bank group from this rank, and finally screen out the fault information of the first bank from this bank group, so as to obtain the first CE that occurs in the first bank. That is, for the CE in the fault log, it is sliced layer by layer with the granularity of rank, bank group, and bank, so as to obtain the CE information of each bank. The first UCE is a fault that occurs in the memory module and can be all the UCEs in the fault log. In a possible solution, to further improve the prediction accuracy, the first UCE can also be the first bank that occurs in the memory module. At this time, the first UCE also needs to be screened, and the screening method is similar to that of the first CE.
[0072] 202. Train a model based on the physical address of the first CE and the physical address of the first UCE.
[0073] After obtaining the fault information, use the physical address of the first CE and the physical address of the first UCE as training data to train a machine learning model. Specifically, use the row address, column address, and bank address of the first CE in the memory module as training samples, and use the row address, column address, and bank address of the first UCE in the memory module as labels to train the machine learning model. In a possible solution, the machine learning model is a deep learning model suitable for processing image sequences, such as a Convolutional Neural Network (CNN) or a Recurrent Neural Network (RNN).
[0074] Before training, the training data needs to be converted into a format that can be processed by the machine learning model. For example, convert the training data into numerical data and then perform normalization or standardization processing to ensure the stability and effectiveness of the model. After converting the training data, the training data can be divided into a training set and a test set. The training set is used for model training, and the test set is used to evaluate the performance and generalization ability of the model.
[0075] In a possible solution, the first bank is divided into multiple small regions, and the address information of the CE and UCE in each small region is used to train the model, which can further improve the accuracy and precision of fault prediction. Specifically, for the first region among the multiple small regions, the second CE belonging to the first region is selected from the first CE, and the second UCE belonging to the first region is selected from the first UCE. The physical address of the second CE is used as the training sample, and the physical address of the second UCE is used as the label to train the model. The accuracy of the trained model is approximately the range of the small region.
[0076] The range of each divided region needs to comprehensively consider the distribution of the first CE and the possible operation and maintenance fault tolerance means. The range of each region can be a cell range of 200*100, or a cell range of 150*50, etc., and no specific limitation is made here. It can be understood that the range of each small region divided cannot be too small, because if the sampling range is too small, the sample quantity may be insufficient. When dividing the first bank, some first CEs can be selected first, and with the selected first CEs as the center, the range of the divided region is determined.
[0077] In a possible solution, the fault information includes the timestamp of the occurred fault. Based on the timestamp of the fault, the training data is divided at a certain time interval, and then the training data belonging to the same time period is used to train the model, which can further improve the accuracy of the model. Specifically, based on the timestamp of the first CE, the third CE that occurred in the first time period is selected from the first CE, and based on the timestamp of the first UCE, the third UCE that occurred in the first time period is selected from the first UCE. The physical address of the third CE is used as the training sample, and the physical address of the third UCE is used as the label to train the model.
[0078] Optionally, the duration of the first time period is not greater than one hour. For example, the duration of the first time period is 10 seconds, 30 seconds, 15 minutes, 30 minutes, 50 minutes, or 1 hour, etc. Making the granularity of the first time period smaller can make the prediction accuracy of the model higher, the prediction result more timely, and it also has more practical application significance for fault tolerance processing. Of course, the duration of the first time period can also be two hours or three hours, etc., but it cannot be a too large time range, such as one day, 5 days, or 10 days. Generally speaking, the duration of the first time period needs to comprehensively consider the characteristics of business tides and the lead of fault tolerance measures, etc., but when the granularity is at the minute level or hour level, fault tolerance processing can be better carried out.
[0079] It can be understood that the temporal division of the training data and the regional division can be combined, that is, the third CE and the third UCE that occur during the first time period are screened out from the second CE and the second UCE in the first region.
[0080] In a possible solution, the training data is processed graphically to convert the training data into a sequence of distribution maps. Specifically, the physical addresses of the first CE are processed graphically to obtain the first sequence of distribution maps. Then, the physical addresses of the first UCE are processed graphically to obtain the second sequence of distribution maps. The first sequence of distribution maps is used as the training sample, and the second sequence of distribution maps is used as the label to train the model.
[0081] It can be understood that for the above possible solutions, they can be applied in combination or separately.
[0082] After the model is trained, the address of the potential UCE can be predicted based on the next occurrence of the CE.
[0083] 203. Determine the fault mode based on the predicted UCE.
[0084] Furthermore, the fault mode can be determined based on the physical address predicted by the model and the distribution of the physical addresses, etc. The fault mode refers to the characteristic mode of a specific type of fault. For example, when only one physical address is predicted, the corresponding fault mode is the single-point error mode. When there are multiple predicted physical addresses and they are in the same row in the bank, the fault mode is the row error mode. When multiple addresses are in the same column in the bank, the fault mode is the column error mode. By identifying the fault mode, the fault type of the memory module can be judged more accurately and targeted processing can be carried out.
[0085] Specifically, first obtain the relevant information of the occurred UCE from the fault log of the memory module, including specifically the physical address of the UCE and the corresponding fault type, etc. Then train a machine learning model through the obtained UCE information. Among them, the physical address and distribution of the UCE, etc. are used as the input of the model, and the fault type is used as the output of the model. After the machine learning model is trained, the corresponding fault mode can be output according to the input UCE information. Among them, the machine learning model used to predict the fault mode can be support vector machines (SVM), Decision Tree, and Random Forest, etc. In addition to the above single-point error mode, row error mode, and column error mode, the fault mode also includes bank-level column error mode, disordered error mode, memory access exception, and Error Correcting Code (ECC) error, etc. The fault mode can be combined with Figures 3a to 3e for understanding. Figure 3a That is, the single-point error mode, where only one cell in the bank fails. Figure 3b It is the row error mode, where multiple cells in a row fail. Figure 3c It is the column error mode, where multiple cells in a column fail. Figure 3d It is the disordered fault mode, that is, multiple cells in the bank all fail, but there is no distribution rule. Figure 3e It is the bank-level column error mode, that is, the fault occurs in the same column of multiple banks.
[0086] 204. Determine the operation and maintenance strategy, which is used to perform fault tolerance processing for potential UCE.
[0087] In a possible solution, determine the operation and maintenance strategy based on the fault mode, and perform targeted fault tolerance processing according to the operation and maintenance strategy. For example, when the fault mode is a row error, mask the entire row where the error occurs and avoid accessing the cells at the address of that row. When the fault mode is a column error, since the range of the column is large, the entire bank where the column is located is masked. When the output fault mode is the bank-level column error, then the entire bank group where the bank is located is masked.
[0088] In another possible solution, the operation and maintenance strategy is determined based on the area where the potential risk area is located in the operating system. Specifically, the physical address output by the model (i.e., the address where UCE may occur) is converted into the system address of the operating system. For example, the physical address of the memory module is resolved through the Application Programming Interface (API) of the Basic Input Output System (BIOS) to obtain the system address of the operating system. The operating system can be divided into three areas, namely the kernel area, the infrastructure area, and the business area. If the possible UCE is in the kernel area, it can be processed through page isolation or memory mirroring; if the possible UCE is in the infrastructure area, fault tolerance can be achieved by re-pulling the management process in advance. Or when the memory error type is not serious, the system service can be allowed to self-heal. If the possible UCE is in the business area, risk address pass-through or early hot migration of the virtual machine can be performed to cope with the possible service suspension. Through the above various refined fault tolerance strategies, high availability without user perception is achieved.
[0089] In addition, the operation and maintenance strategy can also be determined by combining the fault mode and the area where the UCE is located in the operating system. For example, for the UCE in the kernel area, when the corresponding fault mode is a row fault, the operation and maintenance strategy can be page isolation or memory mirroring processing; when the fault mode is a column fault, the bank where the column is located needs to be masked. Considering the two situations together can further achieve refined fault tolerance.
[0090] After obtaining the fault mode and the operation and maintenance strategy, they can be reported to the operation and maintenance personnel, and the operation and maintenance personnel can respond quickly based on the operation and maintenance strategy and take corresponding fault tolerance measures. For the memory module that needs to be replaced, it can also be replaced according to the prediction result to ensure the stability and reliability of the operating system.
[0091] The following combines Figure 4 to give an exemplary description of the process of determining the operation and maintenance strategy. As Figure 4 shown, first, the physical address of the potential UCE output by the model and the historical fault information are fused and input into the fault mode judgment model to obtain the corresponding fault mode. Through the historical fault information, it can be judged whether the address has had multiple faults. If it has had multiple faults, it means that the address has a relatively high probability of failure. Then, it is judged whether the fault mode is a column error or a bank fault. If so, the hardware analog to digital converter (ADC) bank is replaced, the virtual machine for the service is hot migrated, and new services are no longer issued to the bank. And the memory module is marked as in a sub-healthy state and waits for operation and maintenance replacement.
[0092] If the failure mode is not a column error or a bank failure, then determine whether the failure mode is a row error or a single point error. If so, perform a post package repair (PPR) replacement on the potential risk area or a software page offline process. When a failure occurs in a certain area of the memory, the hardware PPR replacement can be an effective solution. In the PPR replacement, the failed area will be marked, and the hardware will move the data in that area to a spare fault-tolerant area or swap it to other available memory areas. In this way, the impact of the failed area will be limited, and the system can continue to run normally. The hardware PPR replacement is the above-mentioned memory mirroring process. The software page offline is another method for handling memory failures. After detecting the memory failure area, the pages in that area can be marked as offline. This means that the failed area will no longer be used, will not be allocated to new tasks or processes, and the existing data will be migrated or discarded. The software page offline is a means of isolating and protecting the failed area to prevent the failure from spreading to other parts and ensuring the stable operation of the system. The software page offline is the above-mentioned page isolation operation.
[0093] After processing the potential risk area, perform a stress test on the potential risk area to check its health. If other failure modes occur during the stress test, mark the memory module where the potential risk area is located as in a sub-healthy state and wait for replacement. If no other failure modes occur during the stress test, mark the memory module as in a normal use state.
[0094] In addition, the historical failure information is simultaneously sent into the machine learning model for memory module failure prediction. If a failure is found in the memory module, trigger a business hot migration and stop issuing new services. Then mark the memory module as a sub-healthy memory module and wait for replacement.
[0095] In this embodiment, by screening out the failure information at the bank granularity to train the model, the prediction accuracy at the bank level can be achieved, that is, it can predict whether a certain area on the bank will have a UCE, greatly improving the prediction accuracy and accuracy. And with the improvement of the prediction accuracy and accuracy, a high accuracy rate can be achieved while achieving a high recall rate, reducing the misjudgment of normal memory. The recall rate is the processing rate of the failed memory. Currently, to achieve a high recall rate, due to the low prediction accuracy, a large amount of normal memory will be misjudged.
[0096] By further segmenting the training data in combination with the time dimension, the dynamic changes of faults in both time and space can be captured, further improving the accuracy of fault prediction. Moreover, in this embodiment, the preferred time interval is at the hour level. That is to say, UCE prediction can be performed at the granularity of hours, and the prediction results are more timely, which has more practical application significance for fault tolerance processing.
[0097] In addition, in this embodiment, based on the predicted UCE information, the fault mode and the operation and maintenance strategy are further determined, which can achieve refined fault tolerance, reduce the replacement of memory modules, and save costs. In this embodiment, with the help of the prediction accuracy at the bank level, the specific fault mode can be determined. For single-point faults and row faults, etc., they can be processed through page isolation and other methods, so as to avoid unnecessary replacement of memory modules, improving the replacement efficiency and cost control.
[0098] Next, in combination with Figure 5 an exemplary description will be given of the memory fault prediction method provided in the embodiments of the present application. As Figure 5 shown, this embodiment includes steps 501 to 505.
[0099] 501. Preprocess the data in the fault log.
[0100] When a CE occurs in the memory is detected, read the fault log of the memory module where the CE is located, and filter out the fault information of the first bank where the CE is located from it.
[0101] Perform image processing on the spatial features (i.e., physical addresses, including row addresses, column addresses, and bank addresses) of the first CE and the first UCE in the fault information to obtain a distribution map sequence of the first CE and the first UCE. The process of image processing can be understood in combination with Figure 6 . As Figure 6 shown in part a of, first distribute the information of the first CE and the first UCE included in the fault information of the first bank in space. Then use a bounding box to segment the fault information to obtain multiple regions (i.e., multiple matrices) shown in part b. According to the rules of spatial transformation, count the information in each region. The fault distribution of the first region is shown in part C, where the hollow part is the UCE fault (i.e., the second UCE), and the solid part is the CE fault (i.e., the first UCE). From Figure 6 part c of, it can be seen that the coverage range of the first region is 20*60.
[0102] Furthermore, the fault information can also be divided in the time dimension. For the division process, please refer to Figure 7 . As Figure 7As shown, where the horizontal axis is the column address, the vertical axis is the time, and the vertical axis is the row address. After dividing the fault information into multiple regions, the multiple regions are further divided at one-hour time intervals, and finally a distribution map sequence of each region in multiple time periods is obtained, and the duration of each time period is one hour. Specifically, assume that for Figure 6 each matrix obtained by dividing in 0 , it is divided into n sub-intervals [w 1 , w n-1 in the time dimension. Let the starting time of the matrix be T o , then w i records the fault information in the time interval [T o+i , T o+1+i . w i is a matrix with the same size as the multiple matrices divided in Figure 6 , and the fault record format in the matrix is the same as that of the matrix divided in Figure 6 .
[0103] 502. Train the model based on the fault distribution map sequence.
[0104] Train the model through the distribution map sequence of the first CE and the first UCE obtained in step 501. The trained model can predict the possible location of UCE according to the input CE information. Specifically, for the newly detected CE, divide the bank where the CE is located as described in step 501 to obtain a potential risk area with a range of 20 * 60. Use two tensors of 60 * 20 to represent the CE and UCE information in the potential risk area in the first time period respectively, that is, one is used to represent the CE as a training sample, and the other is used to represent the UCE as a label. The model training process can be understood in combination with Figure 8 . Figure 8 The model in Figure 8As shown, the tensor containing CE and UCE information is input into the E3D-LSTM. The E3D-LSTM learns the error information of each CE region, which can capture the dynamic changes in time and local features in space, so as to better predict potential faults. Finally, the prediction result is output through the linear layer. Specifically, first, the fault distribution sequence diagrams of different time periods obtained by partitioning (i.e., frames) are input into the three-dimensional convolutional neural network encoder (3D CNN encoder). After being processed by the 3D CNN encoder, the data is input into the fault prediction model E3D-LSTM, and the E3D-LSTM then transfers the output prediction result to the three-dimensional convolutional neural network decoder (3D CNN decoder) and the classifier.
[0105] 503. Predict potential risk areas through the model.
[0106] After the model is trained, the newly detected CE is input into the fault prediction model, and the model can predict the positions where UCE may occur. Specifically, for the newly detected CE, let the maximum row distance be R and the maximum column distance be C. Taking the row and column positions (r, c) of this CE as the center, a matrix of size [2*R + 1, 2*C + 1, 2] is generated. For other positions (r1, c1) on the same bank where faults have occurred, the following processing is carried out:
[0107] 1. If abs(r - r1) < R and abs(c - c1) < C, the relative position of this position in the matrix is p = (r - r1, c - c1). Calculate the number of CE and UCE occurrences at this position and record them at the positions [p[0], p[1], 0] and [p[0], p[1], 1] of the matrix respectively.
[0108] 2. If abs(r - r1) < R and abs(c - c1) >= C, the relative position of this position in the matrix is p = (r - r1, Min(2*C, Max(c - c1, 0))). Calculate the number of CE and UCE occurrences at this position and accumulate them to the positions [p[0], p[1], 0] and [p[0], p[1], 1] of the matrix respectively.
[0109] 3. If abs(r - r1) >= R and abs(c - c1) < C, the relative position of this position in the matrix is p = (Min(2*R, Max(r - r1, 0), c - c1). Calculate the number of CE and UCE occurrences at this position and accumulate them to the positions [p[0], p[1], 0] and [p[0], p[1], 1] of the matrix respectively.
[0110] Finally, the generated matrix is input into the trained model to determine whether a UCE fault will occur within the range of this matrix and the physical address where a UCE may occur.
[0111] The specific process of the model predicting UCE for the input CE can be combined with Figure 9 for understanding. As Figure 9 shown, first perform feature sampling on the input CE, and then input it into the fault prediction model. If the fault prediction model outputs False, it means that this CE will not cause a UCE; if it outputs True, it means that a UCE will occur, and the physical address where the predicted UCE may occur will be output.
[0112] 504. Determine the fault mode based on the potential risk area.
[0113] After the model outputs the physical address of the UCE (i.e., the potential risk area), determine the fault mode based on the physical address of the UCE. Specifically, first establish a fault mode prediction model according to historical fault information. The fault mode can be classified according to the spatial locality of the fault distribution. The fault mode prediction model can be a random forest model. After the model is established, input the information of the possible UCE into the model to obtain the corresponding fault mode. Specifically, it has been described in detail in step 203 of the embodiment shown in Figure 2 and will not be elaborated here.
[0114] 505. Determine the operation and maintenance strategy and perform targeted fault tolerance.
[0115] After determining the fault mode, the specific operation and maintenance fault tolerance solution can be determined according to the fault mode. Specifically, it has been described in detail in step 204 of the embodiment shown in Figure 2 and will not be elaborated here.
[0116] In summary, the memory fault prediction method provided by the embodiments of the present application will be described as a whole in combination with Figure 10 below.
[0117] As Figure 10As shown, the embodiment of the present application mainly includes two stages, wherein the first stage is the memory fault prediction stage, which is used to predict possible UCE. Specifically, when a memory fault (including CE and UCE) occurs, the memory module reports the coordinate information of the fault in the memory module through the advanced configuration and power management interface (ACPI), and the coordinate information includes the central processing unit (CPU) rocket, memory module, rank, bank group, bank, column and row where the CE and UCE are located. The coordinate information of the bank granularity, i.e., row, column and bank, is selected, and the coordinate information is segmented according to a certain time interval in combination with the timeline, and the deep learning model is trained using the final coordinate information. After the model is trained, the information of potential UCE can be predicted based on the coordinate information of CE.
[0118] For the CE that reports the interrupt, the row and column information of the CE in the bank to which it belongs is obtained, that is, the row address, column address and bank address of the CE are obtained. Optionally, after obtaining the row and column information, it can be first input into the convolutional neural network. The convolution operation is usually accompanied by a pooling operation, which can reduce the amount of calculation by reducing the spatial dimension of the data. This is very helpful for reducing the complexity of the model while retaining important information. The convolution operation also has translation invariance, that is, the pattern learned by the convolution kernel can be detected at different positions of the input data. This makes it easier for the neural network to learn reliable features without being affected by its specific position in the input. In addition, the convolution operation also has a sparse interaction property, that is, each output element is only related to the local area of the input. This also helps to reduce the amount of calculation and makes the model easier to train. After the convolution operation, the input data of the model (that is, the physical address information of the CE) can also be time-windowed, that is, the input data is divided according to a certain time interval. Then, the training data of each time period is respectively input into the trained model. The model can be a recurrent neural network combined with a fully connected layer. The model can output the address information of the potential UCE for the input CE information.
[0119] The second stage is the operation and maintenance warning stage, which specifically includes determining the fault mode and operation and maintenance strategy based on the output results of the first stage. For the potential risk addresses (i.e., the physical addresses of potential UCEs) output in the first stage, further analysis is carried out in combination with historical fault information to determine the fault mode and operation and maintenance strategy. Through historical fault information, it can be judged whether a fault has occurred at this address multiple times. If a fault has occurred multiple times, it means that this address has a high probability of failure. Specifically, first collect out-of-band data through the Baseboard Management Controller (BMC), and train a machine learning model for predicting the fault mode through the out-of-band data. This model can be a support vector machine or a decision tree, etc. Then, fuse and input the addresses of potential UCEs output in the first stage and historical fault information into this model, and the corresponding fault mode can be obtained. The fault modes include single-point error mode, row error mode, column error mode, Bank-level column error mode, disordered error mode, memory access exception, and Error Correcting Code (ECC) error, etc.
[0120] Determine the operation and maintenance strategy based on the fault mode and / or the output results of the first stage. Exemplarily, if the potential UCE is in the kernel area, it can be fault-tolerant through page isolation or memory mirroring. If the potential UCE is in the infrastructure area, it can be fault-tolerant by performing a re-pull operation on the management process in advance, or waiting for the system service to self-heal. If the potential UCE is in the business area, risk address pass-through or early hot migration of the virtual machine can be performed to cope with possible service outages. Fine-grained fault tolerance is achieved through targeted operation and maintenance strategies. Combine the fault mode and operation and maintenance strategy to determine the health of the memory module, and replace the memory module when the fault is relatively serious. For example, when the fault mode is the Bank-level column error mode, the memory module is replaced.
[0121] The above has described the embodiments of the present application from the perspective of methods. Next, the related devices in the embodiments of the present application will be introduced from the perspective of specific device implementation.
[0122] Please refer to Figure 11 , a schematic diagram of a memory fault prediction device 1100 is provided in an embodiment of the present application. Among them, the memory fault prediction device 1100 includes an acquisition unit 1101 and a training unit 1102.
[0123] An acquisition unit 1101, configured to acquire fault information, where the fault information includes the physical address of a first correctable error (CE) that occurs in a first memory bank in a memory module, and the physical address of a first uncorrectable error (UCE) that occurs in the memory module. The physical address of the first CE includes the row address, column address, and bank address of the first CE in the memory module, and the physical address of the first UCE includes the row address, column address, and bank address of the first UCE in the memory module.
[0124] A training unit 1102, configured to train a model based on the physical addresses of the first CE and the first UCE. The model is used to predict the physical address of a potential UCE based on the physical address of a CE.
[0125] Optionally, the memory fault prediction device 1100 further includes a determination unit 1103, configured to determine a fault mode based on the physical address of the potential UCE and the distribution of the physical addresses of the potential UCE; the determination unit 1103 is further configured to determine an operation and maintenance strategy based on the fault mode, where the operation and maintenance strategy is used to process the potential UCE.
[0126] Optionally, the memory fault prediction device 1100 further includes a conversion unit 1104, configured to convert the physical address of the potential UCE into a system address of an operating system; the determination unit 1103 is further configured to determine an operation and maintenance strategy based on the area where the system address is located in the operating system.
[0127] Optionally, the training unit 1102 is specifically configured to divide the first bank into multiple regions, where the multiple regions include a first region; determine a second CE in the first region among the first CEs based on the physical address of the first CE; determine a second UCE in the first region among the first UCEs based on the physical address of the first UCE; use the physical address of the second CE as a training sample and the physical address of the second UCE as a label to train the model.
[0128] Optionally, the fault information further includes the timestamps of the first CE and the first UCE. The training unit 1102 is specifically configured to determine a third CE that occurs in a first time period among the first CEs based on the timestamp of the first CE; determine a third UCE that occurs in the first time period among the first UCEs based on the timestamp of the first UCE; use the physical address of the third CE as a training sample and the physical address of the third UCE as a label to train the model.
[0129] Optionally, the duration of the first time period is not greater than one hour.
[0130] Optionally, the training unit 1102 is specifically configured to perform imaging processing on the physical addresses of the first CEs to obtain a first sequence of distribution maps; perform imaging processing on the physical addresses of the first UCEs to obtain a second sequence of distribution maps; use the first sequence of distribution maps as training samples and the second sequence of distribution maps as labels to train a model.
[0131] Optionally, the bank address of the first UCE belongs to the first bank.
[0132] Each unit in the memory fault prediction device 1100 executes the operations of the memory fault prediction device in the foregoing Figure 2 , Figure 5 and Figure 10 shown in the embodiments, and details are not described herein again.
[0133] Next, please refer to Figure 12 , which is a possible structural schematic diagram of a memory fault prediction device 1200 provided in an embodiment of the present application, including a processor 1201, a communication interface 1202, a memory 1203, and a bus 1204. The processor 1201, the communication interface 1202, and the memory 1203 are interconnected through the bus 1204. In the embodiments of the present application, the processor 1201 is used to control and manage the actions of the memory fault prediction device. For example, the processor 1201 is used to execute Figure 2 the steps executed by the memory fault prediction device in the method embodiment shown. The communication interface 1202 is used to support the memory fault prediction device to communicate. The memory 1203 is used to store the program code and data of the memory fault prediction device.
[0134] Among them, the processor 1201 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor 1201 may also be a combination for implementing computing functions, such as a combination including one or more microprocessors, a combination of a digital signal processor and a microprocessor, and so on. The bus 1204 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 12 only a thick line is shown in
[0135] The embodiments of the present application further provide a computer-readable storage medium. The computer-readable storage medium includes instructions that, when running on a computer, cause the computer to execute the methods in the foregoing Figure 2 、 Figure 5 and Figure 10 illustrated embodiments.
[0136] The embodiments of the present application further provide a computer program product containing instructions that, when running on a computer, cause the computer to execute the methods in the foregoing Figure 2 、 Figure 5 and Figure 10 illustrated embodiments.
[0137] The embodiments of the present application further provide a chip system. The chip system includes at least one processor and a communication interface. The communication interface and the at least one processor are interconnected by a line. The at least one processor is configured to run a computer program or instructions to execute the methods in the foregoing Figure 2 、 Figure 5 and Figure 10 illustrated embodiments.
[0138] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0139] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0140] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0141] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0142] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0143] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
Claims
1. A method for predicting memory faults, characterized in that, it includes: Obtain fault information, where the fault information includes the physical address of the first correctable error (CE) that occurs in the first memory bank (bank) of the memory module, and the physical address of the first uncorrectable error (UCE) that occurs in the memory module. The physical address of the first CE includes the row address, column address, and bank address of the first CE in the memory module, and the physical address of the first UCE includes the row address, column address, and bank address of the first UCE in the memory module; Train a model based on the physical address of the first CE and the physical address of the first UCE, where the model is used to predict the physical address of a potential UCE based on the physical address of the CE.
2. The method according to claim 1, characterized in that, the method further includes: Determine a fault mode based on the physical address of the potential UCE and the distribution of the physical address of the potential UCE; Determine an operation and maintenance strategy based on the fault mode, where the operation and maintenance strategy is used to process the potential UCE.
3. The method according to claim 1 or 2, characterized in that, the method further includes: Convert the physical address of the potential UCE into the system address of the operating system; Determine the operation and maintenance strategy based on the area where the system address is located in the operating system.
4. The method according to any one of claims 1 to 3, characterized in that, the training of the model based on the physical address of the first CE and the physical address of the first UCE includes: Divide the first bank into multiple regions, where the multiple regions include a first region; Determine a second CE in the first CE that is in the first region based on the physical address of the first CE; Determine a second UCE in the first UCE that is in the first region based on the physical address of the first UCE; Use the physical address of the second CE as a training sample and the physical address of the second UCE as a label to train the model.
5. The method according to any one of claims 1 to 4, characterized in that, the fault information further includes the timestamps of the first CE and the first UCE, and the training of the model based on the physical address of the first CE and the physical address of the first UCE includes: Determine a third CE in the first CE that occurs in a first time period based on the timestamp of the first CE; Determine a third UCE in the first UCE that occurs in the first time period based on the timestamp of the first UCE; Use the physical address of the third CE as a training sample and the physical address of the third UCE as a label to train the model.
6. The method according to claim 5, characterized in that, the duration of the first time period is not more than one hour.
7. The method according to any one of claims 1 to 6, characterized in that, the training of the model based on the physical address of the first CE and the physical address of the first UCE includes: Perform image processing on the physical address of the first CE to obtain a first distribution map sequence; Image the physical address of the first UCE to obtain a second distribution map sequence; Use the first distribution map sequence as a training sample and the second distribution map sequence as a label to train the model.
8. The method according to any one of claims 1 to 7, wherein, the bank address of the first UCE belongs to the first bank.
9. A memory fault prediction device, wherein, comprises: an acquisition unit configured to acquire fault information, the fault information including the physical address of a first correctable error (CE) that occurs in a first memory bank of a memory module and the physical address of a first uncorrectable error (UCE) that occurs in the memory module, the physical address of the first CE including the row address, column address, and bank address of the first CE in the memory module, and the physical address of the first UCE including the row address, column address, and bank address of the first UCE in the memory module; a training unit configured to train a model based on the physical address of the first CE and the physical address of the first UCE, the model being used to predict the physical address of a potential UCE based on the physical address of the CE.
10. The device according to claim 9, wherein, the device further comprises: a determination unit configured to determine a fault mode based on the physical address of the potential UCE and the distribution of the physical address of the potential UCE; the determination unit is further configured to determine an operation and maintenance strategy based on the fault mode, the operation and maintenance strategy being used to process the potential UCE.
11. The device according to claim 9 or 10, wherein, the device further comprises: a conversion unit configured to convert the physical address of the potential UCE into a system address of an operating system; the determination unit is further configured to determine the operation and maintenance strategy based on the area where the system address is located in the operating system.
12. The device according to any one of claims 9 to 11, wherein, the training unit is specifically configured to: divide the first bank into multiple regions, the multiple regions including a first region; determine a second CE in the first region among the first CEs based on the physical address of the first CE; determine a second UCE in the first region among the first UCEs based on the physical address of the first UCE; use the physical address of the second CE as a training sample and the physical address of the second UCE as a label to train the model.
13. The device according to any one of claims 9 to 12, wherein, the fault information further includes the timestamps of the first CE and the first UCE, and the training unit is specifically configured to: determine a third CE that occurs in a first time period among the first CEs based on the timestamp of the first CE; determine a third UCE that occurs in the first time period among the first UCEs based on the timestamp of the first UCE; use the physical address of the third CE as a training sample and the physical address of the third UCE as a label to train the model.
14. The device according to claim 13, It is characterized in that the duration of the first time period is not greater than one hour.
15. The device according to any one of claims 9 to 14, It is characterized in that the training unit is specifically configured to: perform imaging processing on the physical address of the first CE to obtain a first distribution map sequence; perform imaging processing on the physical address of the first UCE to obtain a second distribution map sequence; use the first distribution map sequence as a training sample and the second distribution map sequence as a label to train the model.
16. The device according to any one of claims 9 to 15, It is characterized in that the bank address of the first UCE belongs to the first bank.
17. A memory fault prediction device, It is characterized in that it includes: a processor and a memory; the memory is used to store instructions; the processor is used to execute the instructions stored in the memory to implement the method according to any one of claims 1 to 8.
18. A computer-readable storage medium, on which a computer program is stored, It is characterized in that when the computer program is executed by one or more processors, the method according to any one of claims 1 to 8 is implemented.
19. A computer program product containing instructions, It is characterized in that when the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1 to 8.
Citation Information
Cited By
Memory fault prediction model training method, memory fault processing method and electronic equipment
CN121210203A