Fault prediction method and apparatus, and related device

By classifying the storage media error data and using machine learning models with high accuracy and coverage for failure prediction, the problem of poor prediction in the existing technology is solved, the overall effect of fault prediction is improved, and the impact of storage media failure on equipment business is reduced.

WO2025138982A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/115829
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-08-30
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, it is difficult for machine learning models to achieve high levels of accuracy and coverage of storage media failure prediction at the same time, resulting in a greater impact on the operation of equipment business.

Method used

By classifying the error data of the storage medium, using machine learning models with high accuracy and machine learning models with high coverage, the overall prediction effect is improved.

Benefits of technology

It achieves the simultaneously improved the accuracy and coverage of fault prediction and reduces the impact of storage media failure on equipment business.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024115829_03072025_PF_FP_ABST
    Figure CN2024115829_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A fault prediction method and apparatus, and a related device, relating to the technical field of computers. The method comprises: acquiring a plurality of pieces of error data, wherein each error data is used for indicating that a CE occurs in a storage medium; and classifying the plurality of pieces of error data, using a first machine learning model to perform fault prediction on the error data of a first category obtained by classification, and using a second machine learning model to perform fault prediction on the error data of a second category obtained by classification, so as to obtain corresponding prediction results for indicating whether a UCE occurs in the storage medium, wherein the precision rate of predicting the UCE by the first machine learning model is higher than a first threshold, and the coverage rate of predicting the UCE by the second machine learning model is higher than a second threshold. Therefore, a machine learning model having a high prediction precision rate and a machine learning model having a high coverage rate are used to perform fault prediction on the error data of different categories, so that the precision rate and coverage rate of fault prediction both can be considered at the same time, thereby effectively improving the overall prediction effect for a plurality of pieces of error data.
Need to check novelty before this filing date? Find Prior Art

Description

Fault prediction method, device and related equipment

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 29, 2023, with application number 202311871806.3 and application name “Fault Prediction Method, Device and Related Equipment”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of computer technology, and in particular to a fault prediction method, apparatus, and related equipment. Background Art

[0003] With the advancement of storage technology, the manufacturing processes of storage media such as memory are shrinking, and the read and write speeds of storage media are constantly improving. At the same time, the probability of storage media failure is also gradually increasing. For example, the impact of residual dirt on storage media is increasing, making it more likely to cause storage media failure.

[0004] Typically, devices use built-in error correction code (ECC) mechanisms to ensure the reliability of data stored on storage media. However, when errors in the storage media exceed the ECC mechanism's correction capabilities, they can result in uncorrectable errors (UCEs), potentially causing the device to freeze, reboot, or even crash, impacting service operations.

[0005] Currently, for errors reported by storage media, the error data is usually input into a machine learning model, and the machine learning model predicts whether the storage medium has a UCE based on the input error data, so that after determining that a UCE has occurred, the storage medium can be fault isolated accordingly. However, in actual application scenarios, it is difficult for the machine learning model to achieve a high level of accuracy and coverage for storage medium fault prediction at the same time. For example, when the accuracy of the machine learning model for fault prediction reaches 70%, the coverage of the machine learning model for fault prediction may only reach 20%. That is, when multiple UCEs actually occur, the machine learning model can only predict a small proportion of UCEs. Therefore, the effect of fault prediction using the machine learning model is poor, which leads to frequent failures to accurately predict storage medium failures in actual application scenarios, making it easy for the UCE caused by storage medium failures to have a significant impact on business operations on the device.

[0006] Summary of the Invention

[0007] The present application provides a fault prediction method to improve the overall prediction effect of the multiple error reporting data. In addition, the present application also provides a fault prediction device, a BMC, a computing device, a computer-readable storage medium, and a computer program product.

[0008] In a first aspect, the present application provides a fault prediction method that can be performed by a corresponding fault prediction device. Specifically, the fault prediction device obtains multiple error data, such as CE data, and each error data in the multiple error data indicates that a CE (correctable error) has occurred in a storage medium, which can be, for example, a memory or other type of storage medium. The fault prediction device then classifies the multiple error data to obtain error data of a first category and error data of a second category. For example, the multiple error data can be classified according to a preset classification rule. Finally, the fault prediction device uses the first machine learning model to perform fault prediction on the first category of error data and obtains a first prediction result. The first prediction result is used to indicate whether UCE (uncorrectable error) occurs in the storage medium. The accuracy of the first machine learning model in predicting UCE is higher than the first threshold, that is, the machine learning model with higher accuracy is used to predict faults for the first category of error data; and the fault prediction device also uses the second machine learning model to perform fault prediction on the second category of error data and obtains a second prediction result. The second prediction result is used to indicate whether UCE occurs in the storage medium. The coverage of the second machine learning model in predicting UCE is higher than the second threshold, that is, the machine learning model with higher coverage is used to predict faults for the second category of error data.

[0009] In this way, the fault prediction device classifies multiple error data corresponding to the storage medium, and uses a machine learning model with a higher prediction accuracy to predict faults for the first category of error data, and uses a machine learning model with a higher coverage to predict faults for the second category of error data. This allows the overall fault prediction results for multiple error data to not only achieve a higher accuracy, but also a higher coverage, that is, taking into account both the accuracy and coverage of fault prediction at the same time, thereby effectively improving the overall prediction effect for the multiple error data, and helping to reduce the impact of UCE occurring in the storage medium on the business operations of the device where the storage device is located.

[0010] In one possible embodiment, after the fault prediction device classifies multiple error data, the positive-to-negative sample ratio corresponding to the obtained first category of error data is smaller than the positive-to-negative sample ratio corresponding to the second category of error data, wherein the positive-to-negative sample ratio is the ratio between the number of positive samples and the number of negative samples, wherein the positive samples are error data generated when the storage medium experiences UCE, and the negative samples are error data generated when the storage medium experiences CE. In this way, using a machine learning model with a high accuracy rate, predicting whether the storage medium has experienced UCE from the first category of error data with a large proportion of positive samples can predict as many positive samples in the first category of error data as possible, that is, predicting as many actual UCEs as possible; using a machine learning model with a high coverage rate, predicting whether the storage medium has experienced UCE from the second category of error data with a small proportion of positive samples can predict as many positive samples in the second category of error data as possible, that is, the results predicted as UCE cover as many actual UCEs as possible. In this way, the fault prediction for the multiple error data takes into account both the accuracy and coverage of the fault prediction.

[0011] In one possible implementation, the first machine learning model and the second machine learning model have different model structures or different hyperparameters.

[0012] In one possible implementation, the fault prediction device can also obtain the actual result corresponding to the first category of error data, and the actual result is used to indicate that UCE or CE has occurred in the storage medium, so that the fault prediction device can update the first machine learning model according to the actual result and the first predicted result corresponding to the first category of error data; or, the fault prediction device can obtain the actual result corresponding to the second category of error data, and the actual result is used to indicate that UCE or CE has occurred in the storage medium, so that the fault prediction device can update the second machine learning model according to the actual result and the second predicted result corresponding to the second category of error data. In this way, the fault prediction device dynamically updates the machine learning model according to the predicted results and the actual results, which can further improve the accuracy of subsequent predictions of the machine learning model, thereby further improving the accuracy and coverage of fault prediction for multiple error data.

[0013] In one possible implementation, multiple error data are classified by a classifier. The fault prediction device can then obtain the true results corresponding to the first category of error data and the true results corresponding to the second category of error data, where the true results are used to indicate that a UCE or CE has occurred on the storage medium. The fault prediction device can then update the classifier based on the true results corresponding to the first category of error data, the true results corresponding to the second category of error data, the first predicted results, and the second predicted results. In this way, by dynamically updating the classifier, the fault prediction device can make the classification results of subsequent classifiers for new error data more accurate, thereby helping to further improve the accuracy and coverage of fault prediction for multiple error data.

[0014] In a possible implementation, multiple error data are classified by a classifier, which includes a first classification rule. Then, when the fault prediction device classifies the multiple error data, it can specifically use the first classification rule in the classifier to classify the multiple error data.

[0015] In one possible embodiment, multiple error data are classified by a classifier, and the classifier includes a first classification rule and a second classification rule; then, when the fault prediction device classifies the multiple error data, it can specifically use the first classification rule in the classifier to classify the multiple error data to obtain first category error data, second category error data, third category error data, and fourth category error data; accordingly, the fault prediction device can also use a third machine learning model to perform fault prediction on the third category error data to obtain a third prediction result, and the third prediction result is used to indicate whether an uncorrectable error (UCE) occurs in the storage medium; and use a fourth machine learning model to perform fault prediction on the fourth category error data to obtain a fourth prediction result, and the fourth prediction result is used to indicate whether a UCE occurs in the storage medium. In this way, the fault prediction device can use different machine learning models to perform fault prediction on the corresponding category error data for multiple different categories, thereby helping to improve the overall prediction effect for the multiple error data.

[0016] In one possible implementation, the fault prediction device may also output multiple candidate classification rules and, in response to a user's selection of one of the candidate classification rules, determine a first classification rule to be used by the classifier to classify the multiple error data. In this way, the fault prediction device can support user selection of the classification rule to be used for classifying the error data, facilitating user intervention and improving the user experience.

[0017] In a second aspect, the present application provides a fault prediction device, which includes various modules for executing the fault prediction method in the first aspect or any possible implementation of the first aspect.

[0018] In a third aspect, the present application provides a BMC, comprising a power supply circuit and a processing circuit. The power supply circuit is used to supply power to the processing circuit, and the processing circuit is used to execute the fault prediction method in the first aspect or any implementation of the first aspect.

[0019] In a fourth aspect, the present application provides a computing device comprising a processor and a memory. The processor and the memory communicate with each other. The processor is configured to execute instructions stored in the memory so that the computing device executes the fault prediction method according to the first aspect or any one of the implementations of the first aspect. It should be noted that the memory may be integrated into the processor or may be independent of the processor. The computing device may further comprise a bus. The processor is connected to the memory via the bus. The memory may comprise a readable memory and a random access memory.

[0020] In a fifth aspect, the present application provides a computer-readable storage medium, which stores instructions. When the computer-readable storage medium is run on a computing device, the computing device executes the operating steps of the fault prediction method described in the first aspect or any implementation of the first aspect.

[0021] In a sixth aspect, the present application provides a computer program product comprising instructions, which, when executed on a computing device, enables the computing device to execute the operating steps of the fault prediction method described in the first aspect or any one of the implementations of the first aspect.

[0022] Based on the implementation methods provided in the above aspects, this application can also be further combined to provide more implementation methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] FIG1 is a schematic diagram of the structure of an exemplary data processing system provided by the present application;

[0024] FIG2 is a schematic diagram of the structure of another exemplary data processing system provided by the present application;

[0025] FIG3 is a flow chart of a fault prediction method provided by the present application;

[0026] Figure 4 is a schematic diagram of predicting different categories of data using machine learning models corresponding to different channels;

[0027] FIG5 is a schematic structural diagram of a fault prediction device provided by the present application;

[0028] FIG6 is a schematic diagram of the hardware structure of a computing device provided in this application. DETAILED DESCRIPTION

[0029] In order to improve the fault prediction effect and reduce the impact of UCE on the business operation of the device, the present application provides a fault prediction method, which classifies multiple error data corresponding to the storage medium and uses a machine learning model adapted to each category of error data (with high prediction accuracy or high prediction coverage) to predict whether UCE occurs in the storage medium. This can achieve both the accuracy and coverage of fault prediction, thereby effectively improving the overall effect of fault prediction for the multiple error data, and helping to reduce the impact of UCE occurring in the storage medium on the business operation of the device.

[0030] The technical solution in this application will be described below in conjunction with the drawings provided in this application.

[0031] Referring to FIG1 , a schematic diagram of the structure of a data processing system is shown. As shown in FIG1 , the data processing system 10 includes a software layer 100 and a hardware layer 200. The software layer 100 may include an operating system (OS) 101 and a basic input / output system (BIOS) 102. The hardware layer 200 may include a memory 201, a processor 202, and a baseboard management controller (BMC) 203. The memory 201, the processor 202, and the BMC 203 may be connected via a bus, such as a peripheral component interconnect express (PCIe) bus or other types of bus connections. Furthermore, the OS 101 and BIOS 102 in the software layer 100 may run on the processor 201.

[0032] OS 101 is used to control and manage resources in software layer 100 and hardware layer 200, such as instructing memory controller 2021 in hardware layer 200 to initiate isolation of a faulty area in memory 201. OS 101 can also run one or more services. FIG1 illustrates the running of services 1 and 2 as examples.

[0033] The BIOS 102 is used to provide low-level hardware configuration and control, such as detecting and reporting whether there is a faulty area in the memory 201 , and instructing the processor 202 to perform fault isolation operations on the memory 201 under the triggering of the BMC 203 .

[0034] Memory 201 is used to store the calculation data of processor 202, including data required to be read when processor 202 performs calculations, result data generated after processor 202 performs calculations, etc. Exemplarily, memory 201 may include one or more dual inline memory modules (DIMMs), wherein a DIMM can be used as a memory bar entity, and each memory bar can have two sides, each side of which is equipped with memory chips. Alternatively, memory 201 can also be double data rate synchronous dynamic random-access memory (DDR), low power double data rate SDRAM (LPDDR), phase change memory (PCM), resistive random access memory (RRAM), magneto resistive random access memory (MRAM), ferroelectric random access memory (FeRAM), nano RAM (NRAM), static random access memory (SRAM), or high bandwidth memory (HBM), etc., without limitation. Optionally, the memory 201 may also be a heterogeneous memory, that is, two or more different types of storage media are used as memory.

[0035] The processor 202 is used to perform corresponding processing operations, such as operations required to execute when executing services in the OS 101. Exemplarily, the processor 202 can be a central processing unit (CPU), or can be another type of processor such as an application-specific integrated circuit (ASIC). In addition, the processor 202 can be configured with a memory controller 2021 for performing corresponding data read and write operations on the memory 201. During the process of reading and writing data, the memory controller 2021 can detect whether there is a fault in the corresponding storage area in the memory 201, and can provide relevant information of the detected fault area to the BIOS 102.

[0036] The BMC 203 is configured to perform fault diagnosis on the memory 201 according to the error data reported by the BIOS 102 to determine whether a UCE occurs in the memory 201 .

[0037] Specifically, when an error occurs in a portion of the storage area in memory 201, memory controller 2021 provides error data corresponding to the storage area to BIOS 102, and BIOS 102 reports the error data to BMC 203. The error data can be used to indicate that a CE has occurred in memory 201. The error data can, for example, include information such as the storage area in memory 201 where the error occurred, the data input / output channel (DQ), and the burst. In this way, BMC 203 can obtain multiple error data reported by BIOS 102 over a period of time and classify the multiple error data to obtain error data of category 1 and error data of category 2. Typically, the error data in each category includes error data generated due to UCE and error data generated due to CE. In error data under category 1, the ratio between error data caused by UCE and error data caused by CE can be relatively small; in error data under category 2, the ratio between error data caused by UCE and error data caused by CE can be relatively large. In this way, BMC 203 can use machine learning model 1, which has a high precision for predicting UCE, to perform fault prediction on error data of category 1 and obtain corresponding prediction result 1, which indicates whether UCE has occurred in memory 201. Furthermore, BMC 203 can use machine learning model 2, which has a high recall for predicting UCE, to perform fault prediction on error data of category 2 and obtain corresponding prediction result 2, which indicates whether UCE has occurred in memory 201.

[0038] In this way, BMC 203 classifies the multiple error data corresponding to memory 201, and uses machine learning model 1 with higher prediction accuracy to predict faults for error data of category 1, and uses machine learning model 2 with higher coverage to predict faults for error data of category 2. This allows the overall fault prediction results for multiple error data to not only achieve a higher accuracy, but also a higher coverage, that is, taking into account both the accuracy and coverage of fault prediction, thereby effectively improving the overall prediction effect for the multiple error data, and helping to reduce the impact of UCE occurring in memory 201 on the operation of business 1 and business 2 in OS101.

[0039] It should be noted that the data processing system 10 shown in FIG1 is merely an example and is not intended to be limiting.

[0040] For example, as shown in Figure 2, the data processing system 20 may include multiple nodes, and Figure 2 is illustrated by taking nodes 301 to 304 as an example. Each node may be a computing node or a storage node, and each node may include a processor and a memory. In addition, the data processing system 20 also includes a server 305, and the server 305 can be connected to multiple nodes. When an error occurs in the memory of each node, the error data can be sent to the server 305, so that the server 305 can classify the multiple error data sent by the multiple nodes, and use the machine learning model adapted to the category to predict whether UCE occurs in the memory of the node for the error data under each category. Its specific implementation can be found in the above-mentioned description of the BMC 203 predicting whether UCE occurs in the memory 201.

[0041] For example, the data processing systems in Figures 1 and 2 above use the example of predicting whether UCE occurs in memory. However, in other data processing systems, it is also possible to predict whether UCE occurs in other types of storage media. For example, the main controller in a hard disk can use the above method to predict whether UCE occurs in the storage medium used for persistent data storage in the hard disk. This storage medium can be, for example, a solid state disk (SSD).

[0042] This application does not limit the specific architecture of the data processing system.

[0043] For ease of understanding, an embodiment of the fault prediction method provided in this application is described below in conjunction with the accompanying drawings.

[0044] Referring to FIG. 3 , FIG. 3 is a flow chart illustrating a fault prediction method provided by the present application. This method can be applied to the data processing system 10 shown in FIG. 1 , or to the data processing system 20 shown in FIG. 2 , or to other applicable data processing systems. For ease of explanation, this embodiment uses the data processing system 10 shown in FIG. 1 as an example. In this embodiment, BMC 203 utilizes multiple machine learning models to predict whether a UCE has occurred in memory 201.

[0045] The fault prediction method shown in FIG3 may specifically include:

[0046] S301 : The BIOS 102 sends a plurality of error data to the BMC 203 , where each error data is used to indicate that a CE occurs in the memory 201 .

[0047] Normally, after the memory 201 has been used for a period of time, some storage areas may fail. The failure in the storage area may be temporary (such as a failure in reading and writing data to the storage area at the software level) or permanent (such as physical damage to the storage area). When the memory controller 2021 reads and writes data to the storage area, if a data read and write error occurs, an error may be reported for the storage area. Specifically, error data may be sent to the BIOS 102. The error data may be, for example, CE data, including the storage area to be reported, error category (such as inspection, reading and writing, etc.), data input and output channels (DQ), bursts, the temperature and voltage of the storage area, and other data, which are not limited to this. Then, the BIOS 102 may send the error data to the BMC 203 so that the BMC 203 can determine whether a UCE occurs in the memory 201 based on the error data.

[0048] BMC 203 may continuously receive multiple error data sent by BIOS 102. For example, multiple storage areas in memory 201 may experience failures simultaneously, so BIOS 102 may send error data for these multiple storage areas to BMC 203. Alternatively, multiple storage areas in memory 201 may experience failures one after another, so BIOS 102 may send each error data to BMC 203 over a period of time (e.g., one hour). In this way, BMC 203 can determine that a CE has occurred in memory 201 based on the received error data, and can further determine whether a UCE has occurred in memory 201 based on the error data.

[0049] S302: The BMC 203 classifies the plurality of error report data to obtain error report data of the first category and error report data of the second category.

[0050] In one possible implementation, BMC 203 may be configured with a classifier, and BMC 203 may use the classifier to classify multiple error data into multiple categories. For ease of understanding, this embodiment uses an example of classifying multiple error data into a first category and a second category. In actual applications, the classifier may classify the multiple error data into more categories, and this is not a limitation.

[0051] In a specific implementation, the classifier may be configured with one or more classification rules, so that the classifier can use the one or more classification rules to classify the error data. For example, the classifier may be a decision tree, or other type of classifier. The classification rules refer to the rules used to classify the error data. The following describes some implementation examples of the classification rules.

[0052] Example 1: The classification rule may be a rule for classifying based on the number of failed DQs.

[0053] For example, when the number of failed DQs indicated in the error data is less than a threshold value of 1 (for example, 2), the BMC 203 may classify the error data into the first category according to the classification rule; and when the number of failed DQs indicated in the error data is greater than or equal to the threshold value of 1, the BMC 203 may classify the data into the second category according to the classification rule.

[0054] Example 2: The classification rule may be a rule for classifying based on the number of failed bursts.

[0055] For example, when the number of burst failures indicated in the error data is less than a threshold value of 2, the BMC 203 may classify the error data into the first category according to the classification rule; and when the number of burst failures indicated in the error data is greater than or equal to the threshold value of 2, the BMC 203 may classify the data into the second category according to the classification rule.

[0056] Example 3: The classification rule may be specifically a rule for classification based on the number of failed rows, where each row may include multiple storage units.

[0057] For example, when the number of failed rows indicated in the error data is less than a threshold value of 3, the BMC 203 may classify the error data into the first category according to the classification rule; and when the number of failed rows indicated in the error data is greater than or equal to the threshold value of 3, the BMC 203 may classify the data into the second category according to the classification rule.

[0058] The thresholds 1 to 3 can be manually configured by a technician. For example, the specific values ​​of thresholds 1 to 3 can be determined by analyzing the correlation between multiple error data generated in a past time period and the UCE. Alternatively, the BMC 203 can initialize thresholds 1 to 3 and dynamically update the specific values ​​of thresholds 1 to 3 through feedback adjustment or other means.

[0059] Example 4: The classification rule may specifically be a rule for classification based on whether the DQ is missing.

[0060] For example, when the error data includes information such as the DQ identifier, the BMC 203 may classify the error data into the first category according to the classification rule; and when the error data does not include (i.e., is missing) information such as the DQ identifier, the BMC 203 may classify the data into the second category according to the classification rule.

[0061] Example 5: The classification rule may be a rule for classification based on whether the burst is missing.

[0062] For example, when the error data includes information such as a burst identifier, the BMC 203 may classify the error data into the first category according to the classification rule; and when the error data does not include (i.e., is missing) information such as a burst identifier, the BMC 203 may classify the data into the second category according to the classification rule.

[0063] Example 6: The classification rule may be a rule for classification based on whether the value in the parity register is valid.

[0064] For example, when the error data includes the value in the parity register and the value is valid (DQ and burst information can be parsed from the value), the BMC 203 may classify the error data into the first category according to the classification rule; and when the error data does not include the value in the parity register or the value is invalid, the BMC 203 may classify the data into the second category according to the classification rule.

[0065] In addition to the above implementation examples, the classification rules configured in the BMC 203 may also be other types of rules, such as user-defined classification rules, etc., which is not limited to this.

[0066] In actual application, the BMC 203 may support the user to configure the classification rules in the classifier.

[0067] For example, the BMC 203 can output multiple candidate classification rules to the processor 202, such as the six classification rules in the above example can be output to the processor 202 as candidate classification rules, so that the processor 202 can use the human-computer interaction device (such as a display screen, etc.) in the data processing system 10 to present the multiple candidate classification rules to the user. Accordingly, the user can select the multiple candidate classification rules presented on the human-computer interaction device, or the user can customize the classification rules on the human-computer interaction device. Then, the human-computer interaction device can notify the BMC 203 of the candidate classification rule selected by the user or the user-defined classification rule through the processor 202. For ease of distinction and description, the candidate classification rule selected by the user or the customized classification rule will be referred to as the first classification rule below. In this way, the BMC 203 can build a classifier based on the first classification rule, and use the subsequent classifier to classify multiple error data.

[0068] Among them, there is a difference in the ratio of positive and negative samples between the error data of the first category classified based on the first classification rule and the error data of the second category. Among them, the positive sample refers to the error data generated when UCE occurs in the memory 201, and the negative sample is the error data generated when CE occurs in the memory 201 (UCE does not occur). Specifically, in the error data of the first category, the ratio between the number of positive samples and the number of negative samples is small. For example, the ratio between the number of positive samples and the number of negative samples is 1:99.1, indicating that the possibility of UCE occurring in the memory 201 indicated by the error data classified into the first category is small, that is, among nearly 100 error data, there is only one error data generated by UCE occurring in the memory 201. In the second category of error data, the ratio between the number of positive samples and the number of negative samples is relatively large. For example, the ratio is 1:5.3. This indicates that the error data classified as the second category is more likely to indicate that UCE has occurred in memory 201. That is, one out of every six error data is caused by UCE in memory 201. In other words, the positive-to-negative sample ratio corresponding to the first category of error data is smaller than the positive-to-negative sample ratio corresponding to the second category of error data.

[0069] For example, assuming that the error data includes the number of failed DQs, when the number of failed DQs is less than 2, the possibility of UCE occurring in the memory 201 is low. That is, when the number of failed DQs is less than 2, the memory 201 will most likely not experience UCE. In this case, the BMC 203 may classify the error data with the number of failed DQs less than or equal to 1 into category 1. When the number of failed DQs is greater than or equal to 2, the possibility of UCE occurring in the memory 201 is high. That is, when the number of failed DQs is greater than or equal to 2, the memory 201 will most likely experience UCE. In this case, the BMC 203 may classify the error data with the number of failed DQs greater than or equal to 2 into category 2.

[0070] S303: Use the machine learning model 1 to perform fault prediction on the first category of error data to obtain a first prediction result, wherein the first prediction result is used to indicate whether UCE occurs in the memory 201. The accuracy of the machine learning model 1 in predicting whether UCE occurs in the memory 201 is higher than the first threshold.

[0071] S304: Use machine learning model 2 to perform fault prediction on the second category of error data to obtain a second prediction result, wherein the second prediction result is used to indicate whether UCE occurs in memory 201, and the coverage rate of the machine learning model 2 predicting whether UCE occurs in memory 201 is higher than the second threshold.

[0072] The prediction accuracy refers to the proportion of actual UCEs in the results predicted by the machine learning model, which can be calculated using the following formula (1):

[0073] Precision refers to the accuracy rate; TP refers to the number of actual UCEs; FP refers to the number of UCEs that did not occur but were judged to have occurred; TP+FP refers to the total number of predicted UCEs.

[0074] The predicted coverage rate refers to the proportion of UCEs that are correctly predicted by the machine learning model among the actual UCEs, which can be calculated using the following formula (2):

[0075] Recall refers to coverage; TP refers to the number of actual UCEs; FN refers to the number of UCEs that occurred but were not predicted by the machine learning model; TP+FN refers to the total number of actual UCEs.

[0076] In actual application, as shown in FIG4 , after the BMC 203 classifies the plurality of error data using a classifier, data of different categories can be input into a machine learning model corresponding to the channel through different channels.

[0077] In this embodiment, BMC 203 can input the error data of the first category into the machine learning model 1 with a higher prediction accuracy. In this way, for error data that is more likely to be generated by CE (no UCE) occurring in the memory 201, the machine learning model 1 with a higher prediction accuracy (the accuracy is greater than the first threshold) can be used to perform fault prediction to improve the accuracy of the fault prediction.

[0078] At the same time, BMC 203 can input the second category of error data into the machine learning model 2 with a higher prediction coverage. In this way, for error data that is more likely to be generated by UCE in memory 201, the machine learning model 2 with a higher coverage (coverage greater than the second threshold) can be used to perform fault prediction to improve the prediction coverage, that is, to identify as many UCEs as possible.

[0079] Each machine learning model can perform fault prediction based on an error data to determine whether the memory 201 generates the error data because of CE or UCE; accordingly, the first prediction result output by the machine learning model 1 (similar to the machine learning model 2) includes the prediction results corresponding to each error data in the first category. Alternatively, each machine learning model can perform fault prediction based on multiple error data at the same time, that is, BMC203 can input two or more (including two) error data into the same machine learning model so that the machine learning model performs fault prediction based on the two or more error data; accordingly, the first prediction result output by the machine learning model 1 (similar to the machine learning model 2) includes the prediction results corresponding to multiple error data in the first category.

[0080] The machine learning model used to predict faults based on different types of error data can use different network structures or different hyperparameters. This is exemplified below.

[0081] Example 1: Machine learning model 1 and machine learning model 2 can use different model structures but have the same hyperparameters.

[0082] For example, machine learning model 1 may be a neural network model, and machine learning model 2 may be a model based on a random forest algorithm; or, machine learning model 1 may be a convolutional neural network (CNN) model, and machine learning model 2 may be a recurrent neural network (RNN). However, machine learning model 1 and machine learning model 2 may have the same hyperparameters, such as using the same threshold to determine whether UCE occurs in memory 201.

[0083] Example 2: Machine learning model 1 and machine learning model 2 can have the same model structure but use different hyperparameters.

[0084] For example, machine learning model 1 and machine learning model 2 are both CNN models. However, the threshold for determining whether UCE occurs in memory 201 in machine learning model 1 is set to 0.8 (the accuracy is improved by setting a larger threshold), and the threshold for determining whether UCE occurs in memory 201 in machine learning model 2 is set to 0.4 (the coverage is improved by setting a smaller threshold).

[0085] Example 3: Machine learning model 1 and machine learning model 2 can adopt different model structures and different hyperparameters.

[0086] For example, machine learning model 1 can be a CNN model, and machine learning model 2 can be an RNN model; and the threshold for determining whether UCE occurs in memory 201 in machine learning model 1 is set to 0.8, and the threshold for determining whether UCE occurs in memory 201 in machine learning model 2 is set to 0.4.

[0087] In this way, for the multiple error reports acquired by BMC 203, machine learning model 1 is used to predict faults for error reports in category 1, and machine learning model 2 is used to predict faults for error reports in category 2. This not only achieves a high accuracy rate, but also a high coverage rate, thereby improving the fault prediction effect for memory 201. At the same time, the false positive rate (false positive rate) of BMC 203's fault prediction for memory 201 can be maintained at a low level. The false positive rate refers to the proportion of CE scenarios that are mistakenly predicted as UCE.

[0088] Moreover, when some dimensional information is missing in the error data, such as missing information on the number of DQs or the number of bursts, BMC 203 can use a machine learning model adapted thereto to predict whether the memory 201 has failed. When some dimensional data is missing in the error data, BMC 203 can use other machine learning models for fault prediction. In this way, it is possible to avoid using a single machine learning model to predict faults for a variety of error data (with and without missing dimensional information) resulting in low accuracy or coverage of fault predictions.

[0089] In actual application, in the following three test scenarios, the overall accuracy and coverage of fault prediction for multiple error data reached a high level.

[0090] In test scenario 1, BMC 203 can predict faults using 89,442 error data. These 89,442 error data include 4,282 error data due to UCE in memory 201 (positive samples) and 85,160 error data due to CE in memory 201 (negative samples). After BMC 203 classifies the 89,442 error data using the rule based on the number of failed DQs, the number of error data in category 1 is 66,453 (the number of failed DQs is less than 2), including 664 positive samples and 65,789 negative samples; the number of error data in category 2 is 22,989 (the number of failed DQs is greater than or equal to 2), including 3,618 positive samples and 19,371 negative samples. Then, machine learning model 1, with a 70% accuracy (10% coverage), was used to predict faults for category 1 error data, and machine learning model 2, with a 70% coverage (50% accuracy), was used to predict faults for category 2 error data. This resulted in a combined accuracy of 69.3% and a coverage of 60.7% for all 89,442 error data points. At the same time, the false positive rate was reduced to 1.4%.

[0091] In test scenario two, BMC 203 can predict faults using 87,301 error data items. These 87,301 error data items include 2,141 error data items (positive samples) due to UCE in memory 201 and 85,160 error data items (negative samples) due to CE in memory 201. After BMC 203 classifies the 89,442 error data items using the rule based on the number of failed bursts, the number of error data items in category 1 is 66,453 (the number of failed bursts is less than 2), including 264 positive samples and 65,789 negative samples. The number of error data items in category 2 is 21,248 (the number of failed bursts is greater than or equal to 2), including 1,877 positive samples and 19,371 negative samples. Then, machine learning model 1, with a 70% accuracy (5% coverage), was used to predict faults for category 1 error data, and machine learning model 2, with a 70% coverage (55% accuracy), was used to predict faults for category 2 error data. This resulted in a combined accuracy of 65% and coverage of 62% for all 89,442 error data points. At the same time, the false positive rate was reduced to 1.27%.

[0092] In test scenario three, BMC 203 can predict faults using 87,301 error data items. These 87,301 error data items include 2,141 error data items (positive samples) due to uncorrupted access (UCE) of memory 201 and 85,160 error data items (negative samples) due to closed access (CE) of memory 201. BMC 203 classifies the 89,442 error data items using a rule based on the validity of the value in the parity register. The results show that category 1 contains 74,115 error data items (valid parity register values), including 1,548 positive samples and 72,567 negative samples. Category 2 contains 13,186 error data items (invalid parity register values), including 593 positive samples and 12,593 negative samples. Then, machine learning model 1, with a 70% accuracy (35% coverage), was used to predict faults for category 1 error data, while machine learning model 2, with a 60% coverage (30% accuracy), was used to predict faults for category 2 error data. This resulted in a combined accuracy of 52% and a coverage of 60% for all 89,442 error data points. Furthermore, the false positive rate was reduced to 1.42%.

[0093] It should be noted that the embodiment shown in FIG2 uses an example in which BMC 203 uses one classification rule (i.e., the first classification rule) to classify multiple error reports into two categories, and then uses two machine learning models to perform fault prediction. In other embodiments, BMC 203 may use multiple classification rules to classify multiple error reports into three or more (including three) categories, and use three or more machine learning models to perform fault prediction.

[0094] Specifically, taking the example of a classifier including a first classification rule and a second classification rule, wherein the first classification rule may be, for example, a rule for classification based on the number of failed DQs, and the second classification rule may be, for example, a rule for classification based on the number of failure bursts. Then, the BMC 203 may use these two classification rules in the classifier to divide multiple error data into four categories. Specifically, when the number of failed DQs included in the error data is less than a threshold value 1, and the number of failure bursts is less than a threshold value 2, the BMC 203 may classify the error data into category 1 and perform a fault prediction on it using the machine learning model 1. When the number of failed DQs included in the error data is greater than or equal to the threshold value 1, and the number of failure bursts is less than the threshold value 2, the BMC 203 may classify the error data into category 2 and perform a fault prediction on it using the machine learning model 2. When the number of failed DQs included in the error data is less than the threshold value 1, and the number of failure bursts is greater than or equal to the threshold value 2, the BMC 203 may classify the error data into category 3 and perform a fault prediction on it using the machine learning model 3. When the error data includes a number of failed DQs greater than or equal to threshold 1 and a number of failed bursts greater than or equal to threshold 2, the BMC 203 may classify the error data into category 4 and use the machine learning model 4 to perform fault prediction on it.

[0095] The accuracy of machine learning model 1 may be higher than that of the other machine learning models, and the coverage of the other machine learning models may be higher than that of the other machine learning models. The conditions for determining whether UCE has occurred in memory 201 based on the error data may differ in machine learning models 2 through 4. For example, the thresholds for determining whether UCE has occurred in memory 201 may differ in machine learning models 2 through 4.

[0096] In actual test scenarios, the overall accuracy and coverage of fault prediction for multiple error data based on the first and second classification rules reached a high level.

[0097] Specifically, BMC 203 can perform fault prediction using 89,442 pieces of error data. These 89,442 pieces of error data include 4,282 pieces of error data due to UCE in memory 201 and 85,160 pieces of error data due to CE in memory 201. BMC 203 classifies the 89,442 pieces of error data using the first and second classification rules, resulting in four categories. Among them, the number of error data under category 1 is 12756 (the number of failed DQs is less than 2, and the number of failed bursts is less than 2), including 163 positive samples and 12593 negative samples; the number of error data under category 2 is 25839 (the number of failed DQs is less than 2, and the failed burst is greater than or equal to 2), including 1852 positive samples and 23987 negative samples; the number of error data under category 3 is 5295 (the number of failed DQs is greater than or equal to 2, and the failed burst is less than 2), including 597 positive samples and 4698 negative samples; the number of error data under category 4 is 45552 (the number of failed DQs is greater than or equal to 2, and the failed burst is greater than or equal to 2), including 1670 positive samples and 43882 negative samples. Then, BMC 203 can use machine learning model 1 with an accuracy of 60% (coverage of 10%) to predict faults for error data of category 1, use machine learning model 2 with a coverage of 65% (accuracy of 50%) to predict faults for error data of category 2, use machine learning model 3 with a coverage of 70% (accuracy of 50%) to predict faults for error data of category 3, and use machine learning model 4 with a coverage of 60% (accuracy of 55%) to predict faults for error data of category 4. In this way, the total accuracy of fault prediction for 89,442 pieces of error data can be as high as 61.4%, and the coverage rate can be as high as 61.7%. At the same time, the false positive rate can be reduced to 1.95%. Among them, the thresholds (i.e., judgment conditions) used to determine whether UCE occurs in memory 201 in machine learning models 1 to 4 can be: 0.9, 0.4, 0.5, and 0.7 respectively.

[0098] Furthermore, when the UCE is predicted to occur in the memory 201, the following non-limiting implementation methods may be used for processing.

[0099] Example 1: After predicting that a UCE occurs in the memory 201, the BMC 203 can perform corresponding fault isolation operations on the storage area in the memory 201 where the fault occurs, as shown in FIG3 . For example, when there is a faulty storage area in the memory 201, the BMC 203 can remap the access of the business in the OS 101 to the faulty storage area to the access to the storage area in the memory 201 or the reserved storage resources in the memory controller 2021, such as remapping the virtual address corresponding to the physical address of the first storage area to the physical address corresponding to the corresponding storage area in the reserved storage resources. Alternatively, the BMC 203 can offline the page corresponding to the faulty storage area in the memory 201 through page table management, so that the multiple pages subsequently accessed by the business in the OS 101 do not include the page that has been offlined, that is, memory page isolation (page offline).

[0100] Example 2: After predicting that a UCE occurs in the memory 201, the BMC 203 can issue an alarm for the UCE, such as generating an alarm message for the UCE and sending the alarm message to the OS 101, so that the OS 101 can perform corresponding operations based on the alarm message. For example, the OS 101 can present the alarm message to the user through a human-computer interaction device so that the user can promptly perceive that the memory 201 has failed and perform maintenance on the memory 201. Alternatively, the BMC 203 can perform fault preprocessing on the memory 201 in advance based on the alarm message, such as migrating the data in the storage area of ​​the memory 201 that generates the UCE to other normal storage areas, or using a data recovery mechanism to recover the data in the storage area, etc. This embodiment does not limit this.

[0101] In Example 3, when the machine learning model outputs the prediction result of the occurrence of UCE, it can also output the confidence level, which is used to measure the credibility of the machine learning model's prediction of the occurrence of UCE in memory 201. In this way, when the confidence level is high, such as when the confidence level is greater than 85%, BMC 203 can directly perform a fault isolation operation on the storage area in memory 201 that generates the UCE, such as the fault isolation operation described in Example 1 above. When the confidence level is low, such as when the confidence level is less than 85%, BMC 203 can issue an alarm for the UCE and send the alarm information to OS 101, or BMC 203 can perform fault preprocessing on memory 201 in advance based on the alarm information.

[0102] In actual application, when it is predicted that the memory 201 has UCE, the BMC 203 may also be implemented in other ways, which is not limited to this.

[0103] In addition, when the BMC 203 uses the machine learning model to predict that a CE occurs in the memory 201 (ie, no UCE occurs), no additional operations need to be performed.

[0104] In a further possible implementation, multiple machine learning models or classifiers in the BMC 203 may also be dynamically updated.

[0105] As an implementation example of updating a machine learning model, after BMC 203 uses machine learning model 1 and machine learning model 2 to perform fault prediction on multiple error data and obtains the fault prediction result, it can obtain the actual result corresponding to the error data. For example, after BMC 203 completes fault prediction based on an error data, if it receives UCE reported by BIOS 102, it can be determined that the actual result corresponding to the error data is that UCE occurs in memory 201; and if it does not receive UCE reported by BIOS 102, it can be determined that the actual result corresponding to the error data is that CE occurs in memory 201 (i.e., no UCE occurs). Then, BMC 203 can update machine learning model 1 based on the actual result corresponding to the first category of error data and the first prediction result output by machine learning model 1, including updating one or more of the hyperparameters, parameters, and model structure in the machine learning model 1. Similarly, BMC 203 can update machine learning model 2 based on the actual result corresponding to the second category of error data and the second prediction result output by machine learning model 2.

[0106] As an implementation example of updating a classifier, after BMC 203 uses machine learning model 1 and machine learning model 2 to perform fault prediction on multiple error data and obtains corresponding prediction results, it can obtain the actual results corresponding to the error data, which are used to indicate whether a UCE has occurred in memory 201 or not. Then, BMC 203 can update the classifier based on the actual results corresponding to the first category of error data, the actual results corresponding to the second category of error data, the first prediction result output by machine learning model 1, and the second prediction result output by machine learning model 2, including adjusting one or more of the classification rules in the classifier, the thresholds in the classification rules, and the classification algorithm used by the classifier.

[0107] It should be noted that in the embodiment shown in FIG3 , the example of BMC 203 using different machine learning models to predict whether UCE has occurred in memory 201 is used for illustrative purposes. In the data processing system 20 shown in FIG2 , the server 305 may also use multiple different machine learning models to predict whether UCE has occurred in the memory of each node; or, in a persistent storage scenario, the main controller may use multiple different machine learning models to predict whether UCE has occurred in the persistent storage medium, etc. The specific implementation process can be found in the relevant description of the embodiment shown in FIG3 above, and will not be repeated here.

[0108] It is worth noting that other reasonable step combinations that can be thought of by those skilled in the art based on the above description also fall within the scope of protection of this application. Secondly, those skilled in the art should also be familiar with that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by this application.

[0109] The fault prediction method provided in the embodiment of the present application is introduced above with reference to FIG. 1 to FIG. 4 . Next, the structure of the fault prediction device and computing equipment provided in the embodiment of the present application is introduced with reference to the accompanying drawings.

[0110] 5 , which shows a schematic diagram of the structure of a processor, the fault prediction device 500 includes:

[0111] An acquisition module 501 is configured to acquire a plurality of error reporting data, each of which is configured to indicate that a correctable error CE occurs on a storage medium.

[0112] A classification module 502 is used to classify the plurality of error reporting data to obtain error reporting data of a first category and error reporting data of a second category;

[0113] The fault prediction module 503 is used to use a first machine learning model to perform fault prediction on a first category of error data to obtain a first prediction result. The first prediction result is used to indicate whether an uncorrectable error (UCE) occurs in the storage medium, and the accuracy of the UCE predicted by the first machine learning model is higher than a first threshold; and to use a second machine learning model to perform fault prediction on a second category of error data to obtain a second prediction result. The second prediction result is used to indicate whether UCE occurs in the storage medium, and the coverage of UCE predicted by the second machine learning model is higher than a second threshold.

[0114] In one possible implementation, the positive-to-negative sample ratio corresponding to the first category of error data is smaller than the positive-to-negative sample ratio corresponding to the second category of error data. The positive-to-negative sample ratio is the ratio between the number of positive samples and the number of negative samples. The positive sample is the error data generated when UCE occurs in the storage medium, and the negative sample is the error data generated when CE occurs in the storage medium.

[0115] In one possible implementation, the first machine learning model and the second machine learning model have different model structures or different hyperparameters.

[0116] In a possible implementation, the fault prediction device 500 further includes an updating module 504, which is configured to:

[0117] Obtaining a true result corresponding to the first category of error data, where the true result is used to indicate that a UCE or CE occurs on the storage medium, and updating the first machine learning model based on the true result corresponding to the first category of error data and the first prediction result;

[0118] Alternatively, the true result corresponding to the second category of error data is obtained, and the true result is used to indicate that UCE or CE occurs in the storage medium, and the second machine learning model is updated based on the true result corresponding to the second category of error data and the second prediction result.

[0119] In a possible implementation, multiple error data are classified by a classifier, and the fault prediction device 500 further includes an updating module 504, which is configured to:

[0120] Obtaining a true result corresponding to the first category of error data and a true result corresponding to the second category of error data, where the true result is used to indicate that a UCE or CE occurs on the storage medium;

[0121] The classifier is updated according to the true results corresponding to the first category of error data, the true results corresponding to the second category of error data, the first prediction result, and the second prediction result.

[0122] In a possible implementation, the plurality of error reporting data are classified by a classifier, wherein the classifier includes a first classification rule and a second classification rule;

[0123] The classification module 502 is configured to classify the plurality of error reporting data using the first classification rule in the classifier to obtain error reporting data of the first category, error reporting data of the second category, error reporting data of the third category, and error reporting data of the fourth category;

[0124] The fault prediction module 503 is further configured to:

[0125] Using a third machine learning model, performing fault prediction on the third category of error reporting data to obtain a third prediction result, where the third prediction result is used to indicate whether an uncorrectable error (UCE) occurs on the storage medium;

[0126] A fourth machine learning model is used to perform fault prediction on the fourth category of error data to obtain a fourth prediction result, where the fourth prediction result is used to indicate whether the UCE occurs in the storage medium.

[0127] In a possible implementation, the fault prediction device 500 further includes:

[0128] Output module 505, used to output multiple candidate classification rules;

[0129] The determination module 506 is configured to determine the first classification rule used by the classifier to classify the plurality of error reporting data in response to a user's selection operation on the plurality of candidate classification rules.

[0130] Since the task processing device 500 shown in FIG5 corresponds to the BMC 203 in the embodiment shown in FIG3 above, the specific implementation of the task processing device 500 shown in FIG5 and the technical effects thereof can be found in the relevant description of the embodiment shown in FIG3 above, and will not be described in detail here.

[0131] FIG6 is a schematic diagram of the hardware structure of a computing device 600 provided in the present application. The computing device 600 can, for example, implement the BMC 203 in the embodiment shown in FIG3 .

[0132] As shown in Figure 6, the computing device 600 includes a processor 601, a memory 602, and a communication interface 603. The processor 601, the memory 602, and the communication interface 603 communicate via a bus 604, and may also communicate via other means such as wireless transmission. The memory 602 is used to store instructions, and the processor 601 is used to execute the instructions stored in the memory 602. Furthermore, the computing device 600 may also include a memory unit 605, and the memory unit 605 may be connected to the processor 601, the storage medium 602, and the communication interface 603 via a bus 604. The memory 602 stores program code, and the processor 601 may call the program code stored in the memory 602 to perform the following operations:

[0133] Acquire a plurality of error reporting data, each error reporting data in the plurality of error reporting data is used to indicate that a correctable error CE occurs in the storage medium;

[0134] Classifying the plurality of error reporting data to obtain error reporting data of a first category and error reporting data of a second category;

[0135] Using a first machine learning model, perform fault prediction on the first category of error data to obtain a first prediction result, where the first prediction result is used to indicate whether an uncorrectable error (UCE) occurs on the storage medium, and an accuracy rate of the first machine learning model in predicting the UCE is higher than a first threshold;

[0136] Using a second machine learning model, fault prediction is performed on the second category of error data to obtain a second prediction result. The second prediction result is used to indicate whether the UCE occurs in the storage medium. The second machine learning model predicts that the coverage of the UCE is higher than a second threshold.

[0137] It should be understood that in this embodiment, the processor 601 may be a CPU, or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete device components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0138] The memory 602 may include a read-only memory and a random access memory, and provides instructions and data to the processor 601. The memory 602 may also include a nonvolatile random access memory.

[0139] The memory 602 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM), and direct rambus RAM (DR RAM).

[0140] The communication interface 603 is used to communicate with other devices connected to the computing device 600. In addition to the data bus, the bus 604 may also include a power bus, a control bus, and a status signal bus. However, for the sake of clarity, various buses are labeled as bus 604 in the figure.

[0141] It should be understood that the computing device 600 according to the embodiment of the present application may correspond to the task processing device 500 in the embodiment of the present application, and may correspond to executing the method executed by the BMC 203 in the method shown in Figure 3 according to the embodiment of the present application. The above-mentioned and other operations and / or functions implemented by the computing device 600 are respectively for implementing the process of the corresponding method in Figure 3. For the sake of brevity, they are not further described here.

[0142] The present application provides a BMC, which includes a power supply circuit and a processing circuit. The power supply circuit is used to supply power to the processing circuit, and the processing circuit is used to execute the method executed by the BMC 203 in the method shown in FIG3 .

[0143] Embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium capable of being stored by a computing device, or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the above-described fault prediction method.

[0144] The present application also provides a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computing device, the computer program product fully or partially generates the process or function described in the present application.

[0145] The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, or data center to another website, computer, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means.

[0146] The computer program product may be a software installation package. When any of the aforementioned fault prediction methods needs to be used, the computer program product may be downloaded and executed on a computing device.

[0147] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0148] The terms used in the above embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and claims of this application, the singular expressions "one", "a kind of", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; the character " / " generally indicates that the objects associated with each other are in an "or" relationship. In the embodiments of the present application. "Simultaneously" means within the same time period, including situations at the same time. The terms "first", "second", etc. in the specification, claims and drawings of this application are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchangeable where appropriate, and this is merely a way of distinguishing objects with the same properties when describing them in the embodiments of the present application.

[0149] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0150] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A fault prediction method, characterized in that, The method includes: Obtaining a plurality of error reporting data, where each error reporting data in the plurality of error reporting data is used to indicate that a correctable error CE has occurred in a storage medium; Classifying the plurality of error reporting data to obtain error reporting data of a first category and error reporting data of a second category; Using a first machine learning model to perform a fault prediction on the error reporting data of the first category to obtain a first prediction result, where the first prediction result is used to indicate whether an uncorrectable error UCE has occurred in the storage medium, and the precision rate of the first machine learning model for predicting the UCE is higher than a first threshold; Using a second machine learning model to perform a fault prediction on the error reporting data of the second category to obtain a second prediction result, where the second prediction result is used to indicate whether the UCE has occurred in the storage medium, and the coverage rate of the second machine learning model for predicting the UCE is higher than a second threshold.

2. The method according to claim 1, wherein The ratio of positive to negative samples corresponding to the error reporting data of the first category is less than the ratio of positive to negative samples corresponding to the error reporting data of the second category. The ratio of positive to negative samples is the ratio between the number of positive samples and the number of negative samples. The positive sample is the error reporting data generated when the UCE occurs in the storage medium, and the negative sample is the error reporting data generated when the CE occurs in the storage medium.

3. The method according to claim 1 or 2, characterized in that, The first machine learning model and the second machine learning model have different model structures or different hyperparameters.

4. The method according to claim 3, characterized in that, The method further includes: Obtaining the true result corresponding to the error reporting data of the first category, where the true result is used to indicate that the UCE or the CE has occurred in the storage medium, and updating the first machine learning model according to the true result corresponding to the error reporting data of the first category and the first prediction result; Or, obtaining the true result corresponding to the error reporting data of the second category, where the true result is used to indicate that the UCE or the CE has occurred in the storage medium, and updating the second machine learning model according to the true result corresponding to the error reporting data of the second category and the second prediction result.

5. The method according to any one of claims 1 to 4, characterized in that, The plurality of error reporting data is classified by a classifier, and the method further includes: Obtaining the true result corresponding to the error reporting data of the first category and the true result corresponding to the error reporting data of the second category, where the true result is used to indicate that the UCE or the CE has occurred in the storage medium; Updating the classifier according to the true result corresponding to the error reporting data of the first category, the true result corresponding to the error reporting data of the second category, the first prediction result, and the second prediction result.

6. The method according to any one of claims 1 to 5, characterized in that, The plurality of error reporting data is classified by a classifier, and the classifier includes a first classification rule; The classifying of the plurality of error reporting data includes: Using the first classification rule in the classifier to classify the plurality of error reporting data.

7. The method according to any one of claims 1 to 5, characterized in that, The plurality of error reporting data is classified by a classifier, and the classifier includes a first classification rule and a second classification rule; The classifying of the plurality of error reporting data includes: Classify the multiple error reporting data by using the first classification rule in the classifier to obtain the error reporting data of the first category, the error reporting data of the second category, the error reporting data of the third category, and the error reporting data of the fourth category; The method further includes: Use a third machine learning model to perform a fault prediction on the error reporting data of the third category to obtain a third prediction result, where the third prediction result is used to indicate whether an uncorrectable error (UCE) has occurred in the storage medium; Use a fourth machine learning model to perform a fault prediction on the error reporting data of the fourth category to obtain a fourth prediction result, where the fourth prediction result is used to indicate whether the UCE has occurred in the storage medium.

8. The method according to claim 6 or 7, characterized in that, The method further includes: Output a plurality of candidate classification rules; In response to a user's selection operation on the plurality of candidate classification rules, determine the first classification rule used by the classifier when classifying the multiple error reporting data.

9. A fault prediction device, characterized in that, The apparatus includes: An acquisition module, configured to acquire multiple error reporting data, where each error reporting data in the multiple error reporting data is used to indicate that a correctable error (CE) has occurred in the storage medium; A classification module, configured to classify the multiple error reporting data to obtain error reporting data of a first category and error reporting data of a second category; A fault prediction module, configured to use a first machine learning model to perform a fault prediction on the error reporting data of the first category to obtain a first prediction result, where the first prediction result is used to indicate whether an uncorrectable error (UCE) has occurred in the storage medium, and the precision rate of the first machine learning model for predicting the UCE is higher than a first threshold; use a second machine learning model to perform a fault prediction on the error reporting data of the second category to obtain a second prediction result, where the second prediction result is used to indicate whether the UCE has occurred in the storage medium, and the coverage rate of the second machine learning model for predicting the UCE is higher than a second threshold.

10. The device according to claim 9, characterized in that, The ratio of positive samples to negative samples corresponding to the error reporting data of the first category is less than the ratio of positive samples to negative samples corresponding to the error reporting data of the second category. The ratio of positive samples to negative samples is the ratio of the number of positive samples to the number of negative samples. The positive sample is the error reporting data generated when the UCE occurs in the storage medium, and the negative sample is the error reporting data generated when the CE occurs in the storage medium.

11. The device according to claim 9 or 10, characterized in that, The first machine learning model and the second machine learning model have different model structures or different hyperparameters.

12. The device according to claim 11, wherein, The apparatus further includes an update module, and the update module is configured to: Obtain the true result corresponding to the error reporting data of the first category, where the true result is used to indicate that the UCE or the CE has occurred in the storage medium, and update the first machine learning model according to the true result corresponding to the error reporting data of the first category and the first prediction result; Alternatively, obtain the true result corresponding to the error reporting data of the second category, where the true result is used to indicate that the UCE or the CE has occurred in the storage medium, and update the second machine learning model according to the true result corresponding to the error reporting data of the second category and the second prediction result.

13. The device according to any one of claims 9 to 12, characterized in that, The multiple error reporting data are classified by a classifier, and the apparatus further includes an updating module, where the updating module is configured to: Obtain the true results corresponding to the error reporting data of the first category and the true results corresponding to the error reporting data of the second category, where the true results are used to indicate that the storage medium has a UCE or a CE; Update the classifier according to the true results corresponding to the error reporting data of the first category, the true results corresponding to the error reporting data of the second category, the first prediction result, and the second prediction result.

14. A computing device, characterized in that, Comprising a processor and a memory; The processor is configured to execute the instructions stored in the memory to cause the computing device to perform the steps of the method according to any one of claims 1 to 8.

15. A computer-readable storage medium, characterized in that, Comprising instructions that, when running on a computing device, cause the computing device to perform the steps of the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Fault prediction method and device and related equipment

    CN120234190A

  • Service risk classifier training method, device and equipment and storage medium

    CN111882426A

  • Memory fault prediction model generation method, memory fault prediction model detection method, memory fault prediction model generation device and memory fault prediction model detection equipment

    CN114443398A

  • Disk fault prediction method and device, electronic equipment and readable storage medium

    CN115858265A

  • Memory error prediction method, device and equipment

    CN116954983A