A Disk Fault Prediction Method for Intelligent Operation and Maintenance of Large-Scale Cloud Data Centers

Through the optimization of time-progressive sampling TPS and Transformer models, the unbalanced data and long-time series problems of disk failure prediction in cloud data centers are solved, and disk failure prediction with high accuracy and low false alarm rate is achieved.

CN115373879BActive Publication Date: 2025-07-11NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211039310.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-29
Publication Date
2025-07-11
Estimated Expiration
2042-08-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict disk failures in large-scale cloud data centers, especially due to high prediction false positive rates and inaccurate fault location caused by imbalanced data and long-time series data.

Method used

The time-gradual sampling TPS method is used to enhance data, combine the Transformer model and optimize attention calculation, enhance local context information through convolutional projection, and build a disk failure prediction model.

Benefits of technology

It improves the accuracy and reliability of disk failure prediction, reduces the false alarm rate, and can achieve excellent results under the highly required F1 value and Matthews correlation coefficient, and is suitable for intelligent operation and maintenance of large-scale cloud data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115373879B_ABST
    Figure CN115373879B_ABST
Patent Text Reader

Abstract

The present invention discloses a disk fault prediction method for intelligent operation and maintenance of large-scale cloud data centers, including: first, performing information entropy feature processing on unbalanced data to select relatively important features; then dividing the processed unbalanced data to extract sample data of the minority class, that is, fault samples; then using the Time Progressive Sampling method (TPS) to perform data augmentation on the fault samples to generate synthetic data, generating more fault sample data through TPS, so that the ratio between the number of healthy samples and the number of fault samples will reach a better balance; then merging the synthetic data with good generation effect and the original data to generate integrated data; finally, inputting the integrated data into a disk fault prediction model for training, and selecting a time window of 7 days to predict whether a fault will occur after 7 days and perform corresponding data marking.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of operation and maintenance of cloud data centers, and specifically relates to a disk fault prediction method for intelligent operation and maintenance of large-scale cloud data centers. Background Art

[0002] Disks are widely used as general and main storage devices for large storage systems in modern large-scale data centers. In such data centers, ensuring high availability and reliability of data center management is a very challenging task because various disk failures occur constantly on site, which is the main factor leading to service interruption in cloud data centers. Without deploying a data redundancy scheme, disk failures may cause temporary data loss, resulting in system unavailability and even permanent data loss. Disk fault prediction, as a key link in the intelligent operation and maintenance of cloud data centers, can predict and speed up the rapid resolution of problems during operation and maintenance, enabling the server to run normally and effectively solving the problems caused by faults.

[0003] With the advent of the big data era and the rapid development of technologies such as machine learning and deep learning, people can, with the support of powerful computing power, use complex neural network models to mine and extract key information from massive data.

[0004] At the same time, with the in-depth study of SMART data, the time series-related characteristics of SMART data have gradually attracted the attention of researchers. Therefore, more and more researchers have tried to use time series processing methods to achieve disk fault prediction.

[0005] However, in practical applications, the disk fault SMART data obtained by the data center is degraded data from the healthy state to the fault state, and the actual disk fault time is unknown. That is to say, for the data of a faulty disk, we can only say that there is faulty data in its data sequence, but we cannot accurately locate it. Facing this problem, a natural solution is to treat the degraded data from the disk as a sample.

[0006] However, the degraded data of disks is often long-term serial data of different lengths. How to classify this non-fixed-length time series data is an important and challenging problem in data mining. Facing such long time series data, even the LSTM neural network cannot play a good role. In addition, during operation, the fault stage of the disk is fast and short. Therefore, the proportion of abnormal data in the life cycle data is very small, making the error information submerged in a large amount of healthy data. This is also called the imbalance problem, which poses a severe challenge to traditional classification methods.

[0007] Currently, disk fault prediction methods are mainly divided into two categories: traditional machine learning methods and deep neural network methods.

[0008] (1) Based on traditional machine learning methods, some research works use SMART attributes and Bayesian networks to predict disk failures. By means of feature selection, binning process and feature creation, a subset of SMART attributes that can best describe the data is selected and used together with a set of trend indicators based on the same SMART attributes. However, it does not well consider the temporal characteristics of the data, and the effect of dynamic Bayesian networks needs to be studied. Some other research works regard fault prediction as a binary classification problem, and at the same time consider the average time between predicted faults and actual faults, and evaluate the model performance according to the fault detection rate (FDR), which is defined as the proportion of faulty drives correctly classified as faults, while the false alarm rate (FAR) is defined as the proportion of good drives misclassified as faults.

[0009] (2) Based on deep neural network methods, some research works use long short-term memory models (LSTM) and different data balancing methods to predict disk failures 5 - 7 days in advance, solve the problem of model aging, and broaden the time range of disk failures. However, the prediction is in units of days, ignoring the IO criteria for issuing alarms, resulting in some false alarms in the prediction. Some other research works focus on predicting disk failures using sequential information. They use a dataset collected from a real-word data center, which contains 3 different disk models (denoted as W, S, and M), and establish prediction models for these disk models respectively. At the same time, they model the sequential SMART data with long-term dependencies and demonstrate the capabilities of their prediction models. Summary of the Invention

[0010] Objective of the present invention: To solve the problems existing in the prior art, the present invention provides a disk fault prediction method for intelligent operation and maintenance of large-scale cloud data centers. By utilizing the advantages of the time progressive sampling (TPS) method, integrating and optimizing the advantages of Transformer, data augmentation of unbalanced data is achieved, and disk faults are classified and predicted.

[0011] Technical solution: To solve the above technical problems, the technical solution adopted by the present invention is as follows:

[0012] In a first aspect, a disk fault prediction method for intelligent operation and maintenance of large-scale cloud data centers is provided, including:

[0013] Step 1: Fill in missing values, normalize the data, and perform information entropy processing on the unbalanced original dataset to obtain the most relevant and frequently changed feature attributes, namely dataset G;

[0014] Step 2: Divide dataset G into source dataset S1 formed by few-class samples with fault labels and source dataset S2 formed by multi-class samples with non-fault labels according to the labels;

[0015] Step 3: Use the Time Progressive Sampling (TPS) method to perform data augmentation on the source dataset S1 to generate synthetic data and obtain the synthetic dataset T.

[0016] Step 4: Integrate the source dataset S1, the source dataset S2, and the synthetic dataset T to form the integrated dataset Q, and divide the integrated dataset Q into a training set M and a test set N.

[0017] Step 5: Use the training set M to train the disk failure prediction model, and use the test set N to test the trained disk failure prediction model until the model prediction effect meets the requirements, and obtain the trained disk failure prediction model.

[0018] Step 6: Input the disk SMART data to be detected into the trained disk failure prediction model.

[0019] Step 7: Determine the disk failure prediction result according to the output of the disk failure prediction model.

[0020] In some embodiments, the missing value filling includes: if there are 2 or more consecutive missing values, use the mode of the SMART items on the disk as the filling value; if only one value is missing, use the average value before and after the value as the filling value.

[0021] In some embodiments, data normalization includes:

[0022] Scale all values to between [0, 1] using the maximum and minimum values in the feature. The scaling formula is as follows:

[0023]

[0024] where x is the original value of the feature, x max and x min are the maximum and minimum values of the feature in the dataset respectively; x' is the scaled feature value.

[0025] In some embodiments, information entropy processing includes: calculating the value of each feature attribute to represent the amount of information. The formula is as follows:

[0026]

[0027] where i represents the i-th sample, and there are a total of n samples; p represents the probability of each value appearing in each SMART attribute; the higher the information entropy value H(U) of the feature, the more information it contains, which means the more significant the volatility of the feature attribute, so as to select the most relevant and frequently changed feature attributes.

[0028] In some embodiments, in step 3, the time progressive sampling (TPS) method is used to perform data augmentation on the source dataset S1, including:

[0029] The time progressive sampling (TPS) method is used to generate and collect fault data for the source dataset S1. Calculate the loss for the generated data and the original data and determine whether it is less than the set threshold. If it is less, collect it; otherwise, repeat this step until the synthetic dataset T is collected.

[0030] In some embodiments, for a given faulty disk, assuming that the disk failure occurs at timestamp t and the prediction operation occurs at timestamp t - i, the time period of length i between the occurrence of the prediction action at t is t - i, and the occurrence of the disk failure at t is denoted as the lead time i;

[0031] During model training, for each faulty disk, TPS will gradually collect more fault data samples within the lead time i, that is, the range of the lead time i is from 1 to I, where I is a hyperparameter of TPS;

[0032] There are also two important parameters in the TPS method:

[0033] The window length h is defined as the time window size of the input data for training the network in each sequence sample, and predict_failure_days is defined as the number of days before the failure.

[0034] Furthermore, in some embodiments, if the window length h is 5, then a training sample will contain the SMART attribute information of the disk within the past 5 days; the value of predict_failure_days is within 5 - 7 days.

[0035] In some embodiments, the disk failure prediction model includes: an input module, an encoder block, a decoder block, and an output module;

[0036] In the input module, using the convolutional Transformer model, the input data L (with appropriate padding) is converted into H different query matrices using a convolutional layer with a kernel size of k (i.e., "Conv,k") and a stride of 1 Key matrix And value matrix where h = 1, …, H, All are learnable parameters;

[0037] When outputting from the encoder block and the decoder block, a feed - forward sub - layer is stacked respectively. The position feed - forward sub - layer has two fully - connected networks and a ReLU activation in the middle. The formula is as follows:

[0038] max(0, XW1 + b1)W2 + b2 (3)

[0039] Among them, X is the input, W1 and W2 are learnable parameters, and b1 and b2 are preset regularization terms used for the conversion of the spatial dimension in the connection layer. The dimension of the output matrix finally obtained by the feed-forward sublayer is the same as that of X;

[0040] By replacing the existing position-based linear projection with a convolutional projection for the attention calculation operation, and using the convolutional projection for the embedding of queries, keys, and values to enhance the attention to local context information.

[0041] In a second aspect, the present invention provides a disk failure prediction device for intelligent operation and maintenance of large-scale cloud data centers, including a processor and a storage medium;

[0042] The storage medium is used to store instructions;

[0043] The processor is used to operate according to the instructions to execute the steps of the method according to the first aspect.

[0044] In a third aspect, the present invention provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to the first aspect are implemented.

[0045] The present invention has the following beneficial effects:

[0046] (1) For the disk failure prediction in the cloud data center, taking advantage of the time progressive sampling method TPS, integrating and optimizing the advantages of the Transformer. This method can make full use of the failure data to extract the relationships between the data, and enable the generated data to have the original features of the failure data and the data distribution and hidden patterns in the latent space, achieving excellent results on the public data, and having good practicability in the disk failure prediction system with high requirements for the F1 value and the Matthews correlation coefficient (MCC).

[0047] (2) In this method for long time series data, we use the multi-head self-attention mechanism and stack multiple encoder-decoder models to learn and obtain the dependence relationships between the time series data, and establish the time correlations between the data of different time steps.

[0048] (3) In this method, since the self-attention calculation method of the original Transformer is insensitive to local information, the model is vulnerable to outliers, bringing potential optimization problems. Therefore, a convolutional projection is used to replace the existing position-based linear projection for the attention calculation operation, and the convolutional projection is used for the embedding of queries, keys, and values to enhance the attention to local context information, making the prediction more accurate.

[0049] (4) In this method, the Time Progressive Sampling (TPS) method is used for data augmentation to solve the data imbalance problem. TPS can generate multiple failed samples for each failed disk, which not only preserves all the characteristics of healthy disks but also brings more failure modes.

[0050] (5) The algorithm structure of this method is simple and has a low time complexity. Description of the Drawings

[0051] Figure 1 It is a flowchart of the disk fault prediction method for intelligent operation and maintenance of large-scale cloud data centers designed in the embodiments of the present invention.

[0052] Figure 2 It is a disk fault prediction model diagram in the embodiments of the present invention.

[0053] Figure 3 It is a design diagram of the Time Progressive Sampling (TPS) method in the embodiments of the present invention. Detailed Embodiments

[0054] The technical solutions of the present invention will be further described in detail below with reference to the drawings in the specification.

[0055] In the description of the present invention, the meaning of several is more than one, the meaning of multiple is more than two, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If there is a description of first and second, it is only for the purpose of distinguishing technical features and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or implicitly indicating the sequence relationship of the indicated technical features.

[0056] In the description of the present invention, the description referring to terms such as "one embodiment", "some embodiments", "illustrative embodiments", "examples", "specific examples", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0057] Embodiment 1

[0058] A disk fault prediction method for intelligent operation and maintenance of large-scale cloud data centers includes:

[0059] Step 1: Fill in missing values, normalize the data, and perform information entropy processing on the imbalanced original data set to obtain the most relevant and frequently changed feature attributes, that is, data set G;

[0060] Step 2: Divide the dataset G into the source dataset S1 formed by few-class samples with fault labels and the source dataset S2 formed by multi-class samples with non-fault labels according to the labels;

[0061] Step 3: Use the Time Progressive Sampling (TPS) method to perform data augmentation on the source dataset S1, generate synthetic data, and obtain the synthetic dataset T;

[0062] Step 4: Integrate the source dataset S1, the source dataset S2, and the synthetic dataset T to form the integrated dataset Q, and divide the integrated dataset Q into the training set M and the test set N;

[0063] Step 5: Use the training set M to train the disk fault prediction model, and use the test set N to test the trained disk fault prediction model until the model prediction effect meets the requirements, and obtain the trained disk fault prediction model;

[0064] Step 6: Input the disk SMART data to be detected into the trained disk fault prediction model;

[0065] Step 7: Determine the disk fault prediction result according to the output of the disk fault prediction model.

[0066] In some embodiments, the missing value filling includes: if there are 2 or more consecutive missing values, use the mode of the SMART items on the disk as the filling value; if only one value is missing, use the average value before and after the value as the filling value.

[0067] In some embodiments, the data normalization includes:

[0068] Scale all values to between [0, 1] using the maximum and minimum values in the features. The scaling formula is as follows:

[0069]

[0070] where x is the original value of the feature, x max and x min are the maximum and minimum values of the feature in the dataset respectively; x′ is the scaled feature value.

[0071] In some embodiments, the information entropy processing includes: calculating the value of each feature attribute to represent the amount of information, and the formula is as follows:

[0072]

[0073] Where \(i\) represents the \(i\)th sample, and there are a total of \(n\) samples; \(p\) represents the probability of each value appearing in each SMART attribute; the higher the information entropy value \(H(U)\) of the feature, the more information it contains, which means the more significant the volatility of the feature attribute, so as to select the most relevant and frequently changed feature attributes.

[0074] In some embodiments, in step 3, the time progressive sampling (TPS) method is used to perform data augmentation on the source dataset \(S1\), including:

[0075] The time progressive sampling (TPS) method is used to generate and collect failure data for the source dataset \(S1\). Calculate the loss for the generated data and the original data and determine whether it is less than the set threshold. If it is less, collect it; otherwise, repeat this step until the synthetic dataset \(T\) is collected.

[0076] In some embodiments, for a given failed disk, assume that the disk failure occurs at timestamp \(t\), and the prediction operation occurs at timestamp \(t - i\). The time period \(t - i\) with a length of \(i\) between the prediction action at \(t\) represents the lead time \(i\) when the disk failure occurs at \(t\).

[0077] During model training, for each failed disk, TPS will gradually collect more failure data samples within the lead time \(i\), that is, the range of the lead time \(i\) is from 1 to \(I\), where \(I\) is a hyperparameter of TPS.

[0078] There are also two important parameters in the TPS method:

[0079] The window length \(h\) is defined as the time window size of the input data for training the network in each sequence sample, and predict_failure_days is defined as the number of days before failure.

[0080] Furthermore, in some embodiments, if the window length \(h\) is 5, then a training sample will contain the SMART attribute information of the disk within the past 5 days; the value of predict_failure_days is within 5 - 7 days.

[0081] In some embodiments, the disk failure prediction model includes: an input module, an encoder block, a decoder block, and an output module;

[0082] In the input module, using the convolutional Transformer model, the input data \(L\) (with appropriate padding) is converted into \(H\) different query matrices using a convolutional layer with a kernel size of \(k\) (i.e., "Conv,k") and a stride of 1. Key matrix And value matrix where \(h = 1,\ldots,H\). are all learnable parameters;

[0083] When the encoder block and the decoder block output, a feed-forward sub-layer is stacked respectively. The position feed-forward sub-layer has two layers of fully connected networks and a ReLU activation in the middle. The formula is as follows:

[0084] max(0, XW1 + b1)W2 + b2 (3)

[0085] where X is the input, W1 and W2 are learnable parameters, and b1 and b2 are preset regularization terms used to perform the conversion of the spatial dimension in the connection layer. The dimension of the output matrix finally obtained by the feed-forward sub-layer is the same as that of X;

[0086] Replace the existing position-based linear projection with a convolutional projection to perform the attention calculation operation, and use the convolutional projection to perform the embedding of queries, keys, and values to enhance the attention to local context information.

[0087] In some specific embodiments, as Figure 1 shown, a disk fault prediction method for intelligent operation and maintenance of large-scale cloud data centers includes the following steps: First, perform information entropy feature processing on the imbalanced data to select more important features; then divide the processed imbalanced data to extract the sample data of the minority class, that is, the fault samples; then use the time progressive sampling method TPS to perform data augmentation on the fault samples to generate synthetic data. By using TPS, more fault sample data are generated, so that the ratio between the number of healthy samples and the number of fault samples will reach a better balance; then merge the synthetic data with good generation effect and the original data to generate integrated data; finally, input the integrated data into the disk fault prediction model for training, and select a time window of 7 days to predict whether a fault will occur after 7 days and perform corresponding data marking. The present invention can make full use of fault data to extract the dependence relationship between time series data, and can make the generated data have the original features of fault data and the data distribution and hidden patterns in the latent space, and has good practicability in a disk fault prediction system with high requirements for the F1 value and the Matthews correlation coefficient (MCC).

[0088] The disk fault prediction method of this embodiment is used to predict the faults of disks for intelligent operation and maintenance of large-scale cloud data centers. In the actual application process, it specifically includes the following steps:

[0089] Step 1, fill in the missing values of the features in the original data set, then perform data normalization to scale all values to the interval of [0, 1], and then perform information entropy feature processing to select the most relevant and frequently changed features, and finally obtain the data set G after feature processing.

[0090] The method for filling missing values is as follows: If there are 2 or more consecutive missing values, the mode of the SMART items on disk is used as the filling value; if only one value is missing, the average of the values before and after it is used as the filling value. Then, data normalization is performed, and all values are scaled to between [0, 1] using the maximum and minimum values in the features. The scaling formula is as follows:

[0091]

[0092] where x is the original value of the feature, x max and x min are the maximum and minimum values of the feature in the dataset respectively, and x′ is the scaled feature value. Immediately afterwards, we perform information entropy processing on the features after missing value filling and data normalization. This method calculates the values of each feature attribute to represent the amount of information. The formula is as follows:

[0093]

[0094] where i represents the i-th sample, and there are a total of n samples; where p represents the probability of each value appearing in each SMART attribute. The higher the information entropy value H(U) of the feature, the more information it contains, which means the more significant the volatility of the feature attribute, and thus the most relevant and frequently changed feature attributes can be selected.

[0095] Step 2: Divide the labeled imbalanced fault data in the dataset G, and select the sample data of the minority class from it, that is, the label is 1, and use it as the source dataset S1. The other class of samples, that is, the label is 0, is used as the source dataset S2;

[0096] Step 3: Adopt the Time Progressive Sampling (TPS) method. Generate and collect fault data from the source dataset S1 (i.e., the minority class samples) through TPS. Then, calculate the loss of the generated data and the original data and determine whether it is less than the set threshold. If it is less, collect it; otherwise, repeat the operation in Step 2. The Time Progressive Sampling (TPS) method is used to perform data augmentation on the minority class sample fault data. Before describing TPS, there is an important concept called lead time. For a given faulty disk, assume that the disk failure occurs at timestamp t, and the prediction operation occurs at timestamp t - i. Then, the time period of length i from t - i to the occurrence of the prediction action at t is the lead time i for the occurrence of the disk failure at t. During model training, for each faulty disk, TPS will gradually collect more fault data samples within the lead time i (i.e., the range of the lead time i is from 1 to I, where I is a hyperparameter of TPS). And there are two important parameters in the TPS method:

[0097] The window length h is defined as the time window size of the input data for training the network in each sequence sample. For example, if h is 5, then a training sample will contain the SMART attribute information of the disk in the past 5 days. h needs to have an appropriate value. If it is too small, less potential information is provided to ConvTrans-TPS. If it is too large, it corresponds to a long time series. Data that is too far from the final failure has little impact on predicting the final failure trend and may even be misleading.

[0098] predict_failure_days is defined as the number of days before failure, which is an alarm boundary. The value of predict_failure_days also needs to be appropriate. Too long or too short a time interval will affect the effectiveness of disk failure handling. A value of predict_failure_days within 5 - 7 days is reasonable. The value of predict_failure_days selected by this method is 7 days.

[0099] Step 4, repeat Step 2 until the generation of the source dataset S1 ends, collect the synthesized data, and use it as the synthesized dataset T;

[0100] Step 5, integrate the source dataset S1, the source dataset S2, and the synthesized dataset T to form the final integrated dataset Q, and divide it into a training set M and a test set N. The label counts in the selected data are shown in Table 1

[0101] Table 1 Statistical table of dataset labels

[0102]

[0103] Step 6, construct and train a disk failure prediction model using the training set M, and predict the test set N. The disk failure prediction model outputs possible failures in the test set N. The disk failure prediction model is modified based on the Transformer model, discarding the attention calculation operation of the position-based linear projection on the original Transformer model, but using convolutional projection for query, key, and value embedding to enhance the attention to local context information, and at the same time including the original encoder block and decoder block. In the original Transformer model, the long-term and short-term dependencies are captured by using the multi-head self-attention mechanism, and different attention heads learn and focus on different aspects of the time pattern.

[0104] In the self-attention layer, the multi-head self-attention sublayer (applying the same model at each time step, so some symbols are used to simplify the formula) simultaneously transforms the input data L into H different query matrices Key matrix And value matrix where h = 1, …, H, and are all learnable parameters. After these linear projections, scaled dot-product attention computes the vector output sequence:

[0105]

[0106] where the masking matrix M is used to avoid future information leakage by setting all upper triangular elements to -∞, and d k is the number of columns of the Q and K matrices, i.e., the vector dimension. After that, A1, A2, …, A H will be concatenated and linearly projected again. At the output, a feed-forward sub-layer is stacked, which has two fully-connected networks with a ReLU activation in the middle, as shown in the formula:

[0107] max(0, XW1 + b1)W2 + b2 (4)

[0108] where X is the input, and the dimension of the output matrix finally obtained by this feed-forward sub-layer is the same as that of X.

[0109] However, in the original Transformer model, the similarity between queries and keys is calculated based on their dot product, which may lead to abnormal focus of data. The original calculation method cannot consider the information of the current data and is insensitive to local context, resulting in the attention score only reflecting the correlation between single time points, which is different from the original intention of time series prediction. It may confuse whether the observed values of the self-attention module are outliers, change points, or part of the pattern, and bring potential optimization problems.

[0110] Therefore, a convolutional self-attention mechanism is used to alleviate this problem. It uses a convolutional layer with a kernel size of k and a stride of 1 to convert the input (with appropriate padding) into queries and keys, instead of using a kernel size of 1 and a stride of 1 (matrix multiplication). The attention calculation operation is performed by replacing the existing position-based linear projection with a convolutional projection, and the convolutional projection is used for the embedding of queries, keys, and values to enhance the attention to local context information.

[0111] Step 7: Use the output of the disk failure prediction model to label all instances in the test set N, and obtain the labeling result c. When the labeling value of the labeling result c is 0, it means the instance is non-faulty; when the labeling value is 1, it means the instance is faulty. Among them, in step 7, the labeling result c is calculated using equation (5):

[0112]

[0113] In equation (4), is the i-th feature of sample k, D is all the data sets available for model training, Dk is a subset of D, y j is the eigenvalue of sample j, a is a parameter, and p is a prior value.

[0114] Step 8: Output a fault according to the marking result c.

[0115] A positioning system for a disk fault prediction method for intelligent operation and maintenance of large-scale cloud data centers according to this embodiment includes:

[0116] A dataset feature preprocessing module for filling missing values, normalizing data, and performing information entropy processing on the original dataset to obtain the most relevant and frequently changed feature attributes;

[0117] A dataset partitioning module for screening out minority-class samples from the imbalanced dataset to form a source dataset S1 and majority-class samples to form a source dataset S2;

[0118] A Time Progressive Sampling (TPS) module for data augmentation of the source dataset S1 to generate high-quality synthetic data as a synthetic dataset T;

[0119] A disk fault prediction model for training a training set M partitioned from the final integrated dataset Q and predicting the test set N to output the faults of each instance in the test set N;

[0120] A marking module for marking all instances in the test set N to obtain a marking result;

[0121] A display module for predicting and displaying faults in the network according to the marking result.

[0122] The Time Progressive Sampling (TPS) module generates synthetic data T and merges it with the source dataset S1 and the source dataset S2 to form an integrated dataset Q. The integrated dataset Q is used to build a disk fault prediction model.

[0123] Embodiment 2

[0124] Second, this embodiment provides a disk fault prediction device for intelligent operation and maintenance of large-scale cloud data centers, including a processor and a storage medium;

[0125] The storage medium is used to store instructions;

[0126] The processor is used to operate according to the instructions to execute the steps of the method according to Embodiment 1.

[0127] Embodiment 3

[0128] Third, this embodiment provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to Embodiment 1 are implemented.

[0129] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0130] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0131] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that realizes the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0132] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or a plurality of flows and / or blocks

[0133] The above is only the preferred embodiment of the present invention, and the protection scope of the present invention is not limited to the above embodiment. Any equivalent modification or change made by those of ordinary skill in the art according to the disclosure of the present invention should be included in the protection scope recorded in the claims.

Claims

1. A disk failure prediction method, characterized in that, Including: Step 1: Fill in missing values, normalize the data, and perform information entropy processing on the unbalanced original dataset to obtain the most relevant and frequently changed feature attributes, namely dataset G; Step 2: Divide dataset G into source dataset S1 formed by few-class samples with fault labels and source dataset S2 formed by multi-class samples with non-fault labels according to the labels; Step 3: Use the Time Progressive Sampling (TPS) method to perform data augmentation on source dataset S1 to generate synthetic data and obtain synthetic dataset T, including: using the TPS method to generate and collect fault data from source dataset S1, calculating the loss between the generated data and the original data, and determining whether it is less than the set threshold. If it is less, collect it; otherwise, repeat this step until synthetic dataset T is obtained; Step 4: Integrate source dataset S1, source dataset S2, and synthetic dataset T to form integrated dataset Q, and divide integrated dataset Q into training set M and test set N; Step 5: Use training set M to train the disk fault prediction model, and use test set N to test the trained disk fault prediction model until the model prediction effect meets the requirements to obtain the trained disk fault prediction model; Step 6: Input the disk SMART data to be detected into the trained disk fault prediction model; Step 7: Determine the disk fault prediction result according to the output of the disk fault prediction model; The disk fault prediction model includes: an input module, an encoder block, a decoder block, and an output module; In the input module, using a convolutional Transformer model, with a convolutional layer having a kernel size of k and a stride of 1, the input data L is converted into H different query matrices Key matrix Sum value matrix where h = 1, …, H, are all learnable parameters; When the encoder block and the decoder block output, a feed-forward sub-layer is stacked respectively. The feed-forward sub-layer has two fully connected networks and a ReLU activation in the middle. The formula is as follows: max(0, XW1 + b1)W2 + b2 where X is the input, W1 and W2 are learnable parameters, and b1 and b2 are preset regularization terms used for the conversion of the spatial dimension in the connection layer. The dimension of the output matrix finally obtained by the feed-forward sub-layer is the same as that of X.

2. The disk failure prediction method according to claim 1, wherein, The missing value filling includes: if there are 2 or more consecutive missing values, use the pattern of the SMART item on the disk as the filling value; if only one value is missing, use the average value before and after this value as the filling value.

3. The disk failure prediction method according to claim 1, wherein Data normalization includes: Scale all values to between [0, 1] using the maximum and minimum values in the features. The scaling formula is as follows: where x is the original value of the feature, x max and x min are the maximum and minimum values of the feature in the dataset respectively; x' is the scaled feature value.

4. The disk failure prediction method according to claim 1, wherein The information entropy processing includes: calculating the value of each feature attribute to represent the amount of information. The formula is as follows: where i represents the i-th sample, and there are n samples in total; p represents the probability of each value appearing in each SMART attribute; the higher the information entropy value H(U) of the feature, the more information it contains, which means the more significant the volatility of the feature attribute, so as to select the most relevant and frequently changed feature attributes.

5. The disk failure prediction method according to claim 1, wherein For a given faulty disk, assuming that the disk fault occurs at timestamp t and the prediction operation occurs at timestamp t - i, the time period of length i between the prediction action at t - i and the occurrence of the disk fault at t is represented as the lead time i; During model training, for each failed disk, the TPS will gradually collect more failure data samples within the lead time i, where the range of the lead time i is from 1 to I, and I is a hyperparameter of the TPS; There are also two important parameters in the TPS method: The window length h is defined as the size of the time window for the input data of the training network in each sequence sample, and predict_failure_days is defined as the number of days before failure.

6. The disk failure prediction method according to claim 5, wherein If the window length h is 5, then a training sample will contain the SMART attribute information of the disk within the past 5 days; the value of predict_failure_days is within 5 - 7 days.

7. A disk failure prediction device, characterized in that, Including a processor and a storage medium; The storage medium is used to store instructions; The processor is used to operate according to the instructions to execute the steps of the method according to any one of claims 1 to 6.

8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Fault handling method, apparatus, computer apparatus, and storage medium

    CN109144829A

  • Disk fault prediction method, apparatus and device, and storage medium

    CN111782491A