Cloud platform disk fault prediction system based on lightweight adaptive model

By designing a disk failure prediction system based on lightweight adaptive models on the cloud platform, problems such as difficulty in data acquisition and long model training time in traditional technology are solved, and efficient and accurate disk failure prediction and the ability to quickly respond to business changes are achieved.

CN120011193APending Publication Date: 2025-05-16EAST CHINA NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145155.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing disk failure prediction technology has problems such as difficulty in collecting data, long model training time, strong invasiveness to hardware, and inability to adjust dynamically, which affects the prediction accuracy and user experience.

Method used

A cloud platform disk failure prediction system based on lightweight adaptive model was designed. By analyzing the disk system parameter characteristics, the timing prediction model and outlier detection method are used to predict disk performance and detect failure risks. The adaptive module can automatically issue retraining commands to adapt to the upper-level read and write frequency changes and shorten the training time.

Benefits of technology

It realizes non-hardware intrusion disk performance data acquisition, uses lightweight prediction models to improve prediction accuracy, and the adaptive module enables the model to respond quickly to business changes and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011193A_ABST
    Figure CN120011193A_ABST
Patent Text Reader

Abstract

The invention discloses a cloud platform disk fault prediction system based on a lightweight self-adaptive model, and aims to solve the problems of large required original data volume, long model training time, high invasiveness to bottom hardware and the like in an existing disk fault prediction method. According to the system, disk performance parameters are collected through the data acquisition module, fault prediction is performed by using the lightweight time sequence prediction model and the abnormal value detection module, and the prediction model is dynamically adjusted through the adaptive module to adapt to upper-layer business changes, so that prediction accuracy and user experience are improved. The method comprises the following steps: firstly, collecting performance data of a disk operating system by using a data acquisition module, and inputting a time sequence prediction module training model for prediction; secondly, the abnormal value detection module compares input prediction data and actual operation data at the same time and judges whether an abnormal risk exists or not; and finally, quickly identifying the read-write frequency change of the upper-layer business by utilizing a self-adaptive detection module, and retraining the prediction model for prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disk failure prediction, and in particular to a cloud platform disk failure prediction system based on a lightweight adaptive model. Background Art

[0002] In today's digital age, information technology plays a key role in the operation of enterprises and organizations. With the rapid growth of data volume and the increasing complexity of IT environment, traditional operation and maintenance management methods can no longer meet the needs, and intelligent operation and maintenance (AIOPS) has emerged. Among them, disk failure prediction is an important part of AIOPS. As a basic device for data storage, disks are widely used in data centers, servers, and personal computers. However, their sensitive, precise and complex structure make failures difficult to avoid. There are many reasons for disk failures. For example, high temperature, humidity, mechanical wear, read and write operation frequency and other factors interact with each other, making the failure mode complex and the prediction more difficult. Traditional fault prediction methods are based on fixed thresholds and empirical judgments, and have obvious limitations such as only being able to take action when a fault occurs or is about to occur and being prone to false alarms.

[0003] There are two main methods for predicting disk failures: prediction schemes based on SVM machine learning algorithms and prediction schemes that combine deep learning with big data analysis. For the former, it often relies on analyzing and learning the SMART (Self-Monitoring Analysis and Reporting Technology) information of a large number of hard disks to generate a model that can accurately predict hard disk failures. However, SMART data collection usually requires intrusion into the underlying hardware, and the collection methods for machines from different manufacturers are also different. Data collection is often a huge problem in actual predictions, and the actual effect of cloud platforms built by servers from multiple manufacturers is not good. For the latter, there are problems such as long training time for deep learning networks, high resource requirements, and large amounts of raw data required. RNN (recurrent neural network) calculations rely on time steps, and its sequential calculation method makes it difficult to parallelize, resulting in a long training time. CNN (convolutional neural network) often requires a large number of convolution kernels and multi-layer structures, which will result in a large number of model parameters. A large number of parameters will increase the training time of the model. Generally, deep learning networks rely on a large amount of historical disk data. In actual production applications, it may happen that a model that takes a long time to train is not applicable after production, and data needs to be collected again. This situation often occurs when the upper-level business model changes due to new requirements, resulting in changes in the disk read and write frequency.

[0004] In summary, the two current mainstream disk failure prediction technologies face different problems, which greatly affect the prediction accuracy and user experience. The main problems are that the amount of raw data required is large, the model training time is long, the invasiveness to the underlying hardware is strong, and it cannot be dynamically adjusted according to changes in the business model. Summary of the invention

[0005] In order to overcome the problems existing in the above-mentioned technologies, the purpose of the present invention is to provide a cloud platform disk failure prediction system based on a lightweight adaptive model. The system analyzes the parameter characteristics of the cloud platform disk system and designs an adaptive lightweight prediction model and an outlier judgment method according to the disk performance rules of the upper-level application system under different production cycles. The system predicts the performance of the disk within the production cycle window and determines whether the disk is faulty according to the outlier detection method. At the same time, the adaptive module can automatically issue a fast retraining command so that the prediction model can adapt to the changes in the upper-level read and write frequency, greatly shortening the training time and improving the user experience.

[0006] The specific technical solution for achieving the purpose of the present invention is: A cloud platform disk failure prediction system based on a lightweight adaptive model, comprising: The data collection module collects the cloud platform disk timing performance data, including historical data and online data; the historical data is input into the timing prediction module and trained to obtain a lightweight prediction model; the online data is simultaneously input into the timing prediction module and the anomaly detection module for prediction and real-time detection; The time series prediction module uses the lightweight prediction model trained with the historical data to predict the disk performance data of the next time window based on the online data, and so on to perform real-time performance prediction, and inputs the prediction results into the anomaly detection module; The anomaly detection module compares the disk performance prediction data with the online data to determine whether there is a risk of disk failure and inputs the judgment result into the adaptive detection module; The adaptive detection module makes adaptive adjustments based on the results of anomaly detection to determine whether the business read and write frequency has changed. If it is determined that the read and write frequency has changed, an instruction is issued to the timing prediction module to retrain the lightweight prediction model using the latest collected data. If it is determined that the read and write frequency remains the same, the abnormal result is reported as a disk failure warning based on the judgment result of the anomaly detection module.

[0007] Furthermore, the data acquisition module collects the disk timing performance data through operating system commands, which includes average waiting time, I / O utilization, number of read requests per second, number of write requests per second, number of requests in the queue, and number of responses of the device, instead of the underlying SMART data. There is no need to rely on the disk manufacturer's disk hardware monitoring program collection interface.

[0008] Furthermore, the timing prediction module adopts the SparseTSF prediction model, which utilizes the prior periodicity of disk read and write performance and reduces the scale of model parameters and dependence on original data through cross-period sparse prediction. By setting the prediction step size, the prediction data for the next step is continuously obtained based on the performance data of each step size.

[0009] Furthermore, the anomaly detection module sets a sliding detection time window, uses a weighted comparison method to compare the predicted performance data with the actual collected data, normalizes each parameter first, and then obtains the difference by weighted averaging; if the difference exceeds the fault threshold, it is recorded as a detection failure, and the disk failure risk is reported only when each identification within the detection window fails.

[0010] Furthermore, the adaptive detection module counts and finds that if more than half of the disks on the same server report fault warnings in the same cycle, it determines that the read and write frequency changes; at this time, the model training flag is modified to retrain the lightweight prediction model to improve the prediction accuracy and ease of use.

[0011] The present invention trains a time series prediction model by collecting disk system parameters, defines whether the disk is faulty by sliding time windows and similar distance methods, and designs an adaptive module by combined analysis and comparative verification. The time series prediction model used in the present invention presets the disk performance regularity cycle in advance, so fewer training parameters and initial data can be used to achieve a good prediction effect. The adaptive module can detect whether the performance regularity has changed due to changes in the upper-level business model, automatically issue retraining instructions, and make the model automatically adapt to the upper-level model, thereby improving the model prediction accuracy and user experience.

[0012] Compared with the prior art, the benefit of the present invention is to realize a cloud platform disk failure prediction system based on a lightweight adaptive model, collect disk performance through non-hardware intrusion, reasonably use the lightweight prediction model to predict disk failure, and through the adaptive module, the prediction model parameters can be quickly updated as the application is put into production, thereby improving the model's ease of use and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 It is a structural diagram of the present invention; Figure 2 Sample graph of disk collection data indicators; Figure 3 This is the flow chart of the anomaly detection module; Figure 4 This is the flow chart of the adaptive module. DETAILED DESCRIPTION

[0014] The present invention is described in detail below with reference to the accompanying drawings.

[0015] The operation process of the cloud platform disk failure prediction system based on a lightweight adaptive model proposed in the present invention is as follows: First, during the operation of the system, the system performance parameters of the disk are collected, and the prediction model is trained after reaching a certain order of magnitude. Secondly, after the model training is completed, the system performance of the next step is predicted based on the currently collected data. The outlier detection method will compare the predicted performance with the actual performance and report the disk failure warning. Finally, the adaptive module will detect whether the current model needs to be adjusted (such as Figure 1 It can be divided into three aspects.

[0016] First, a disk performance data collection and prediction model training mechanism is provided, including: disk performance data collection mainly collects system-level operating parameters rather than underlying SMART parameters, including average waiting time, I / O utilization, number of read requests per second, number of write requests per second, number of requests in the queue, number of device responses, etc. These parameters can be collected using operating system commands, without the need to collect them through the disk manufacturer's disk hardware monitoring program interface, thereby avoiding dependence on hardware manufacturers and intrusion on physical devices (such as Figure 2 As shown). When the raw data is collected to a certain order of magnitude, the prediction model will be trained. The prediction model uses the time series prediction model SparseTSF, which relies on fewer model parameters and raw data, consumes lower computing resources but can still achieve good prediction accuracy. Given that the data to be predicted usually exhibits a constant prior periodicity (for example, the disk read and write performance is usually a daily period when the upper-level application model remains unchanged), SparseTSF uses cross-period sparse prediction technology to enhance the extraction of long-term sequential dependencies while reducing the parameter scale of the model. A single linear layer is used within the model framework to model the LTSF task. The previously collected disk performance data is input into the model. When the loss function reaches the target value, the method has a good prediction effect, as judged by the mean square error (MSE) and mean absolute error (MAE).

[0017] On the second aspect, a disk performance data anomaly detection based on predicted values ​​is provided, including: first, according to the prediction step L (for example, assuming that the prediction step L is one hour, that is, inputting the current one hour data, and predicting the disk performance of the next hour through the model), it is necessary to set a detection sliding time window W (for example, assuming that the time window is 3L, that is, testing is performed on 3 prediction results), and each identification within the detection time window needs to reach the failure threshold before the disk failure risk is reported. Secondly, based on the importance of the disk performance data, a weighted comparison method is applied to determine whether there is a failure risk. That is, the performance data obtained by prediction is compared with the actual collected performance data, and each parameter is normalized and weighted averaged after the difference is made, and finally compared with the failure threshold. If the failure threshold is exceeded, a detection failure is recorded. (such as Figure 3 shown) On the third aspect, an adaptive module is provided to detect whether the current model needs to be adjusted, including: first, the prediction requires a priori assumption that the data has a periodic pattern in order to achieve accurate prediction results. This happens to be in line with actual production. In a production environment, changes are not implemented all the time. Usually, applications involving changes in read and write frequency are put into production on a monthly or quarterly basis. Therefore, based on the accuracy of the prediction results, it is possible to reversely judge whether there is a situation where the read and write frequency changes are caused by the production. Secondly, for cloud platform storage clusters, the load is usually scattered on all disks of each machine, so the predictions of all disks on the same machine can be compared. If more than half of the disks report a fault warning at the same time during the anomaly detection cycle, it is considered that the upper-level application write model has changed, rather than a possible failure of a single disk. At this time, the adaptive module will modify the model training flag to retrain. At this time, the model collects and updates the data, and retrains using the data from the most recent period (such as Figure 4 as shown).

[0018] like Figure 1 As shown, the present invention focuses on designing four modules in the disk failure prediction system: data acquisition module, time series prediction module, abnormal value detection module, and adaptive detection module. It is mainly implemented in the following steps: 1. During the operation of the distributed storage device on the cloud platform, the data acquisition module collects disk performance data through operating system commands to form a performance data time series, including historical data and online data.

[0019] 2. The time series prediction module uses historical data to train the prediction model. The model can achieve good prediction results with a small number of parameters and original data by presetting the data period. When the model training is completed, the online time series is predicted.

[0020] 3. The anomaly detection module compares the predicted data and actual operation data input at the same time to determine whether there is an abnormal risk. If there is, the anomaly detection counter is incremented by one, and the next result is determined based on the sliding window. If there is a risk in three consecutive anomaly detections, the disk failure risk is reported externally.

[0021] 4. As the system runs, if the upper-layer read and write frequency changes due to business production, the adaptive module will analyze the comprehensive abnormal detection results of multiple disks on the same server and determine that there is a change in the upper-layer read and write frequency, thereby triggering model retraining. While ensuring the accuracy of prediction, the actual usability is greatly improved.

Claims

1. A cloud platform disk failure prediction system based on a lightweight adaptive model, characterized in that: include: The data collection module collects the cloud platform disk timing performance data, including historical data and online data; the historical data is input into the timing prediction module and trained to obtain a lightweight prediction model; the online data is simultaneously input into the timing prediction module and the anomaly detection module for prediction and real-time detection; The time series prediction module uses the lightweight prediction model trained with the historical data to predict the disk performance data of the next time window based on the online data, and so on to perform real-time performance prediction, and inputs the prediction results into the anomaly detection module; The anomaly detection module compares the disk performance prediction data with the online data to determine whether there is a risk of disk failure and inputs the judgment result into the adaptive detection module; The adaptive detection module makes adaptive adjustments based on the results of anomaly detection to determine whether the business read and write frequency has changed. If it is determined that the read and write frequency has changed, an instruction is issued to the timing prediction module to retrain the lightweight prediction model using the latest collected data. If it is determined that the read and write frequency remains the same, the abnormal result is reported as a disk failure warning based on the judgment result of the anomaly detection module.

2. The cloud platform disk failure prediction system according to claim 1, characterized in that: The data collection module collects the disk timing performance data through operating system commands, including average waiting time, I / O utilization, number of read requests per second, number of write requests per second, number of requests in the queue and number of device responses.

3. The cloud platform disk failure prediction system according to claim 1, characterized in that: The timing prediction module adopts the SparseTSF prediction model, utilizes the prior periodicity of disk read and write performance, and reduces the scale of model parameters and the dependence on original data through cross-period sparse prediction; by setting the prediction step length, the prediction data of the next step length is continuously obtained based on the performance data of each step length.

4. The cloud platform disk failure prediction system according to claim 1, characterized in that: The anomaly detection module sets a sliding detection time window, uses a weighted comparison method to compare the predicted performance data with the actual collected data, normalizes each parameter first, and then obtains the difference by weighted averaging; if the difference exceeds the fault threshold, it is recorded as a detection failure, and the disk failure risk is reported only when each identification within the detection window fails.

5. The cloud platform disk failure prediction system according to claim 1, characterized in that: The adaptive detection module counts that if more than half of the disks on the same server report fault warnings in the same cycle, it determines that the read and write frequency changes; at this time, the training flag is modified to retrain the lightweight prediction model to improve the prediction accuracy and ease of use.