Server fault prediction system based on deep learning

Through a server fault prediction system based on deep learning, the data acquisition interval is dynamically adjusted, which solves the problems of wasted storage space and inaccurate prediction results in server fault prediction, and improves the accuracy and accuracy of prediction.

CN120371634AInactive Publication Date: 2025-07-25刘正海
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510419707.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-25
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the data acquisition interval for server failure prediction is fixed, resulting in waste of storage space or insufficient prediction results.

Method used

The server fault prediction system based on deep learning is adopted to calculate the relationship value and stability of the monitoring data and the prediction results, and dynamically adjust the acquisition interval of the monitoring data to reduce the amount of data and improve the data accuracy.

Benefits of technology

It achieves the accuracy and accuracy of server failure prediction while reducing storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120371634A_ABST
    Figure CN120371634A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of server fault management, and discloses a server fault prediction system based on deep learning, which comprises a data collection module and a fault prediction module, the data collection module is used for obtaining monitoring data of a preset type in the operation process of the server based on the collection interval of the monitoring data; the fault prediction module is used for performing fault prediction on the server based on the monitoring data acquired by the data collection module and a deep learning technology to obtain a prediction result; and the data collection module is also used for calculating the collection interval of each type of monitoring data according to the prediction result. According to the method, the larger acquisition interval can be adopted when the influence of the monitoring data on the prediction result is smaller, and the larger acquisition interval can be adopted when the stability degree of the monitoring data is higher, so that the quantity of the acquired data can be reduced, and the storage space is effectively saved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of server fault management, and particularly to a server fault prediction system based on deep learning. Background Art

[0002] Before predicting server faults, it is usually necessary to collect various operation data of the server, including CPU usage rate, memory usage rate, disk write volume, disk read volume, network upload volume, network download volume, etc. In the prior art, these data are usually acquired based on a fixed acquisition interval, which results in too large an acquisition interval for data that has little impact on the prediction result, wasting storage space, while for data that has a greater impact on the prediction result, the acquisition interval is small, resulting in insufficient data accuracy and inaccurate prediction results. Summary of the Invention

[0003] The purpose of the present invention is to disclose a server fault prediction system based on deep learning to solve the technical problems proposed in the background art.

[0004] To achieve the above purpose, the present invention provides the following technical solutions:

[0005] The present invention provides a server fault prediction system based on deep learning, including a data collection module and a fault prediction module;

[0006] The data collection module is used to obtain preset types of monitoring data during the operation process of the server based on the acquisition interval of the monitoring data;

[0007] The fault prediction module is used to perform fault prediction on the server based on the monitoring data obtained by the data collection module and deep learning technology to obtain a prediction result;

[0008] The data collection module is further used to calculate the acquisition interval of each type of monitoring data according to the prediction result, including:

[0009] For the monitoring data of type A, the calculation formula for its acquisition interval is:

[0010]

[0011] d A represents the acquisition interval of the monitoring data of type A, corr A represents the relationship value between the monitoring data of type A and the prediction result, N represents the number of the monitoring data of type A in su, value i represents the value of the monitoring data i of type A in the set su, and su represents the set of monitoring data obtained within a preset time range; value maxrepresents the maximum value of the monitoring data of type A in su, ds represents the preset acquisition interval, and corr max represents the preset relationship value comparison coefficient, and λ is the data weight.

[0012] Preferably, the preset time range is [ts - T, ts], where ts is the moment when the fault prediction module last started fault prediction, and T is the preset first duration.

[0013] Preferably, the preset first duration is 1 hour.

[0014] Preferably, the preset acquisition interval is 1 minute, and the preset relationship value comparison coefficient is 1.

[0015] Preferably, it further includes a model training module, which is used to train the deep learning model to obtain a deep learning model for fault prediction.

[0016] Preferably, based on the monitoring data obtained by the data collection module and deep learning technology, fault prediction is performed on the server to obtain a prediction result, including:

[0017] Obtain the monitoring data for prediction;

[0018] Preprocess the monitoring data for prediction to obtain the preprocessed monitoring data;

[0019] Input the preprocessed monitoring data into the deep learning model for fault prediction to obtain a prediction result.

[0020] Preferably, obtaining the monitoring data for prediction includes:

[0021] Use tp to represent the moment when starting to obtain the monitoring data for prediction, and use the monitoring data of the preset type whose acquisition moment falls within the time range [tp - Tp, tp] as the monitoring data for prediction, where Tp represents the preset second duration.

[0022] Preferably, preprocessing the monitoring data for prediction to obtain the preprocessed monitoring data includes:

[0023] Perform anomaly detection processing on the monitoring data for prediction of each type respectively to obtain the preprocessed monitoring data of each type.

[0024] Preferably, the prediction result includes the existence of a fault or the non - existence of a fault.

[0025] Preferably, the deep learning model includes an LSTM model.

[0026] Beneficial effects:

[0027] In the process of obtaining various types of monitoring data for server fault prediction, the present invention does not directly use a fixed collection interval to obtain monitoring data. Instead, it obtains the collection interval based on the relationship value and the stability of the numerical amplitude of the monitoring data. In this way, a larger collection interval can be used when the influence of the monitoring data on the prediction result is smaller, and a larger collection interval can also be used when the stability of the monitoring data is greater, thereby reducing the amount of data obtained and effectively saving storage space. When the influence of the monitoring data on the prediction result is greater and the stability of the monitoring data is smaller, a smaller collection interval is used to achieve the acquisition of more highly accurate data and improve the accuracy of the prediction result. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0029] Figure 1 It is a schematic diagram of a server fault prediction system based on deep learning according to the present invention.

[0030] Figure 2 It is another schematic diagram of a server fault prediction system based on deep learning according to the present invention.

[0031] Figure 3 It is a schematic diagram of the process of fault prediction for a server according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0033] The present invention provides a server fault prediction system based on deep learning, including a data collection module and a fault prediction module;

[0034] The data collection module is used to obtain preset types of monitoring data during the operation of the server based on the collection interval of the monitoring data;

[0035] The fault prediction module is used to perform fault prediction on the server based on the monitoring data obtained by the data collection module and deep learning technology to obtain a prediction result;

[0036] The data collection module is also used to calculate the collection interval of each type of monitoring data according to the prediction result, including:

[0037] For the monitoring data of type A, the calculation formula for its collection interval is:

[0038]

[0039] d A represents the collection interval of the monitoring data of type A, corr A represents the relationship value between the monitoring data of type A and the prediction result, N represents the number of the monitoring data of type A in su, value i represents the value of the monitoring data of type A, i, in the set su, and su represents the set of monitoring data obtained within a preset time range; value max represents the maximum value of the monitoring data of type A in su, ds represents the preset collection interval, corr max represents the preset relationship value comparison coefficient, and λ is the data weight.

[0040] The collection interval of the present invention is not a fixed value, and can change with the change of the relationship value and the stability degree of the data. Under the condition that other conditions remain unchanged, if the value of the relationship value is larger, the collection interval is smaller; under the condition that other conditions remain unchanged, if the stability degree of the data is lower, that is, the calculated variance is larger, the collection interval is smaller. Therefore, the collection interval of the present invention can achieve a good balance between the relationship value and the collection interval, so as to ensure the accuracy of the obtained monitoring data while reducing the data volume of the obtained monitoring data, that is, the monitoring data is collected at a smaller collection interval.

[0041] Preferably, the preset type of monitoring data includes CPU usage rate, memory usage rate, disk write volume, disk read volume, network upload volume, network download volume, cumulative running time, temperature and humidity of the environment, etc.

[0042] Preferably, the preset time range is [ts - T, ts], where ts is the moment when the fault prediction module starts the fault prediction for the last time, and T is the preset first duration.

[0043] Here, the first duration is a relatively small value, mainly to analyze the stability of the numerical change by analyzing the monitoring data within a certain time range.

[0044] Preferably, the preset first duration is 1 hour.

[0045] Specifically, the preset first duration can also be other values, which can be adjusted according to requirements, but should be controlled within a relatively appropriate value, so as to more accurately reflect the stability of the numerical changes of the monitoring data in the recent period.

[0046] Preferably, the preset acquisition interval is 1 minute, and the preset relational value comparison coefficient is 1.

[0047] Specifically, according to the type of monitoring data, corresponding acquisition intervals can be set for each type of monitoring data. For example, for data such as CPU, memory, disk, and network traffic, the preset acquisition interval can be set smaller, such as 1 minute; while for environmental temperature, humidity, etc., the preset acquisition intervals for these data can be set larger, such as 10 minutes.

[0048] Preferably, the data weight is

[0049] Preferably, as Figure 2 , it further includes a model training module, and the model training module is used to train the deep learning model to obtain a deep learning model for fault prediction.

[0050] In the present invention, training the deep learning model to obtain a deep learning model for fault prediction includes:

[0051] The first step is to perform data cleaning on the monitoring data for training, and divide the data obtained after data cleaning into a training set, a validation set, and a test set;

[0052] The second step is to construct a deep learning model. For example, for an LSTM model, the construction process includes:

[0053] Constructing an LSTM model includes an input layer, an LSTM layer, a fully connected layer, and an output layer. The LSTM layer is responsible for capturing long-term dependencies in time series data.

[0054] Hyperparameter setting: Determine hyperparameters such as input dimension, hidden layer size, output dimension, and number of LSTM layers. For example, the hidden layer size and number of layers need to be adjusted according to the task complexity and data volume.

[0055] The input dimension is the same as the total number of types of monitoring data. The hidden layer size and the number of LSTM layers can be gradually adjusted starting from a small value to find the optimal value, and the output dimension is related to the number of output results. For example, if only judging whether a fault will occur, the output dimension is 1;

[0056] In the third step, determine the loss function and optimizer, and select a suitable loss function (such as MSELoss or SmoothL1Loss) and optimizer (such as Adam or RMSprop);

[0057] In the fourth step, input the data of the training set into the LSTM model for forward propagation;

[0058] Calculate the loss value and update the model parameters through backpropagation;

[0059] Use gradient clipping to prevent gradient explosion.

[0060] Evaluate the model performance on the validation set every certain number of rounds (such as 20 rounds) and save the model. During the training process, if the loss on the validation set does not improve within a certain number of rounds (e.g., 5 rounds), stop the training early to prevent overfitting.

[0061] Preferably, as Figure 3 , perform fault prediction on the server based on the monitoring data obtained by the data collection module and deep learning technology, and obtain prediction results, including:

[0062] Obtain the monitoring data for prediction;

[0063] Preprocess the monitoring data for prediction to obtain preprocessed monitoring data;

[0064] Input the preprocessed monitoring data into the deep learning model for fault prediction to obtain prediction results.

[0065] By preprocessing the monitoring data, the quality of the data input into the model can be improved, and the accuracy of the prediction results can be enhanced.

[0066] Preferably, obtaining the monitoring data for prediction includes:

[0067] Let tp represent the moment when starting to obtain the monitoring data for prediction, and use the monitoring data of the preset type whose collection moment falls within the time range [tp - Tp, tp] as the monitoring data for prediction, where Tp represents the preset second duration.

[0068] The collection moment refers to the moment when the data collection module obtains the monitoring data

[0069] The preset second duration of the present invention can be a relatively large value, such as 1 day.

[0070] Preferably, preprocessing the monitoring data for prediction to obtain preprocessed monitoring data includes:

[0071] Anomaly detection processing is performed on each type of monitoring data for prediction to obtain preprocessed monitoring data of each type.

[0072] By performing anomaly detection processing, abnormal monitoring data can be repaired to improve the data quality.

[0073] Preferably, the process of performing outlier detection processing on the monitoring data for prediction includes:

[0074] Let M represent the total number of monitoring data for prediction. These M monitoring data are numbered in the order of acquisition time from early to late. The range of the numbering is [1, M], and the interval between the numbers is 1.

[0075] For the m-th monitoring data, calculate its detection value:

[0076]

[0077] difc m represents the detection value of the m-th monitoring data. um represents the set of the first K monitoring data whose acquisition time is before the acquisition time of the m-th monitoring data and is the closest to the acquisition time of the m-th monitoring data. K is a set positive integer. val i represents the value of the monitoring data i in um, w i represents the weight of the monitoring data i, val max represents the maximum value of the monitoring data in um, valpre m represents the value obtained by prediction based on the data in um. λ1 and λ2 are the cumulative weight and the deviation weight respectively; m ∈ [1, M]; val m represents the value of the m-th monitoring data;

[0078] Judge whether the detection value is greater than the anomaly comparison value. If so, it means that the m-th monitoring data is abnormal, and a preset correction algorithm is used to obtain the preprocessed monitoring data.

[0079] In the process of calculating the detection value, the present invention does not directly calculate based on the weighted result of the m-th monitoring data and the surrounding data, but also introduces a predicted value, and comprehensively calculates the detection value through the gap between the weighted result and the predicted value. In this way, when the m-th monitoring data is abnormal, a detection value with a larger value can be calculated in the first time, improving the recognition efficiency. At the same time, the introduction of the weighted value can reduce the deviation degree of the detection value, making the detection value more in line with the law of data distribution.

[0080] Preferably, the calculation formula of wi is:

[0081]

[0082] The weight of the monitoring data in the present invention is calculated based on the numerical difference between the m-th monitoring data and the detection data i. The greater the numerical difference, the smaller the weight, and the smaller the impact on the calculation result of the detection value.

[0083] Preferably, valpre m is obtained as follows:

[0084] Input the monitoring data in um into the trained ARIMA-LSTM hybrid model for prediction to obtain valpre m .

[0085] When detecting outliers in the monitoring data of the present invention, a threshold with a fixed value is not used for detection. Since the monitoring data is constantly changing, if a fixed threshold is set, it cannot be dynamically adjusted according to the change in the data distribution. If the statistical characteristics of the data (such as mean, variance) change, the fixed threshold may lead to misjudgment or missed judgment, and it cannot be dynamically adjusted according to the change in the data distribution. If the statistical characteristics of the data (such as mean, variance) change, the fixed threshold may lead to misjudgment or missed judgment. Therefore, calculating the threshold based on the recently obtained monitoring data is beneficial to obtaining a better detection threshold.

[0086] The ARIMA-LSTM hybrid model is a method that combines traditional time series analysis and deep learning, aiming to make full use of the ability of the ARIMA model to capture linear trends and the ability of LSTM to model non-linear patterns, thereby improving the accuracy of time series prediction.

[0087] The specific implementation steps of the hybrid model are roughly as follows:

[0088] Use the ARIMA model to fit the time series to obtain the predicted value of the linear part.

[0089] Subtract the predicted value of ARIMA from the original time series to obtain the residual series.

[0090] Use the LSTM model to model the residual series to capture non-linear patterns not captured by ARIMA.

[0091] Add the predicted value of ARIMA and the predicted value of LSTM to obtain the final prediction result.

[0092] Preferably, the abnormal comparison value is one-fifth of the mean of the monitoring data in um.

[0093] Preferably, a preset correction algorithm is used to obtain the preprocessed monitoring data, including:

[0094] The m-th monitoring data is calculated using the following formula:

[0095]

[0096] fval m represents the preprocessed monitoring data obtained after calculating val m od i represents the number of the monitoring data i in um.

[0097] The calibration algorithm obtains the preprocessed monitoring data by weighting the monitoring data in um. When the number of the monitoring data in um is smaller, its influence on the calculation result is correspondingly smaller. Therefore, it is possible to focus more on using other monitoring data with smaller differences in acquisition times to calibrate the m-th monitoring data.

[0098] um is a subset of M monitoring data for prediction. Therefore, each monitoring data in um has a corresponding number.

[0099] In another embodiment, obtaining the preprocessed monitoring data using a preset calibration algorithm includes:

[0100] Taking the average value of the first 5 monitoring data with the closest acquisition time to the m-th monitoring data as the preprocessed monitoring data.

[0101] The set positive integer of the present invention can be 20.

[0102] The cumulative weight and deviation weight of the present invention can be 0.7 and 0.3 respectively.

[0103] Preferably, the prediction result includes the presence of a fault or the absence of a fault.

[0104] Preferably, the process of obtaining the relationship value between the monitoring data of type A and the prediction result includes:

[0105] Using 0 to represent the absence of a fault and 1 to represent the presence of a fault;

[0106] Normalizing the monitoring data of type A within the time range [ts - Ts, ts] of the acquisition time to obtain a data sequence C; Ts represents a preset third time duration;

[0107] Taking the prediction result within the time range [ts - Ts, ts] of the output time of the prediction result as the data sequence B;

[0108] Calculating the cross-correlation coefficient between the data sequences C and B, and taking the cross-correlation coefficient as the relationship value.

[0109] The preset third time duration of the invention can be 1 week. The third time duration needs to be set relatively long so that the cross-correlation coefficient can be obtained through a relatively large number of data to determine the relationship value.

[0110] Preferably, the deep learning model includes an LSTM model.

[0111] The deep learning model of the present invention may also include an RNN model.

[0112] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can understand and utilize the present invention well. The present invention is only limited by the claims and their full scope and equivalents.

Claims

1. A server fault prediction system based on deep learning, characterized in that It includes a data collection module and a fault prediction module; The data collection module is used to obtain the monitoring data of a preset type in the running process of the server based on the collection interval of the monitoring data; The fault prediction module is used to perform fault prediction on the server based on the monitoring data obtained by the data collection module and deep learning technology to obtain a prediction result; The data collection module is also used to calculate the collection interval of each type of monitoring data according to the prediction result, including: For the monitoring data of type A, the calculation formula of its collection interval is: d A represents the acquisition interval of the monitoring data of type A, corr A represents the relationship value between the monitoring data of type A and the prediction result, N represents the number of the monitoring data of type A in su, value i represents the value of the monitoring data i of type A in the set su, su represents the set of the monitoring data obtained within a preset time range; value max represents the maximum value of the monitoring data of type A in su, ds represents the preset acquisition interval, corr max represents the preset relationship value comparison coefficient, and λ is the data weight.

2. The server fault prediction system based on deep learning according to claim 1, wherein, The preset time range is [ts - T, ts], where ts is the moment when the fault prediction module starts the most recent fault prediction, and T is the preset first duration.

3. The server fault prediction system based on deep learning according to claim 2, characterized in that, The preset first duration is 1 hour.

4. A server fault prediction system based on deep learning according to claim 1, characterized in that, The preset collection interval is 1 minute, and the preset relationship value comparison coefficient is 1.

5. The server fault prediction system based on deep learning according to claim 1, characterized in that, It also includes a model training module, which is used to train the deep learning model to obtain a deep learning model for fault prediction.

6. The server fault prediction system based on deep learning according to claim 1, characterized in that Performing fault prediction on the server based on the monitoring data obtained by the data collection module and deep learning technology to obtain a prediction result, including: Obtaining the monitoring data for prediction; Preprocessing the monitoring data for prediction to obtain the preprocessed monitoring data; Inputting the preprocessed monitoring data into the deep learning model for fault prediction to obtain a prediction result.

7. The server fault prediction system based on deep learning according to claim 6, characterized in that, Obtaining the monitoring data for prediction, including: Using tp to represent the moment when the monitoring data for prediction starts to be obtained, and taking the monitoring data of the preset type whose collection moment falls within the time range [tp - Tp, tp] as the monitoring data for prediction, where Tp represents the preset second duration.

8. The server fault prediction system based on deep learning according to claim 6, wherein, Preprocessing the monitoring data for prediction to obtain the preprocessed monitoring data, including: Performing anomaly detection processing on each type of monitoring data for prediction respectively to obtain the preprocessed monitoring data of each type.

9. The server fault prediction system based on deep learning according to claim 6, characterized in that, The prediction result includes the existence of a fault or the non - existence of a fault.

10. A server fault prediction system based on deep learning according to claim 5, characterized in that, The deep learning model includes an LSTM model.