Fault prediction method and device of computer equipment, server and storage medium
By automatically collecting and preprocessing the operation and maintenance data of computer equipment, extracting features and inputting the fault prediction model, the problem of low timeliness of computer equipment failure prediction is solved, and efficient and accurate fault prediction is achieved.
Patent Information
- Application Number
- CN202510866991.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, computer equipment failure prediction is low in time, and relying on manual monitoring and analysis leads to troubleshooting delays.
By automatically collecting operation and maintenance data of computer equipment, pre-processing, extracting periodic and health status characteristics, and inputting them to the fault prediction model, outputting fault categories and abnormal modes.
It improves the efficiency and accuracy of fault prediction, reduces the workload of manual monitoring and analysis, reduces subjective bias and errors in human judgment, and improves the timeliness of fault prediction.
Smart Images

Figure CN120372459A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of device management, and particularly to a method, apparatus, server, and storage medium for predicting faults of computer devices. Background Art
[0002] In the context of the big data era, the demand for data detection, processing, and analysis has increased sharply, and the requirements for the stability of computer devices for data processing in various industries have been continuously improved. Therefore, it is particularly important to predict faults of computer devices.
[0003] Currently, related technologies usually rely on manual operations, and fault prediction is carried out by manually monitoring and analyzing the data generated by computer devices. However, the manual processing speed lags behind the data generation speed, which may lead to delays in fault troubleshooting and reduce the timeliness of fault prediction. Summary of the Invention
[0004] This application provides a method, apparatus, server, and storage medium for predicting faults of computer devices to at least solve the problem of low timeliness of fault prediction in related technologies.
[0005] This application provides a method for predicting faults of computer devices, including:
[0006] Collecting operation and maintenance data of computer devices.
[0007] Preprocessing the operation and maintenance data to obtain preprocessed operation and maintenance data.
[0008] Extracting periodic features and health status features from the preprocessed operation and maintenance data.
[0009] Inputting the periodic features and health status features into a fault prediction model, where the fault prediction model includes a fault category identification module and an abnormal pattern identification module, and the fault prediction model is used to perform the following steps:
[0010] Outputting a fault category through the fault category identification module according to the periodic features and health status features;
[0011] Outputting an abnormal pattern of the fault through the abnormal pattern identification module according to the periodic features and health status features.
[0012] This application also provides a device for predicting faults of computer devices, including:
[0013] A collection module, configured to collect operation and maintenance data of computer devices.
[0014] A preprocessing module, configured to preprocess the operation and maintenance data to obtain preprocessed operation and maintenance data.
[0015] An extraction module, configured to extract periodic features and health status features from the preprocessed operation and maintenance data.
[0016] An execution module, configured to input the periodic features and health status features into a fault prediction model, where the fault prediction model includes a fault category recognition module and an abnormal pattern recognition module, and the fault prediction model is used to perform the following steps:
[0017] The execution module includes:
[0018] A first output unit, configured to output a fault category through the fault category recognition module according to the periodic features and health status features.
[0019] A second output unit, configured to output an abnormal pattern of the fault through the abnormal pattern recognition module according to the periodic features and health status features.
[0020] This application also provides a server, including: a memory, configured to store a computer program; a processor, configured to implement the steps of the fault prediction method of any one of the above computer devices when executing the computer program.
[0021] This application also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, the steps of the fault prediction method of any one of the above computer devices are implemented.
[0022] This application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the fault prediction method of any one of the above computer devices are implemented.
[0023] Through this application, by automatically collecting the operation and maintenance data of computer devices and preprocessing the operation and maintenance data of computer devices, the operation and maintenance data is made convenient for subsequent processing, reducing the workload of manual monitoring and analysis; from the preprocessed operation and maintenance data, periodic features and health status features are extracted. The periodic features and health status features are input into the fault prediction model, and the fault prediction model automatically outputs the fault category and the abnormal pattern of the fault, improving the efficiency of fault prediction, thereby improving the timeliness of fault prediction. In addition, the fault prediction model can objectively analyze the periodic features and health status features, reducing the subjective bias and errors of human judgment and improving the accuracy of fault prediction. In addition, converting unstructured data into structured data facilitates feature extraction and model prediction. Structured data reduces noise and redundant information, and the features extracted based on structured data can make the fault prediction model focus more on key information, avoiding prediction deviations caused by chaotic data formats. Description of the Drawings
[0024] To more clearly illustrate the embodiments of the present application, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0025] Figure 1 Schematic diagram of the scenario of the fault prediction method for the computer device provided by the embodiment of the present application;
[0026] Figure 2 Schematic flow diagram of the fault prediction method for the computer device provided by the embodiment of the present application;
[0027] Figure 3 Schematic structural diagram of the fault prediction device for the computer device provided by the embodiment of the present application;
[0028] Figure 4 Schematic structural diagram of the server provided by the embodiment of the present application. Detailed implementation manners
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present application.
[0030] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0031] In the context of the big data era, the demand for data detection, processing, and analysis has increased sharply, and the requirements for the stability of computer devices for data processing in various industries have been continuously improved. Therefore, it is particularly important to predict faults in computer devices. Currently, related technologies usually rely on manual operations, and fault prediction is carried out by manually monitoring and analyzing the data generated by computer devices. However, the manual processing speed lags behind the data generation speed, which may lead to delays in fault troubleshooting and reduce the timeliness of fault prediction.
[0032] To solve the technical problems in the related art, the present application proposes the following technical concept: Considering that fault prediction by manually monitoring and analyzing the data generated by computer devices will reduce the timeliness of fault prediction. The inventor thought of automatically collecting the operation and maintenance data of computer devices and automatically preprocessing the operation and maintenance data, reducing the workload of manual monitoring and analysis; extracting periodic features and health status features from the preprocessed operation and maintenance data. Inputting the periodic features and health status features into the fault prediction model, and the fault prediction model automatically outputs the fault category and the abnormal mode of the fault, improving the efficiency of fault prediction and thus improving the timeliness of fault prediction.
[0033] To enable those skilled in the art of the present technology to better understand the solution of the present application, the following further detailed description of the present application will be given in conjunction with the accompanying drawings and specific embodiments.
[0034] Combined with the specific application environment architecture or specific hardware architecture on which the execution of the fault prediction method of computer devices depends, the specific application environment architecture or specific hardware architecture will be described herein.
[0035] Reference Figure 1 , Figure 1 is a schematic diagram of the scenario of the fault prediction method of the computer device provided by the embodiment of the present application. As Figure 1 shown, it includes: a receiving device 101, a processing device 102, and a display device 103.
[0036] It can be understood that the structure schematically shown in the embodiment of the present application does not constitute a specific limitation on the fault prediction method of computer devices. In other feasible embodiments of the present application, the above architecture may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements, which can be specifically determined according to the actual application scenario and will not be limited herein. Figure 1 The components shown can be implemented in hardware, software, or a combination of software and hardware.
[0037] In the specific implementation process, the receiving device 101 can be an input / output interface or a communication interface, and can receive the operation and maintenance data of the collected computer device.
[0038] The processing device 102 can perform a series of processes on the operation and maintenance data of the computer device to obtain the fault category and the abnormal mode of the fault.
[0039] The display device 103 can be used to display the fault category and the abnormal mode of the fault.
[0040] In addition, the network architecture and business scenarios described in the embodiments of this application are used to more clearly illustrate the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those of ordinary skill in the art will know that with the evolution of the network architecture and the emergence of new business scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0041] Figure 2 is a schematic flowchart of a method for predicting faults of a computer device provided by an embodiment of this application. As Figure 2 shown, an embodiment of this application provides a method for predicting faults of a computer device, and the method will be described in detail as follows:
[0042] S201: Collect operation and maintenance data of the computer device.
[0043] Optionally, the operation and maintenance data is collected in real time through methods such as an Application Programming Interface (API), a log collection tool, and a Simple Network Management Protocol (SNMP). The operation and maintenance data of the computer device includes CPU usage rate, memory usage rate, disk I / O, network traffic, error logs, etc. collected from servers, network devices, databases, and application logs. Among them, the API interface is provided by an application or a data storage system.
[0044] Optionally, the log collection tool can be the ELK Stack, and the ELK Stack includes Elasticsearch, Logstash, and Kibana. Specifically, by configuring the collection path of Filebeat, all log files under the collection path are collected, and the collected log data is sent to Logstash; Logstash processes the log data and saves it to Elasticsearch after the processing is completed. Optionally, the log data can be saved in JSON format.
[0045] Optionally, the log data can also be visualized and analyzed through Kibana.
[0046] Optionally, by running an SNMP agent, such as the snmpd service of Linux, the snmpwalk or snmpget command is used to obtain the operation and maintenance data.
[0047] Exemplarily, the snmpwalk or snmpget command is used to obtain the operation and maintenance data as follows:
[0048] Collect system information: snmpwalk -v 2c -c public localhost .1.3.6.1.2.1.1;
[0049] Collect CPU load: snmpwalk -v 2c -c public localhost.1.3.6.1.4.1.2021.10.1.3;
[0050] Collect memory usage: snmpwalk -v 2c -c public localhost.1.3.6.1.4.1.2021.4;
[0051] Collect disk usage: snmpwalk -v 2c -c public localhost.1.3.6.1.4.1.2021.9;
[0052] Get system description: snmpget -v 2c -c public localhost .1.3.6.1.2.1.1.1.0;
[0053] Get system startup time: snmpget -v 2c -c public localhost.1.3.6.1.2.1.1.3.0;
[0054] Get CPU load: snmpget -v 2c -c public localhost.1.3.6.1.4.1.2021.10.1.3.1.
[0055] S202: Preprocess the operation and maintenance data to obtain the preprocessed operation and maintenance data.
[0056] Specifically, step S202 includes S2021~S2024:
[0057] S2021: Delete the duplicate data in the operation and maintenance data to obtain the first operation and maintenance data.
[0058] S2022: Fill in the missing values in the first operation and maintenance data to obtain the second operation and maintenance data.
[0059] In this embodiment, obtain the missing values in the first operation and maintenance data, and fill in the missing values by the time series interpolation method.
[0060] Optionally, forward filling can be used, filling with the previous valid data of the missing value, which is suitable for scenarios where the data trend is relatively stable; backward filling can also be used, filling with the next valid data of the missing value, which is suitable for data scenarios with hysteresis.
[0061] Exemplarily, Table 1 shows an example of filling in the missing values of CPU usage provided by the embodiments of the present application. As shown in Table 1, the CPU usage for 10 dates was collected, and there are 3 missing values among them. For the missing value of CPU usage with serial number 2, the forward fill is 22.0 and the backward fill is 25.0; for the missing value of CPU usage with serial number 4, the forward fill is 25.0 and the backward fill is 30.0; for the missing value of CPU usage with serial number 7, the forward fill is 32.0 and the backward fill is 35.0.
[0062] Table 1 Example of Filling in Missing Values of CPU Usage
[0063]
[0064] S2023: Convert the unstructured data in the second operation and maintenance data into structured data to obtain the third operation and maintenance data.
[0065] Specifically, step S2023 includes Sa~Se:
[0066] Sa: Obtain the unstructured data in the second operation and maintenance data.
[0067] Exemplarily, the collected unstructured log data is as follows:
[0068] 2025-02-18T12:00:00Z INFO System started successfully.
[0069] 2025-02-18T12:05:00Z ERROR Disk usage on / dev / sda1 exceeded 95%.
[0070] 2025-02-18T12:10:00Z WARNING CPU load average is high: 4.5(threshold: 2.0).
[0071] 2025-02-18T12:15:00Z INFO Backup completed successfully.
[0072] Sb: Split the unstructured data to obtain multiple text data blocks.
[0073] Exemplarily, the unstructured log data is split by line, and each piece of unstructured log data is used as an independent text data block.
[0074] Sc: For each text data block, use regular expressions to extract key fields from each text data block.
[0075] Optionally, the key fields can be timestamp, log level, message content, etc.
[0076] Exemplarily, use regular expressions to extract timestamp, log level, message content, etc. from each unstructured log data.
[0077] Sd: For each text data block, use named entity recognition technology to extract key entities from each text data block.
[0078] Optionally, the key entities can be disk path, numerical metrics, device name, etc.
[0079] Exemplarily, use named entity recognition technology to extract disk path, numerical metrics, device name, etc. from each unstructured log data.
[0080] Exemplarily, the extracted disk path is / dev / sda1, etc., the numerical metrics are 95% and 4.5, etc., and the device names are CPU and Disk, etc.
[0081] Se: Package the key fields and key entities into structured data to obtain the third operation and maintenance data.
[0082] Optionally, package the extracted key fields and key entities into structured data, such as JSON format and CSV format, etc.
[0083] S2024: Normalize the third operation and maintenance data to obtain the preprocessed operation and maintenance data.
[0084] In this embodiment, the numerical ranges of different operation and maintenance data vary greatly. For example, the CPU usage rate is 0 - 100%, and the memory usage is 0 - 16GB. If directly input into the fault prediction model, large-scale features will dominate the loss function, resulting in the neglect of small-scale features. Normalization processing is required to scale them to the same scale.
[0085] Optionally, common normalization methods include: min-max normalization method and standardization method, etc.
[0086] S203: Extract periodic features and health status features from the preprocessed operation and maintenance data.
[0087] Specifically, step S203 includes S2031~S2032:
[0088] S2031: Input the preprocessed operation and maintenance data into a convolutional neural network, so that the convolutional layer of the convolutional neural network extracts periodic features from the preprocessed operation and maintenance data.
[0089] In this embodiment, the structure of the convolutional neural network includes a convolutional layer, a pooling layer, a dimensionality reduction function, and a fully connected layer. The parameters of each layer are set to complete the construction of the convolutional neural network. The convolutional neural network is compiled and trained to obtain a trained convolutional neural network for extracting periodic features.
[0090] In this embodiment, after the preprocessed operation and maintenance data is input into the convolutional neural network, the convolutional neural network first normalizes the data to ensure again that the data processed by the convolutional neural network is of the same scale. The data shape of the preprocessed operation and maintenance data is adjusted. The data shape is: the number of samples, the time step, and the number of features. Exemplarily, the time step can be set to 1440 and the number of features is 1.
[0091] In this embodiment, periodic features are output through the convolutional layer of the convolutional neural network.
[0092] Exemplarily, the time series data of the CPU usage rate is input into the convolutional neural network, so that the convolutional layer of the convolutional neural network extracts periodic features from the time series data of the CPU usage rate, such as the peak periods daily or weekly.
[0093] S2032: Input the preprocessed operation and maintenance data into the recurrent neural network, so that the recurrent neural network extracts health status features from the preprocessed operation and maintenance data.
[0094] Optionally, the recurrent neural network can be a Long Short-Term Memory Networks (LSTM), and through the LSTM, health status features reflecting the health status are extracted.
[0095] In this embodiment, the construction and training process of the long short-term memory network is: import the core library, preprocess the input data, divide the training set and the test set, construct the LSTM model, compile the model, and train the model.
[0096] In this embodiment, the preprocessing includes: converting the timestamp to a time series index, normalizing the data, and creating a time series window. The timestamp is converted to the datetime type and set as the index to facilitate processing the data in chronological order; the data is normalized to eliminate the scale difference; a time series window is created to divide the data into sliding windows of a fixed length, and each window contains the data of the first 30 time points for predicting the data of the 31st time point.
[0097] In this embodiment, the first layer of the LSTM includes 64 memory units, retaining the hidden states of all time steps. The second layer of the LSTM includes 32 memory units, retaining only the hidden state of the last time step. 20% of the neuron outputs are randomly dropped to prevent overfitting. The output and the input feature dimensions are set to be the same.
[0098] In this embodiment, after the LSTM is trained using the training set, the test set is used for prediction, and the error between the predicted value and the actual value is calculated. The error is minimized by the Adam optimizer. A threshold is set. When the prediction error exceeds the threshold, it may be the time point of abnormal health status. The sensitivity of anomaly detection can be controlled by adjusting the threshold.
[0099] Optionally, the error of the test set can be visualized.
[0100] S204: Input the periodic features and health status features into the fault prediction model, where the fault prediction model includes a fault category recognition module and an abnormal pattern recognition module, and the fault prediction model is used to perform the following steps.
[0101] Specifically, step S204 includes S2041~S2042:
[0102] S2041: Through the fault category recognition module, according to the periodic features and health status features, output the fault category.
[0103] In this embodiment, the fault category recognition module is a fault category recognition model, and the fault category recognition model predicts the fault category through the random forest algorithm, where the fault categories include having a fault and no fault. In subsequent embodiments, the training process of the fault category recognition model is introduced.
[0104] S2042: Through the abnormal pattern recognition module, according to the periodic features and health status features, output the abnormal pattern of the fault. The abnormal patterns of the fault include excessive CPU usage, excessive memory usage, and disk I / O anomalies, etc.
[0105] In this embodiment, the abnormal pattern recognition module is an abnormal pattern recognition model, and the abnormal pattern recognition model predicts the abnormal pattern of the fault through the clustering algorithm. In subsequent embodiments, the training process of the abnormal pattern recognition model is introduced.
[0106] In summary, by automatically collecting the operation and maintenance data of computer devices and preprocessing the operation and maintenance data of computer devices, the operation and maintenance data is made convenient for subsequent processing, reducing the workload of manual monitoring and analysis; from the preprocessed operation and maintenance data, periodic features and health status features are extracted. The periodic features and health status features are input into the fault prediction model, and the fault prediction model automatically outputs the fault category and the abnormal pattern of the fault, improving the efficiency of fault prediction, thereby improving the timeliness of fault prediction. In addition, the fault prediction model can objectively analyze the periodic features and health status features, reducing the subjective bias and errors of human judgment and improving the accuracy of fault prediction. In addition, converting unstructured data into structured data facilitates feature extraction and model prediction. Structured data reduces noise and redundant information, and features extracted based on structured data can make the fault prediction model focus more on key information, avoiding prediction deviations caused by chaotic data formats.
[0107] Based on the above embodiments, in this embodiment, the training process of the fault prediction model is introduced, where the fault category recognition module is the fault category recognition model; the abnormal pattern recognition module is the abnormal pattern recognition model, which is described in detail as follows:
[0108] S301: Obtain the historical fault data of the computer device; where the historical fault data includes historical fault labels.
[0109] In this embodiment, the historical fault labels include 1 and 0, where 1 represents a fault and 0 represents normal.
[0110] S302: Obtain the historical periodic features and historical health status features in the historical fault data.
[0111] In this embodiment, the process of obtaining the historical periodic features and historical health status features will not be elaborated again.
[0112] S303: Divide the historical periodic features and historical health status features into a training set and a test set.
[0113] Optionally, the division ratio of the training set and the test set can be set to 4:1.
[0114] S304: Adopt a supervised learning algorithm, and according to the training set and the historical fault labels, recursively construct multiple decision trees according to the preset splitting decision to obtain an initial fault category recognition model.
[0115] In this embodiment, a supervised learning algorithm is adopted to label the training set according to historical fault labels, and multiple decision trees are recursively constructed according to a preset splitting decision to obtain an initial fault category recognition model. For each decision tree, a subset is randomly selected from the training set, and the training data sets of each decision tree are different, enhancing the diversity of the fault category recognition model; only part of the features are considered during the splitting of each node to avoid overfitting of a single decision tree. Multiple decision trees are recursively constructed according to a preset splitting decision to obtain an initial fault category recognition model.
[0116] Optionally, 100 decision trees can be selected for construction, and the maximum depth of a single decision tree is 10.
[0117] Optionally, the preset splitting decisions include information gain and Gini impurity. Information gain is used to select splitting features, and Gini impurity is used to measure the impurity of data.
[0118] Among them, the formula for information gain is:
[0119]
[0120] In the formula, represents the information gain, represents the training data set before splitting, represents the entropy of the training data set before splitting; represents the feature considered during splitting, represents all possible values of the feature considered during splitting, represents that the considered feature value is the corresponding training data set, represents that the considered feature value is the entropy of the corresponding training data set.
[0121] Among them, the formula for Gini impurity is:
[0122]
[0123] In the formula, represents the Gini impurity, represents the number of different feature categories in the training data set during splitting, represents the category the proportion in the training data set.
[0124] Among them, the periodic features and health status features include different features, such as CPU usage rate, memory usage rate, and disk I / O, etc.
[0125] S305: Use the training set to train the initial fault category recognition model to obtain a trained fault category recognition model.
[0126] S306: Use the test set to optimize the trained fault category recognition model to obtain the final fault category recognition model.
[0127] Specifically, input the test set into the trained fault category recognition model, so that the trained fault category recognition model outputs the predicted fault category for the test set; evaluate the trained fault category recognition model according to the true fault category and the predicted fault category of the test set to obtain an evaluation result; according to the evaluation result, use the backpropagation algorithm to adjust the model parameters of the trained fault category recognition model to optimize the trained fault category recognition model and obtain the final fault category recognition model.
[0128] Optionally, the performance of the trained fault category recognition model can be improved by selecting different algorithms, such as switching the random forest algorithm to the decision tree algorithm.
[0129] In this embodiment, the fault category recognition model can be evaluated by analyzing the accuracy rate, precision rate, recall rate, and F1 score, etc., or by displaying the comparison between the true fault category and the predicted fault category of the test set through a confusion matrix.
[0130] In this embodiment, the final output result is determined by a voting mechanism, and the formula is as follows:
[0131]
[0132] In the formula, represents the output result, that is, the fault category, represents the prediction result of the first decision tree, represents the prediction result of the second decision tree, represents the prediction result of the kth decision tree, is a number of functions, and the fault category with the most occurrences is selected.
[0133] S307: Adopt an unsupervised learning algorithm, set the neighborhood radius and the minimum number of samples in the neighborhood, and construct an initial abnormal pattern recognition model.
[0134] In this embodiment, the abnormal pattern recognition model identifies the abnormal pattern of the fault through a clustering algorithm. Such as the DBSCAN algorithm. The principle is to mark the points with insufficient density as abnormal points in the feature space.
[0135] In this embodiment, the neighborhood radius defines the distance threshold for two points to be adjacent; the points below the minimum number of samples in the neighborhood are regarded as noise, that is, abnormal points.
[0136] In this embodiment, there may be problems with the marked abnormal points, and the abnormal patterns include excessive CPU usage, excessive memory usage, disk I / O anomalies, etc.
[0137] S308: Use the training set to train the initial abnormal pattern recognition model to obtain a trained abnormal pattern recognition model.
[0138] S309: Use the test set to optimize the trained abnormal pattern recognition model to obtain the final abnormal pattern recognition model.
[0139] In this embodiment, the model parameters of the abnormal pattern recognition model are continuously adjusted through the backpropagation algorithm to optimize the prediction accuracy of the abnormal pattern recognition model.
[0140] Optionally, LSTM can be used to predict the load trend of the server, and cross-validation can be used to evaluate the performance of the abnormal pattern recognition model.
[0141] In summary, the supervised learning method is used to train the fault category recognition model so that the fault category recognition model can predict known faults; the unsupervised learning method is used to train the abnormal pattern recognition model so that the abnormal pattern recognition model can predict unknown faults. The combination of the two covers known and unknown fault scenarios and improves the accuracy of fault prediction. In addition, the fault category recognition model and the abnormal pattern recognition model are optimized to adapt to different operation and maintenance scenarios, further improving the accuracy of fault prediction, thereby improving the timeliness of fault prediction.
[0142] Based on the above embodiments, in this embodiment, after the abnormal pattern of the fault is output, an early warning notification is triggered according to the fault category and the abnormal pattern of the fault; the early warning notification is sent to the operation and maintenance end so that the operation and maintenance end can give an early warning.
[0143] In this embodiment, if the fault category is a fault, it means that the computer device may fail within the next few hours; the abnormal pattern of the fault indicates the possible fault problems. According to the fault category and the abnormal pattern of the fault, an early warning notification is triggered to notify the operation and maintenance end to give an early warning. Optionally, a fault report can be generated and sent to the operation and maintenance end.
[0144] In this embodiment, the computer device can be optimized by adjusting resource allocation and dynamically adjusting the CPU and memory allocation of virtual machines.
[0145] In summary, by triggering an early warning notification according to the fault category and the abnormal pattern of the fault, notifying the operation and maintenance end to give an early warning, an early warning is given before potential problems occur, avoiding the failure of the computer device, and improving the stability of the computer device.
[0146] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method.
[0147] Figure 3 This is a schematic structural diagram of a fault prediction device for a computer device provided by an embodiment of the present application. As Figure 3 shown, an embodiment of the present application also provides a fault prediction device for a computer device, including: an acquisition module 301, a preprocessing module 302, an extraction module 303, and an execution module 304. Among them, the execution module 304 includes a first output unit 3041 and a second output unit 3042.
[0148] The acquisition module 301 is used to acquire operation and maintenance data of the computer device.
[0149] The preprocessing module 302 is used to preprocess the operation and maintenance data to obtain preprocessed operation and maintenance data.
[0150] The extraction module 303 is used to extract periodic features and health status features from the preprocessed operation and maintenance data.
[0151] The execution module 304 is used to input the periodic features and health status features into a fault prediction model, where the fault prediction model includes a fault category recognition module and an abnormal pattern recognition module, and the fault prediction model is used to perform the following steps:
[0152] The execution module 304 includes:
[0153] The first output unit 3041 is used to output a fault category through the fault category recognition module according to the periodic features and health status features.
[0154] The second output unit 3042 is used to output an abnormal pattern of the fault through the abnormal pattern recognition module according to the periodic features and health status features.
[0155] In a possible implementation manner, the preprocessing module 302 includes:
[0156] A deletion unit for deleting duplicate data in the operation and maintenance data to obtain first operation and maintenance data;
[0157] A filling unit for filling in missing values in the first operation and maintenance data to obtain second operation and maintenance data;
[0158] A conversion unit for converting unstructured data in the second operation and maintenance data into structured data to obtain third operation and maintenance data;
[0159] A normalization unit, configured to normalize the third operation and maintenance data to obtain preprocessed operation and maintenance data.
[0160] In a possible implementation manner, the conversion unit includes:
[0161] An acquisition subunit, configured to acquire unstructured data in the second operation and maintenance data;
[0162] A segmentation subunit, configured to segment the unstructured data to obtain a plurality of text data blocks;
[0163] A first extraction subunit, configured to, for each text data block, use a regular expression to extract key fields from each text data block;
[0164] A second extraction subunit, configured to, for each text data block, use a named entity recognition technology to extract key entities from each text data block;
[0165] An encapsulation subunit, configured to encapsulate the key fields and key entities into structured data to obtain the third operation and maintenance data.
[0166] In a possible implementation manner, the extraction module 303 includes:
[0167] A first extraction unit, configured to input the preprocessed operation and maintenance data into a convolutional neural network, so that the convolutional layer of the convolutional neural network extracts periodic features from the preprocessed operation and maintenance data;
[0168] A second extraction unit, configured to input the preprocessed operation and maintenance data into a recurrent neural network, so that the recurrent neural network extracts health status features from the preprocessed operation and maintenance data.
[0169] In a possible implementation manner, the fault category recognition module is a fault category recognition model; the abnormal pattern recognition module is an abnormal pattern recognition model; the fault prediction device of the computer device further includes a training module, and the training module includes:
[0170] A first acquisition unit, configured to acquire historical fault data of the computer device; wherein the historical fault data includes historical fault labels;
[0171] A second acquisition unit, configured to acquire historical periodic features and historical health status features in the historical fault data;
[0172] A partitioning unit, configured to partition the historical periodic features and historical health status features into a training set and a test set;
[0173] A first construction unit, configured to use a supervised learning algorithm to recursively construct multiple decision trees according to the training set and the historical fault labels according to a preset segmentation decision to obtain an initial fault category recognition model;
[0174] A first training unit, configured to use a training set to train an initial fault category recognition model to obtain a trained fault category recognition model;
[0175] A first optimization unit, configured to use a test set to optimize the trained fault category recognition model to obtain a final fault category recognition model;
[0176] A second construction unit, configured to use an unsupervised learning algorithm to set a neighborhood radius and a minimum number of samples in the neighborhood to construct an initial abnormal pattern recognition model;
[0177] A second training unit, configured to use a training set to train the initial abnormal pattern recognition model to obtain a trained abnormal pattern recognition model;
[0178] A second optimization unit, configured to use a test set to optimize the trained abnormal pattern recognition model to obtain a final abnormal pattern recognition model.
[0179] In a possible implementation manner, the first optimization unit includes:
[0180] A prediction subunit, configured to input the test set into the trained fault category recognition model, so that the trained fault category recognition model outputs a predicted fault category for the test set;
[0181] An evaluation subunit, configured to evaluate the trained fault category recognition model according to the true fault category and the predicted fault category of the test set to obtain an evaluation result;
[0182] An adjustment subunit, configured to adjust the model parameters of the trained fault category recognition model by using a backpropagation algorithm according to the evaluation result to optimize the trained fault category recognition model to obtain a final fault category recognition model.
[0183] In a possible implementation manner, the fault prediction device of the computer device further includes a warning module, and the warning module includes:
[0184] A triggering unit, configured to trigger a warning notification according to the fault category and the abnormal pattern of the fault;
[0185] A warning unit, configured to send the warning notification to the operation and maintenance end so that the operation and maintenance end gives a warning.
[0186] For the description of the features in the embodiment corresponding to the fault prediction device of the computer device, reference may be made to the relevant description in the embodiment corresponding to the fault prediction method of the computer device, which will not be elaborated here one by one.
[0187] Figure 4 This is a schematic structural diagram of the server provided by the embodiment of the present application. AsFigure 4 As shown in the figure, the server provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the server further includes a communication component 403. Among them, the processor 401, the memory 402, and the communication component 403 are connected through a bus.
[0188] In the specific implementation process, at least one processor 401 executes the computer-executable instructions stored in the memory 402, so that at least one processor 401 executes the fault prediction method embodiment of the above computer device.
[0189] For the specific implementation process of the processor 401, reference can be made to the above method embodiment. Its implementation principle and technical effect are similar, and will not be elaborated here in this embodiment.
[0190] In the above embodiment, it should be understood that the processor may be a central processing unit (Central Processing Unit, abbreviated as CPU), or other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0191] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0192] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, the bus in the attached drawings of this application is not limited to only one bus or one type of bus.
[0193] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. Among them, the computer program is set to execute the steps in any one of the above-mentioned fault prediction method embodiments of the computer device when running.
[0194] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), external hard drives, magnetic disks, or optical discs.
[0195] The embodiment of the present application also provides a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in the embodiment of any one of the above-mentioned fault prediction methods for computer devices.
[0196] The embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in the embodiment of any one of the above-mentioned fault prediction methods for computer devices.
[0197] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0198] The above has introduced in detail a fault prediction method, device, server, and storage medium for a computer device provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for predicting faults of a computer device, characterized in that, Including: Collecting the operation and maintenance data of computer devices; Preprocessing the operation and maintenance data to obtain preprocessed operation and maintenance data; Extracting periodic features and health status features from the preprocessed operation and maintenance data; Inputting the periodic features and the health status features into a fault prediction model, where the fault prediction model includes a fault category recognition module and an abnormal pattern recognition module, and the fault prediction model is used to perform the following steps: Outputting a fault category through the fault category recognition module according to the periodic features and the health status features; Outputting an abnormal pattern of the fault through the abnormal pattern recognition module according to the periodic features and the health status features.
2. The method according to claim 1, characterized in that, The preprocessing the operation and maintenance data to obtain preprocessed operation and maintenance data includes: Deleting duplicate data in the operation and maintenance data to obtain first operation and maintenance data; Filling in missing values in the first operation and maintenance data to obtain second operation and maintenance data; Converting unstructured data in the second operation and maintenance data into structured data to obtain third operation and maintenance data; Normalizing the third operation and maintenance data to obtain preprocessed operation and maintenance data.
3. The method according to claim 2, wherein The converting unstructured data in the second operation and maintenance data into structured data to obtain third operation and maintenance data includes: Obtaining unstructured data in the second operation and maintenance data; Splitting the unstructured data to obtain multiple text data blocks; For each text data block, using regular expressions to extract keyword fields from each text data block; For each text data block, using named entity recognition technology to extract key entities from each text data block; Encapsulating the keyword fields and the key entities into structured data to obtain third operation and maintenance data.
4. The method according to claim 1, characterized in that, The extracting periodic features and health status features from the preprocessed operation and maintenance data includes: Inputting the preprocessed operation and maintenance data into a convolutional neural network, so that the convolutional layer of the convolutional neural network extracts periodic features from the preprocessed operation and maintenance data; Inputting the preprocessed operation and maintenance data into a recurrent neural network, so that the recurrent neural network extracts health status features from the preprocessed operation and maintenance data.
5. The method according to claim 1, characterized in that The fault category recognition module is a fault category recognition model; the abnormal pattern recognition module is an abnormal pattern recognition model; Correspondingly, before collecting the operation and maintenance data of computer devices, it further includes: Obtaining historical fault data of computer devices; where the historical fault data includes historical fault labels; Obtaining historical periodic features and historical health status features in the historical fault data; Dividing the historical periodic features and the historical health status features into a training set and a test set; Using a supervised learning algorithm, according to the training set and the historical fault labels, recursively constructing multiple decision trees according to a preset splitting decision to obtain an initial fault category recognition model; Using the training set to train the initial fault category recognition model to obtain a trained fault category recognition model; Using the test set, optimize the trained fault category recognition model to obtain the final fault category recognition model; Adopt an unsupervised learning algorithm, set the neighborhood radius and the minimum number of samples in the neighborhood, and construct an initial abnormal pattern recognition model; Use the training set to train the initial abnormal pattern recognition model to obtain a trained abnormal pattern recognition model; Use the test set to optimize the trained abnormal pattern recognition model to obtain the final abnormal pattern recognition model.
6. The method according to claim 5, wherein The step of using the test set to optimize the trained fault category recognition model to obtain the final fault category recognition model includes: Input the test set into the trained fault category recognition model, so that the trained fault category recognition model outputs the predicted fault category for the test set; Evaluate the trained fault category recognition model according to the true fault category of the test set and the predicted fault category to obtain an evaluation result; According to the evaluation result, adopt the backpropagation algorithm to adjust the model parameters of the trained fault category recognition model to optimize the trained fault category recognition model and obtain the final fault category recognition model.
7. The method according to any one of claims 1-6, characterized in that, After outputting the abnormal pattern of the fault, it further includes: Trigger a warning notification according to the fault category and the abnormal pattern of the fault; Send the warning notification to the operation and maintenance end so that the operation and maintenance end gives a warning.
8. A fault prediction device for a computer device, characterized in that, It includes: A collection module for collecting operation and maintenance data of computer devices; A preprocessing module for preprocessing the operation and maintenance data to obtain preprocessed operation and maintenance data; An extraction module for extracting periodic features and health status features from the preprocessed operation and maintenance data; An execution module for inputting the periodic features and the health status features into a fault prediction model, where the fault prediction model includes a fault category recognition module and an abnormal pattern recognition module, and the fault prediction model is used to perform the following steps: The execution module includes: A first output unit for outputting a fault category through the fault category recognition module according to the periodic features and the health status features; A second output unit for outputting an abnormal pattern of the fault through the abnormal pattern recognition module according to the periodic features and the health status features.
9. A server, characterized in that, It includes: A memory for storing a computer program; A processor for implementing the steps of the fault prediction method of the computer device according to any one of claims 1-7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the fault prediction method of the computer device according to any one of claims 1-7 when executed by a processor.
Citation Information
Patent Citations
Operation and maintenance automation system and method
CN105323111A
Predictive operation and maintenance scheme generation method and device of equipment, terminal equipment and medium
CN116258484A
Method and system for monitoring health state of transformer of power station
CN117630758A
Fault processing method and device of cloud computing platform, electronic equipment and storage medium
CN119557134A
Software operation anomaly detection method, system, medium, product and equipment
CN119557629A