System Monitoring Method, Device, Electronic Device and Storage Medium
The LSTM neural network preprocesses and model tuning of the application system's historical data is solved, and the problem of low prediction accuracy in high distributed systems is achieved, accurate prediction of future performance changes trends is achieved, and the stability and prediction efficiency of the system are improved.
Patent Information
- Application Number
- CN202210729168.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-06-24
AI Technical Summary
In the operational state prediction of high-distributed application systems, the prediction results cannot provide sufficient operational state change information, resulting in high false alarm rate and low accuracy of traditional fault prediction methods.
The time series prediction technology based on LSTM neural network is adopted to preprocess historical running data, train the initial model and tune it to obtain the target model, and dynamically optimize the model parameters to improve the prediction accuracy.
It improves the accuracy of prediction of the performance trend of the application system, can predict abnormal states in a timely manner, reduce the rate of false alarms, helps operation and maintenance personnel to take preventive measures in advance, and improve system stability.
Smart Images

Figure CN115115030B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of system monitoring, and in particular, to a system monitoring method, device, electronic device and storage medium. Background Art
[0002] Currently, in the process of predicting the running state of a highly distributed application system, the prediction results usually cannot convey sufficient information about the changes in the running state to the task scheduler. And most traditional fault prediction methods are to set a threshold, and then determine whether an abnormality has occurred by judging whether the prediction result exceeds the range of this threshold, which will cause a relatively high false alarm rate. Summary of the Invention
[0003] In view of the above, it is necessary to propose a system monitoring method, device, electronic device and storage medium, which can predict the performance of the application system based on the time series prediction technology of the LSTM neural network, obtain the performance change trend of the application system in the next period of time, and improve the accuracy of anomaly prediction.
[0004] The first aspect of the present invention provides a system monitoring method, and the method includes: obtaining historical operation data and real-time operation data of the application system;
[0005] Preprocessing the historical operation data, and dividing the preprocessed historical operation data into a training set and a test set;
[0006] Training a long short-term memory (LSTM) neural network based on the training set to obtain an initial model, testing the initial model using the test set, and tuning the model parameters of the initial model according to the test results until a target model meeting the preset requirements is obtained;
[0007] Inputting the real-time operation data of the application system into the target model for prediction, obtaining the predicted operation data of the application system within a preset time period, and the predicted operation data includes the time and probability of the occurrence of the predicted abnormal state;
[0008] Calculating the accuracy rate of the real-time prediction of the target model, and dynamically optimizing the target model according to the accuracy rate.
[0009] According to an optional embodiment of the present invention, the preprocessing includes: data integration, data cleaning, data standardization and data segmentation;
[0010] The data integration is a process of synthesizing data from multiple sources into one data for storage;
[0011] The data cleaning includes denoising the historical operation data based on the 3σ criterion of the normal distribution;
[0012] The data standardization includes normalizing the historical operation data;
[0013] The data segmentation includes segmenting the historical operation data according to a preset first time length, taking the data within each preset first time length after segmentation as a group of data, and taking the first data arranged in chronological order after any group of data as the label of the any group of data, where the label represents the predicted data of the any group of data.
[0014] According to an optional embodiment of the present invention, training the long short-term memory (LSTM) neural network based on the training set to obtain an initial model includes:
[0015] Grouping the training set according to a preset second time length to obtain a plurality of sub-training sets;
[0016] Using each sub-training set to batch-train the LSTM neural network, so that the initial model performs multi-step point value prediction, where the multi-step point value prediction includes: predicting the initial prediction data of the third time length of each sub-training set based on each sub-training set.
[0017] According to an optional embodiment of the present invention, using each sub-training set to batch-train the LSTM neural network, so that the initial model performs multi-step point value prediction, includes:
[0018] Based on any sub-training set and the initial model, performing point value prediction on the abnormal state of the application system to obtain a first prediction point value corresponding to any sub-training set of the application system, where the first prediction point value includes the occurrence time point of the abnormal state predicted by the any sub-training set;
[0019] Determining the third time length based on the first prediction point value includes: taking the time length within a preset first percentage before and after the first prediction point value as the third time length, and taking the prediction data of the third time length including the first prediction point value as the first confidence interval of the first prediction point value.
[0020] According to an optional embodiment of the present invention, the model parameters of the initial model include the first percentage, testing the initial model using the test set, and optimizing the model parameters of the initial model according to the test results until a target model meeting the preset requirements is obtained, includes:
[0021] Grouping the test set according to the second time length to obtain a plurality of sub-test sets;
[0022] Input any sub - test set into the initial model, and use the initial model to output the second predicted point value corresponding to the any sub - test set and the second confidence interval containing the second predicted point value;
[0023] Take the time point when the actual abnormal state occurs in the second confidence interval as the first actual point value corresponding to the any sub - test set, and calculate the first difference between the first actual point value and the second predicted point value as the error corresponding to the any sub - test set;
[0024] Calculate the root - mean - square error of the errors of all sub - test sets, and determine whether the root - mean - square error meets the preset requirements. The preset requirements include: the root - mean - square error is less than or equal to a preset error threshold;
[0025] When the root - mean - square error does not meet the preset requirements, exhaustively update the first percentage;
[0026] When the root - mean - square error meets the preset requirements, take the updated minimum percentage that meets the preset requirements as the second percentage, and obtain the target model that meets the preset requirements.
[0027] According to an optional implementation manner of the present invention, calculating the real - time prediction accuracy rate of the target model and dynamically optimizing the target model according to the accuracy rate includes:
[0028] Calculate the accuracy rate once every preset fourth time period, including: calculating the second difference between the time when the predicted abnormal state occurs and the time when the actual abnormal state occurs in each fourth time period, and the third difference between the probability that the predicted abnormal state occurs and 1, and taking the product of the second difference and the third difference as the accuracy rate;
[0029] Determine whether the accuracy rate reaches the range of a preset accuracy rate threshold;
[0030] When the accuracy rate does not reach the range of the preset accuracy rate threshold, optimize the target model, including: updating the historical operation data with the data in the previous fourth time period, using the updated historical operation data as the updated test set, and tuning the model parameters of the target model based on the updated test set until an updated target model that meets the preset requirements is obtained. The model parameters of the target model include the second percentage.
[0031] According to an optional implementation manner of the present invention, the method further includes:
[0032] Use a preset visualization interface to display the operating status of the application system during the preset time period, and send the predicted abnormal status to preset associated personnel to notify the maintenance of the application system.
[0033] The second aspect of the present invention provides a system monitoring device, which includes an acquisition module, a processing module, a training module, a prediction module, and an optimization module:
[0034] The acquisition module is used to acquire the historical operation data and real-time operation data of the application system;
[0035] The processing module is used to preprocess the historical operation data and divide the preprocessed historical operation data into a training set and a test set;
[0036] The training module is used to train a long short-term memory (LSTM) neural network based on the training set to obtain an initial model, test the initial model using the test set, and optimize the model parameters of the initial model according to the test results until a target model that meets the preset requirements is obtained;
[0037] The prediction module is used to input the real-time operation data of the application system into the target model for prediction, and obtain the predicted operation data of the application system within a preset time period, where the predicted operation data includes the time and probability of the occurrence of a predicted abnormal status;
[0038] The optimization module is used to calculate the accuracy rate of the real-time prediction of the target model and dynamically optimize the target model according to the accuracy rate.
[0039] The third aspect of the present invention provides an electronic device, which includes a processor and a memory. The processor is used to implement the system monitoring method when executing a computer program stored in the memory.
[0040] The fourth aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. The computer program is used to implement the system monitoring method when executed by a processor.
[0041] In summary, for the system monitoring method, device, electronic device, and storage medium according to the present invention, first, the collected data is preprocessed to delete redundant and invalid data. An initial model is trained using an LSTM neural network, and the initial model is optimized to obtain a target model. The target model is dynamically optimized by an active detection method, thereby improving the accuracy of predicting the future change trend of the application system performance by the model. Applying the time series prediction technology based on the LSTM neural network to the application system performance monitoring can not only obtain the overall operating condition of the application system but also predict the performance change trend of the application system in the future for a period of time, helping the operation and maintenance personnel to take corresponding preventive measures before the system fails. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 FIG. 1 is a flowchart of the system monitoring method provided in Embodiment 1 of the present invention.
[0043] Figure 2 FIG. 2 is a structural diagram of the system monitoring device provided in Embodiment 2 of the present invention.
[0044] Figure 3 FIG. 3 is a schematic structural diagram of the electronic device provided in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] In order to more clearly understand the above objects, features, and advantages of the present invention, the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, without conflict, the embodiments of the present invention and the features in the embodiments can be combined with each other.
[0046] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the specification of the present invention herein are only for the purpose of describing an optional embodiment and are not intended to limit the present invention.
[0047] The system monitoring method provided in the embodiment of the present invention is executed by an electronic device. Correspondingly, the system monitoring device runs in the electronic device.
[0048] Embodiment 1
[0049] Figure 1 FIG. 1 is a flowchart of the system monitoring method provided in Embodiment 1 of the present invention. The system monitoring method specifically includes the following steps. According to different requirements, the order of the steps in this flowchart can be changed, and some steps can be omitted.
[0050] S11, obtaining historical operation data and real-time operation data of the application system.
[0051] In an alternative embodiment, the solution provided by the embodiments of the present application can be applied to monitor multiple application systems to predict the time and probability of the occurrence of abnormal states of the application systems, so as to respond to and maintain the abnormal states of the application systems in a timely manner, and improve the stability and working efficiency of the application systems.
[0052] In an alternative embodiment, the operation data (or performance data) of the application system includes: the response time, the number of concurrent users, the throughput, the error rate, the page loading time, the data of the load status, etc. of the application system at each time node (or called time point), wherein the load status includes the resource usage of the application system, such as the memory occupancy rate of the hardware storage device, etc.
[0053] In an alternative embodiment, the operation data can be obtained or received from the monitoring platform or operation log of the application system at fixed intervals. For example, it is obtained every 0.1 seconds, and the data at the moment when the acquisition operation is executed is used as the real-time operation data, and the data before the acquisition operation is executed is used as the historical operation data.
[0054] In an alternative embodiment, the historical operation data is stored in a preset database, and as time goes by, after each execution of the acquisition operation, the latest obtained operation data is stored in the database and the historical data is updated.
[0055] S12, preprocess the historical operation data, and divide the preprocessed historical operation data into a training set and a test set.
[0056] In an alternative embodiment, a large amount of operation data or performance parameter data will be generated during the actual operation of the application system. These data may contain a large amount of noise and abnormal data. In addition, the collected time series data may be incomplete. If these data are directly used to train the model, it will seriously affect the training efficiency of the model and the accuracy of model prediction. Therefore, it is necessary to preprocess the historical operation data.
[0057] In an alternative embodiment, the preprocessing includes: data integration, data cleaning, data standardization, data segmentation, etc. Data integration is the process of synthesizing data from multiple sources and storing them as one data.
[0058] In an alternative embodiment, the data cleaning includes denoising the historical operation data based on the 3σ criterion of the normal distribution. The data cleaning can delete the data irrelevant to abnormal prediction, duplicate data, etc. in the historical operation data, and handle the missing values and abnormal values in the historical operation data.
[0059] For example, the load status of the application system is not constant. Network communication of the application system, switching of processes in the host kernel, etc. will cause fluctuations in the load status. If the historical operation data of the load status is not preprocessed, it will affect the convergence speed of the historical operation data and the prediction accuracy of the subsequent trained model.
[0060] Therefore, before training and using a Long Short-Term Memory (LSTM) network for prediction, the method provided by the embodiments of this application uses the 3σ criterion of the normal distribution to perform data cleaning and smoothing on the historical operation data. If the historical operation data follows a normal distribution, assuming that the mean of the historical operation data is μ and the standard deviation is b, then the probability that the data falls outside [μ - 3σ, μ + 3σ] is less than three thousandths. Therefore, the data outside this range is called small-probability data and is deleted as noise data.
[0061] In an optional implementation manner, the data standardization includes normalizing the historical operation data. Data unification mainly performs normalization processing on the data and converts the data into the standard form of the model data set uniformly.
[0062] Performing data standardization on the historical operation data before training the LSTM model can accelerate the convergence of the loss function of the LSTM model and facilitate the calculation of gradients. The method provided by the embodiments of this application adopts the normalization (min-max) method.
[0063] For example, if the time series of the abnormal status of the load status in the historical operation data is L = {L1, L2, L3,..., Ln}, then the formula used for min-max normalization of L is: Li' = (Li - Lmin) / (Lmax - Lmin), where Li' represents the data after standardization, i represents an integer greater than or equal to 1 and less than or equal to n, Lmax represents the largest number in L, and Lmin represents the smallest number in L.
[0064] In an optional implementation manner, the data segmentation includes segmenting the historical operation data according to a preset first time length, taking the data within each preset first time length after segmentation as a group of data, and taking the first data arranged in chronological order after any group of data as the label of the any group of data, where the label represents the predicted data of the any group of data.
[0065] In an alternative embodiment, in order to adapt to the characteristics of the input of the hidden layer of the LSTM neural network, the data needs to be segmented and converted into time series data. The principle of the LSTM neural network for time series prediction is to infer future data from the characteristics of historical data. Similarly, LSTM also needs to perform a similar data conversion.
[0066] For example, assume that the historical operation data is the data set S = {A, B, C, D, E, F}, which can be segmented according to a preset first time length, such as 0.3 seconds. Each segment within the first time length contains 3 historical data, and 3 historical data are used to predict a future data. Then the data set S is segmented or constructed in the following format: S = [[A, B, C]->[D], [B, C, D]->[E], [C, D, E]->[F]]. Among them, after the conversion is completed, the data set S can be regarded as a list. There are three elements in the list, and each element is divided into two parts: data and label. For the element [A, B, C]->[D], [A, B, C] is the data input to the LSTM algorithm, and D is the label or output data of [A, B, C], that is, D is the predicted data of [A, B, C].
[0067] In an alternative embodiment, the preprocessed historical operation data can be divided into a training set and a test set according to a preset ratio. For example, one-third of the preprocessed historical operation data is randomly selected as the test set, and the rest is used as the training set.
[0068] S13. Train a long short-term memory (LSTM) neural network based on the training set to obtain an initial model, test the initial model using the test set, and optimize the model parameters of the initial model according to the test results until a target model that meets the preset requirements is obtained.
[0069] In an alternative embodiment, the training of the long short-term memory (LSTM) neural network based on the training set to obtain an initial model includes:
[0070] Group the training set according to a preset second time length to obtain a plurality of sub-training sets;
[0071] Use each sub-training set to batch-train the LSTM neural network, so that the initial model performs multi-step point value prediction. The multi-step point value prediction includes: predicting the initial prediction data of the third time length of each sub-training set based on each sub-training set.
[0072] In an alternative embodiment, the training set data is grouped or batched according to a preset second time length to obtain a plurality of sub-training sets. Using the sub-training sets to batch-train the LSTM neural network can improve the speed of model training, reduce the time required for the loss function of the model to converge without sacrificing the model accuracy. The loss function represents the difference between the predicted accuracy of the model and the actual accuracy (the value is 1); the smaller the preset value to which the loss function converges, the higher the accuracy of the target detection model.
[0073] In an alternative embodiment, using each sub-training set to batch-train the LSTM neural network enables the initial model to perform multi-step point value prediction, including:
[0074] Based on any one of the sub-training sets and the initial model, perform point value prediction on the abnormal state of the application system to obtain a first predicted point value corresponding to any one of the sub-training sets of the application system. The first predicted point value includes the occurrence time point or time range of the abnormal state predicted by any one of the sub-training sets;
[0075] Determine the third time length based on the first predicted point value, including: taking the time length within a preset first percentage before and after the first predicted point value as the third time length, and using the prediction data of the third time length including the first predicted point value as the first confidence interval of the first predicted point value.
[0076] For example, use the LSTM model to predict the future load status of the application system through multi-step point value prediction. In actual situations, a single or single-step point value is usually an ideal estimated value, which can represent the estimated value of the average load status of the application system at a given time node or time range, but cannot reflect the change trend of the load status of the application system at this time.
[0077] In a highly distributed application system, a single or single-step point value usually cannot convey sufficient change information of the load status to the task scheduler, and most traditional fault prediction methods based on single or single-step point values set a threshold, and then determine whether an anomaly has occurred by judging whether the predicted result exceeds the range of this threshold, which will result in a relatively high false alarm rate.
[0078] If a better interval range can be determined to predict the load status within this better range, that is, to determine the first confidence interval of the prediction of the load status, then the change trend within the confidence interval can be obtained from the load status at the current time. Among them, the confidence interval represents an interval with a relatively high confidence that the actual value falls around the prediction result.
[0079] In an alternative embodiment, the framework and structure of the initial model are the same as those of a general LSTM neural network. The first percentage can be preset to 20%, and the first 20% range of the first prediction point value in the sub-training set and the last 20% range of the first prediction point value in the sub-training set are used as the third time length. Subsequently, in subsequent steps, the first percentage is updated, and the parameters of the initial model are optimized to update the initial model.
[0080] In an alternative embodiment, the model parameters of the initial model include the first percentage. Testing the initial model using the test set and optimizing the model parameters of the initial model according to the test results until a target model meeting the preset requirements is obtained, including:
[0081] Grouping the test set according to the second time length to obtain a plurality of sub-test sets;
[0082] Inputting any one of the sub-test sets into the initial model, and using the initial model to output the second prediction point value corresponding to the any one of the sub-test sets and a second confidence interval including the second prediction point value;
[0083] Taking the time point at which the actual abnormal state occurs in the second confidence interval as the first actual point value corresponding to the any one of the sub-test sets, and calculating the first difference between the first actual point value and the second prediction point value as the error corresponding to the any one of the sub-test sets;
[0084] Calculating the root mean squared error (RMSE) of the errors of all sub-test sets, and determining whether the root mean squared error meets the preset requirements. The preset requirements include: the root mean squared error is less than or equal to a preset error threshold;
[0085] When the root mean squared error does not meet the preset requirements, exhaustively update the first percentage;
[0086] When the root mean squared error meets the preset requirements, taking the updated minimum percentage that meets the preset requirements as the second percentage to obtain the target model that meets the preset requirements.
[0087] S14, inputting the real-time operation data of the application system into the target model for prediction to obtain the predicted operation data of the application system within a preset time period. The predicted operation data includes the time and probability of the occurrence of a predicted abnormal state.
[0088] In an alternative embodiment, the preset time period includes a fifth time length obtained based on the second percentage. The predicted abnormal state includes, for example, an abnormal state such as the predicted load state of the application system exceeding the load threshold.
[0089] S15. Calculate the accuracy rate of the real-time prediction of the target model, and dynamically optimize the target model according to the accuracy rate.
[0090] In an alternative embodiment, the calculating the accuracy rate of the real-time prediction of the target model and dynamically optimizing the target model according to the accuracy rate includes:
[0091] Calculate the accuracy rate once every preset fourth time length, including: calculating a second difference between the time when the predicted abnormal state occurs and the time when the actual abnormal state occurs in each fourth time length, and a third difference between the probability of the predicted abnormal state occurring and 1, and taking the product of the second difference and the third difference as the accuracy rate;
[0092] Determine whether the accuracy rate reaches the range of a preset accuracy rate threshold;
[0093] When the real-time prediction accuracy rate does not reach the range of the preset accuracy rate threshold, optimize the target model, including: updating the historical operation data with the data in the previous fourth time length, using the updated historical operation data as an updated test set, and tuning the model parameters of the target model based on the updated test set until an updated target model meeting the preset requirements is obtained. The model parameters of the target model include the second percentage.
[0094] Specifically, the embodiment of the present application proposes an adaptive parameter adjustment method, which adjusts the parameters of the target model during the prediction process, so that the third confidence interval generated based on the second percentage can ensure the interval span (corresponding to the fifth time length), and further improve the accuracy rate of the target model. When using the target model of the LSTM algorithm based on the confidence interval for prediction, the most suitable third confidence interval is predicted by adjusting the size of the second percentage.
[0095] Set the fourth time length as an adaptive interval parameter, where the adaptive interval parameter indicates that an adaptive process is triggered every fourth time length to verify the accuracy of the target model. If the accuracy does not reach the range of the preset accuracy threshold, adjust and update the second percentage based on the updated test set, and obtain an updated third confidence interval and an updated fifth time length according to the updated second percentage, so as to obtain an updated target model, and use the updated target model to monitor the application system within the next fourth time length, and predict the availability of the application system within the next fourth time length in real time.
[0096] In an alternative embodiment, the method further includes: using a preset visualization interface to display the running status of the application system within the preset time period, and sending the predicted abnormal status to a preset associated person to notify the maintenance of the application system. For example, in the visualization interface, display the load status of the application system in the form of a table or a curve graph, etc., and send the predicted abnormal status of the load status to the app of the terminal communication device of the operation and maintenance personnel. The operation and maintenance personnel can also directly refer to the prediction results in the visualization interface to optimize the application system, etc.
[0097] In an alternative embodiment, the system monitoring method provided by the embodiments of the present application first preprocesses the collected data, deletes redundant and invalid data, trains an initial model using an LSTM neural network, and tunes the initial model to obtain a target model, and dynamically optimizes the target model by an active detection method, so as to improve the accuracy of the model in predicting the future change trend of the application system performance. Applying the time series prediction technology based on the LSTM neural network to the application system performance monitoring can not only obtain the overall running status of the application system, but also predict the performance change trend of the application system in the future for a period of time, helping the operation and maintenance personnel to take corresponding preventive measures before the system fails.
[0098] Embodiment 2
[0099] Figure 2 It is a structural diagram of the system monitoring device provided by Embodiment 2 of the present invention.
[0100] In some embodiments, the system monitoring device 20 may include a plurality of functional modules composed of computer program segments. The computer programs of each program segment in the system monitoring device 20 may be stored in the memory of the electronic device and executed by at least one processor to perform the functions of system monitoring (see Figure 1 description).
[0101] In this embodiment, the system monitoring device 20 can be divided into multiple functional modules according to the functions it performs. The functional modules may include: an acquisition module 201, a processing module 202, a training module 203, a prediction module 204, and an optimization module 205. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0102] The acquisition module 201 is used to acquire the historical operation data and real-time operation data of the application system.
[0103] In an optional implementation manner, the solution provided in the embodiment of the present application can be applied to monitor multiple application systems to predict the time and probability of the occurrence of abnormal states of the application systems, so as to respond to and maintain the abnormal states of the application systems in a timely manner, and improve the stability and working efficiency of the application systems.
[0104] In an optional implementation manner, the operation data (or performance data) of the application system includes: the response time, the number of concurrent users, the throughput, the error rate, the page loading time, the data of the load status, etc. of the application system at each time node (or called time point), where the load status includes the resource usage of the application system, such as the memory occupancy rate of the hardware storage device, etc.
[0105] In an optional implementation manner, the operation data can be acquired or received from the monitoring platform or operation log of the application system at fixed intervals. For example, acquisition is performed every 0.1 seconds, and the data at the moment of performing the acquisition operation is used as the real-time operation data, and the data before performing the acquisition operation is used as the historical operation data.
[0106] In an optional implementation manner, the historical operation data is stored in a preset database, and as time goes by, after each acquisition operation is performed, the latest obtained operation data is stored in the database and the historical data is updated.
[0107] The processing module 202 is used to preprocess the historical operation data and divide the preprocessed historical operation data into a training set and a test set.
[0108] In an optional implementation manner, during the actual operation of the application system, a large amount of operation data or performance parameter data will be generated. These data may contain a large amount of noise and abnormal data. In addition, the collected time series data may be incomplete. If these data are directly used to train the model, it will seriously affect the efficiency of training the model and the accuracy of model prediction. Therefore, it is necessary to preprocess the historical operation data.
[0109] In an alternative embodiment, the preprocessing includes: data integration, data cleaning, data standardization, data segmentation, etc. The data integration is a process of combining data from multiple sources into one data for storage.
[0110] In an alternative embodiment, the data cleaning includes denoising the historical operation data based on the 3σ criterion of the normal distribution. The data cleaning can delete data irrelevant to anomaly prediction, duplicate data, etc. in the historical operation data, and handle missing values and outliers in the historical operation data.
[0111] For example, the load status of the application system is not constant. Fluctuations in the load status will occur due to network communication of the application system, switching of processes in the host kernel, etc. If the historical operation data of the load status is not preprocessed, it will affect the convergence speed of the historical operation data and the prediction accuracy of the subsequent trained model.
[0112] Therefore, before training and using a Long Short-Term Memory (LSTM) network for prediction, the method provided by the embodiments of the present application uses the 3σ criterion of the normal distribution to perform data cleaning and smoothing on the historical operation data. If the historical operation data follows a normal distribution, assuming the mean of the historical operation data is μ and the standard deviation is b, then the probability that the data falls outside [μ - 3σ, μ + 3σ] is less than 3 per thousand. Therefore, data exceeding this range is called small-probability data and is deleted as noise data.
[0113] In an alternative embodiment, the data standardization includes normalizing the historical operation data. Data unification mainly performs normalization processing on the data and converts the data into the standard form of the model data set uniformly.
[0114] Performing data standardization on the historical operation data before training the LSTM model can accelerate the convergence of the loss function of the LSTM model and facilitate the calculation of gradients. The method provided by the embodiments of the present application adopts the normalization (min-max) method.
[0115] For example, if the time series of the occurrence of the abnormal status of the load status in the historical operation data is L = {L1, L2, L3,..., Ln}, then the formula used for min-max standardization of L is: Li' = (Li - Lmin) / (Lmax - Lmin), where Li' represents the data after standardization, i represents an integer greater than or equal to 1 and less than or equal to n, Lmax represents the largest number in L, and Lmin represents the smallest number in L.
[0116] In an alternative embodiment, the data segmentation includes segmenting the historical operation data according to a preset first time length, taking the data within each preset first time length after segmentation as a set of data, and taking the first data arranged in chronological order after any set of data as the label of the any set of data, where the label represents the predicted data of the any set of data.
[0117] In an alternative embodiment, in order to adapt to the characteristics of the input of the hidden layer of the LSTM neural network, it is necessary to segment the data and convert it into time series data. The principle of the LSTM neural network for time series prediction is to infer future data from the characteristics of historical data. Similarly, the LSTM also needs to perform a similar data conversion.
[0118] For example, assuming that the historical operation data is the data set S = {A, B, C, D, E, F}, it can be segmented according to a preset first time length, such as 0.3 seconds. Each segment within the first time length after segmentation contains 3 historical data, and 3 historical data are used to predict a future data. Then the data set S is segmented or constructed in the following format: S = [[A, B, C]->[D], [B, C, D]->[E], [C, D, E]->[F]]. Among them, after the conversion is completed, the data set S can be regarded as a list. There are three elements in the list, and each element is divided into two parts: data and label. For the element [A, B, C]->[D], [A, B, C] is the data input to the LSTM algorithm, and D is the label or output data of [A, B, C], that is, D is the predicted data of [A, B, C].
[0119] In an alternative embodiment, the preprocessed historical operation data can be divided into a training set and a test set according to a preset ratio. For example, one-third of the preprocessed historical operation data is randomly selected as the test set, and the rest is used as the training set.
[0120] The training module 203 is configured to train a long short-term memory (LSTM) neural network based on the training set to obtain an initial model, test the initial model using the test set, and optimize the model parameters of the initial model according to the test results until a target model meeting the preset requirements is obtained.
[0121] In an alternative embodiment, the training of the long short-term memory (LSTM) neural network based on the training set to obtain an initial model includes:
[0122] Grouping the training set according to a preset second time length to obtain a plurality of sub-training sets;
[0123] Use each sub-training set to batch-train the LSTM neural network, so that the initial model performs multi-step point value prediction. The multi-step point value prediction includes: predicting the initial prediction data of the third time length of each sub-training set based on each sub-training set.
[0124] In an alternative embodiment, the training set data is grouped or batched according to a preset second time length to obtain multiple sub-training sets. Using the sub-training sets to batch-train the LSTM neural network can improve the speed of model training and reduce the time required for the loss function of the model to converge without sacrificing the model accuracy. The loss function represents the difference between the prediction accuracy of the model and the actual accuracy (the value is 1); the smaller the preset value to which the loss function converges, the higher the accuracy of the target detection model.
[0125] In an alternative embodiment, the using each sub-training set to batch-train the LSTM neural network so that the initial model performs multi-step point value prediction includes:
[0126] Based on any sub-training set and the initial model, perform point value prediction on the abnormal state of the application system to obtain the first predicted point value corresponding to any sub-training set of the application system. The first predicted point value includes the occurrence time point or time range of the abnormal state predicted from any sub-training set.
[0127] Determining the third time length based on the first predicted point value includes: taking the time length within a preset first percentage before and after the first predicted point value as the third time length, and taking the prediction data of the third time length including the first predicted point value as the first confidence interval of the first predicted point value.
[0128] For example, use the LSTM model to predict the future load status of the application system through multi-step point value prediction. In actual situations, a single or single-step point value is usually an ideal estimated value, which can represent the estimated value of the average load status of the application system at a given time node or time range, but cannot reflect the change trend of the load status of the application system at this time.
[0129] In a highly distributed application system, a single or single-step point value usually cannot convey sufficient change information of the load status to the task scheduler, and most traditional fault prediction methods based on single or single-step point values set a threshold and then determine whether an abnormality has occurred by judging whether the prediction result exceeds the range of this threshold, which will result in a relatively high false alarm rate.
[0130] If a better interval range can be determined to predict the load status within this better range, that is, to determine the first confidence interval for the prediction of the load status, then the change trend within the confidence interval can be obtained from the load status at the current time. Among them, the confidence interval represents an interval with a relatively high confidence level where the actual value falls around the prediction result.
[0131] In an alternative embodiment, the framework and structure of the initial model are the same as those of a general LSTM neural network. The first percentage can be preset to 20%, and the range of the first 20% of the first prediction point value in the sub-training set and the range of the last 20% of the first prediction point value in the sub-training set are used as the third time length. Then, in subsequent steps, the first percentage is updated, and the parameters of the initial model are tuned to optimize and update the initial model.
[0132] In an alternative embodiment, the model parameters of the initial model include the first percentage. Testing the initial model using the test set and tuning the model parameters of the initial model according to the test results until a target model meeting the preset requirements is obtained, including:
[0133] Group the test set according to the second time length to obtain a plurality of sub-test sets;
[0134] Input any one of the sub-test sets into the initial model, and use the initial model to output the second prediction point value corresponding to the any one of the sub-test sets, and a second confidence interval including the second prediction point value;
[0135] Take the time point when the actual abnormal state occurs in the second confidence interval as the first actual point value corresponding to the any one of the sub-test sets, and calculate the first difference between the first actual point value and the second prediction point value as the error corresponding to the any one of the sub-test sets;
[0136] Calculate the root mean squared error (RMSE) of the errors of all sub-test sets, and determine whether the root mean squared error meets the preset requirements. The preset requirements include: the root mean squared error is less than or equal to a preset error threshold;
[0137] When the root mean squared error does not meet the preset requirements, exhaustively update the first percentage;
[0138] When the root mean squared error meets the preset requirements, take the updated minimum percentage that meets the preset requirements as the second percentage to obtain the target model that meets the preset requirements.
[0139] The prediction module 204 is configured to input the real-time operation data of the application system into the target model for prediction, so as to obtain the predicted operation data of the application system within a preset time period, where the predicted operation data includes the time and probability of the occurrence of a predicted abnormal state.
[0140] In an optional embodiment, the preset time period includes a fifth time length obtained based on the second percentage. The predicted abnormal state includes, for example, an abnormal state where the predicted load state of the application system exceeds a load threshold.
[0141] The optimization module 205 is configured to calculate the accuracy rate of the real-time prediction of the target model, and dynamically optimize the target model according to the accuracy rate.
[0142] In an optional implementation manner, calculating the accuracy rate of the real-time prediction of the target model and dynamically optimizing the target model according to the accuracy rate includes:
[0143] Calculating the accuracy rate once every preset fourth time length, including: calculating a second difference between the time when the predicted abnormal state occurs and the time when the actual abnormal state occurs in each fourth time length, and a third difference between the probability of the occurrence of the predicted abnormal state and 1, and taking the product of the second difference and the third difference as the accuracy rate;
[0144] Determining whether the accuracy rate reaches a preset accuracy rate threshold range;
[0145] When the real-time prediction accuracy rate does not reach the preset accuracy rate threshold range, optimizing the target model, including: updating the historical operation data with the data in the previous fourth time length, using the updated historical operation data as an updated test set, and tuning the model parameters of the target model based on the updated test set until an updated target model that meets the preset requirements is obtained, where the model parameters of the target model include the second percentage.
[0146] Specifically, the embodiment of the present application proposes an adaptive parameter adjustment method, which adjusts the parameters of the target model during the prediction process, so that the third confidence interval generated based on the second percentage can ensure the interval span (corresponding to the fifth time length), and further improve the accuracy rate of the target model. When using the target model of the LSTM algorithm based on the confidence interval for prediction, the most suitable third confidence interval is predicted by adjusting the size of the second percentage.
[0147] Set the fourth time length as an adaptive interval parameter, where the adaptive interval parameter indicates that an adaptive process is triggered every fourth time length to verify the accuracy of the target model. If the accuracy does not reach the range of the preset accuracy threshold, adjust and update the second percentage based on the updated test set, and obtain an updated third confidence interval and an updated fifth time length according to the updated second percentage, so as to obtain an updated target model, and use the updated target model to monitor the application system within the next fourth time length, and predict the availability of the application system within the next fourth time length in real time.
[0148] In an alternative embodiment, the method further includes: using a preset visualization interface to display the operating state of the application system within the preset time period, and sending the predicted abnormal state to a preset associated person to notify the maintenance of the application system. For example, in the visualization interface, display the load state of the application system in the form of a table or a curve graph, etc., and send the predicted abnormal state of the load state to the app of the terminal communication device of the operation and maintenance personnel. The operation and maintenance personnel can also directly refer to the prediction results in the visualization interface to optimize the application system, etc.
[0149] In an alternative embodiment, the system monitoring method provided by the embodiments of the present application first preprocesses the collected data, deletes redundant and invalid data, trains an initial model using an LSTM neural network, and tunes the initial model to obtain a target model, and dynamically optimizes the target model by means of active detection, so as to improve the accuracy of the model in predicting the future change trend of the application system performance. Applying the time series prediction technology based on the LSTM neural network to the application system performance monitoring can not only obtain the overall operating condition of the application system, but also predict the performance change trend of the application system in the future for a period of time, helping the operation and maintenance personnel to take corresponding preventive measures before the system fails.
[0150] Embodiment III
[0151] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the embodiments of the above system monitoring method are implemented, such as Figure 1 S11-S15 shown:
[0152] S11, obtain the historical operation data and real-time operation data of the application system;
[0153] S12, preprocess the historical operation data, and divide the preprocessed historical operation data into a training set and a test set;
[0154] S13. Train a long short-term memory (LSTM) neural network based on the training set to obtain an initial model, test the initial model using the test set, and optimize the model parameters of the initial model according to the test results until a target model that meets the preset requirements is obtained;
[0155] S14. Input the real-time operation data of the application system into the target model for prediction to obtain the predicted operation data of the application system within a preset time period, where the predicted operation data includes the time and probability of the occurrence of a predicted abnormal state;
[0156] S15. Calculate the accuracy of the real-time prediction of the target model, and dynamically optimize the target model according to the accuracy.
[0157] Alternatively, when the computer program is executed by a processor, it implements the functions of each module / unit in the above device embodiment. For example Figure 2 the modules 201-205 in
[0158] The obtaining module 201 is configured to obtain the historical operation data and real-time operation data of the application system;
[0159] The processing module 202 is configured to preprocess the historical operation data and divide the preprocessed historical operation data into a training set and a test set;
[0160] The training module 203 is configured to train a long short-term memory (LSTM) neural network based on the training set to obtain an initial model, test the initial model using the test set, and optimize the model parameters of the initial model according to the test results until a target model that meets the preset requirements is obtained;
[0161] The prediction module 204 is configured to input the real-time operation data of the application system into the target model for prediction to obtain the predicted operation data of the application system within a preset time period, where the predicted operation data includes the time and probability of the occurrence of a predicted abnormal state;
[0162] The optimization module 205 is configured to calculate the accuracy of the real-time prediction of the target model, and dynamically optimize the target model according to the accuracy.
[0163] Embodiment 4
[0164] Refer to Figure 3 shown in the structural schematic diagram of the electronic device provided in Embodiment 3 of the present invention. In a preferred embodiment of the present invention, the electronic device 3 includes a memory 31, at least one processor 32, at least one communication bus 33, and a transceiver 34.
[0165] Those skilled in the art should understand, Figure 3The structure of the illustrated electronic device does not constitute a limitation of the embodiments of the present invention. It can be a bus structure or a star structure. The electronic device 3 may also include more or fewer other hardware or software than shown, or different component arrangements.
[0166] In some embodiments, the electronic device 3 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. Its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits, programmable gate arrays, digital signal processors, and embedded devices, etc. The electronic device 3 may also include a client device, and the client device includes, but is not limited to, any electronic product that can perform human-computer interaction with the client through means such as a keyboard, mouse, remote control, touchpad, or voice control device. For example, personal computers, tablet computers, smart phones, digital cameras, etc.
[0167] It should be noted that the electronic device 3 is only an example. Other existing or future electronic products that can be adapted to the present invention should also be included within the protection scope of the present invention and are incorporated herein by reference.
[0168] In some embodiments, a computer program is stored in the memory 31, and when the computer program is executed by the at least one processor 32, all or part of the steps in the system monitoring method as described are implemented. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium that can be used to carry or store data.
[0169] Furthermore, the computer-readable storage medium may mainly include a storage program area and a storage data area. Among them, the storage program area may store an operating system, application programs required for at least one function, etc.; the storage data area may store data created according to the use of the blockchain node, etc.
[0170] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithm. Blockchain, in essence, is a decentralized database, a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity of the information (anti-counterfeiting) and generate the next block. The blockchain can include the blockchain underlying platform, the platform product service layer, and the application service layer, etc.
[0171] In some embodiments, the at least one processor 32 is the control core (Control Unit) of the electronic device 3, connecting various components of the entire electronic device 3 through various interfaces and circuits. By running or executing programs or modules stored in the memory 31, and by calling data stored in the memory 31, it performs various functions of the electronic device 3 and processes data. For example, when the at least one processor 32 executes the computer program stored in the memory, it implements all or part of the steps of the system monitoring method described in the embodiments of the present invention; or implements all or part of the functions of the system monitoring device. The at least one processor 32 can be composed of integrated circuits. For example, it can be composed of a single packaged integrated circuit, or can be composed of multiple integrated circuits with the same or different functions packaged together, including a combination of one or more central processing units (CPU), microprocessors, digital processing chips, graphics processors, and various control chips, etc.
[0172] In some embodiments, the at least one communication bus 33 is set to enable connection communication between the memory 31 and the at least one processor 32, etc.
[0173] Although not shown, the electronic device 3 may further include a power source (such as a battery) for powering each component. Preferably, the power source can be logically connected to the at least one processor 32 through a power management device, so as to implement functions such as management of charging, discharging, and power consumption management through the power management device. The power source may also include any components such as one or more DC or AC power sources, a recharge device, a power failure detection circuit, a power converter or inverter, a power status indicator, etc. The electronic device 3 may also include various sensors, a Bluetooth module, a Wi-Fi module, a camera device, etc., which will not be elaborated here.
[0174] The integrated units implemented in the form of software functional modules as described above can be stored in a computer-readable storage medium. The above-mentioned software functional modules are stored in a storage medium and include several instructions for causing a computer device (which can be a personal computer, an electronic device, or a network device, etc.) or a processor to execute a part of the methods described in various embodiments of the present invention.
[0175] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned modules is only a logical function division, and there may be other division methods in actual implementation.
[0176] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units. They can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0177] In addition, in various embodiments of the present invention, the functional modules can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of hardware plus software functional modules.
[0178] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above-mentioned exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, it is intended to cover all changes within the meaning and scope of the equivalent elements of the claims in the present invention. Any reference signs in the claims should not be regarded as limiting the claimed rights. In addition, obviously, the word "including" does not exclude other units, and the singular does not exclude the plural. The multiple units or devices described in the specification can also be implemented by one unit or device through software or hardware. The words such as "first" and "second" are used to represent names and do not represent any specific order.
[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A system monitoring method, characterized in that, The method includes: Obtaining the historical operation data and real-time operation data of the application system; Preprocessing the historical operation data, and dividing the preprocessed historical operation data into a training set and a test set; Training a long short-term memory (LSTM) neural network based on the training set to obtain an initial model, testing the initial model using the test set, and tuning the model parameters of the initial model according to the test results until a target model meeting the preset requirements is obtained; Inputting the real-time operation data of the application system into the target model for prediction, to obtain the predicted operation data of the application system within a preset time period, where the predicted operation data includes the time and probability of the occurrence of a predicted abnormal state; Calculating the accuracy of the real-time prediction of the target model, and dynamically optimizing the target model according to the accuracy, including: Calculating the accuracy once every preset fourth time length, including: calculating a second difference between the time of the occurrence of the predicted abnormal state and the time of the actual occurrence of the abnormal state in each fourth time length, and a third difference between the probability of the occurrence of the predicted abnormal state and 1, and taking the product of the second difference and the third difference as the accuracy; Determining whether the accuracy reaches the range of a preset accuracy threshold; When the accuracy does not reach the range of the preset accuracy threshold, optimizing the target model, including: updating the historical operation data using the data in the previous fourth time length, using the updated historical operation data as an updated test set, and tuning the model parameters of the target model based on the updated test set until an updated target model meeting the preset requirements is obtained.
2. The system monitoring method according to claim 1, characterized in that, The preprocessing includes: data integration, data cleaning, data standardization, and data segmentation; The data integration is a process of synthesizing data from multiple sources into one data for storage; The data cleaning includes denoising the historical operation data based on the 3σ criterion of the normal distribution; The data standardization includes normalizing the historical operation data; The data segmentation includes segmenting the historical operation data according to a preset first time length, taking the data within each preset first time length after segmentation as a group of data, and taking the first data arranged in chronological order after any group of data as the label of the any group of data, where the label represents the predicted data of the any group of data.
3. The system monitoring method according to claim 1, wherein The training of the long short-term memory (LSTM) neural network based on the training set to obtain an initial model includes: Grouping the training set according to a preset second time length to obtain multiple sub-training sets; Using each sub-training set to batch-train the LSTM neural network, so that the initial model performs multi-step point value prediction, where the multi-step point value prediction includes: predicting the initial prediction data of the third time length of each sub-training set based on each sub-training set.
4. The system monitoring method according to claim 3, wherein The use of each sub-training set to batch-train the LSTM neural network, so that the initial model performs multi-step point value prediction, includes: Based on any sub-training set and the initial model, perform point value prediction on the abnormal state of the application system to obtain a first predicted point value corresponding to any sub-training set of the application system, where the first predicted point value includes the time point of the occurrence of the abnormal state predicted by any sub-training set; Determining the third time length based on the first predicted point value includes: taking the time length within a preset first percentage before and after the first predicted point value as the third time length, and using the prediction data of the third time length including the first predicted point value as the first confidence interval of the first predicted point value.
5. The system monitoring method according to claim 4, wherein The model parameters of the initial model include the first percentage. Testing the initial model using the test set and optimizing the model parameters of the initial model according to the test results until a target model meeting the preset requirements is obtained, including: Grouping the test set according to the second time length to obtain a plurality of sub-test sets; Inputting any sub-test set into the initial model, and using the initial model to output a second predicted point value corresponding to any sub-test set and a second confidence interval including the second predicted point value; Taking the time point of the actual occurrence of the abnormal state in the second confidence interval as the first actual point value corresponding to any sub-test set, and calculating the first difference between the first actual point value and the second predicted point value as the error corresponding to any sub-test set; Calculating the root mean square error of the errors of all sub-test sets, and determining whether the root mean square error meets the preset requirements, where the preset requirements include: the root mean square error is less than or equal to a preset error threshold; When the root mean square error does not meet the preset requirements, exhaustively update the first percentage; When the root mean square error meets the preset requirements, taking the updated smallest percentage that meets the preset requirements as the second percentage to obtain the target model that meets the preset requirements.
6. The system monitoring method according to claim 5, characterized in that The model parameters of the target model include the second percentage.
7. The system monitoring method according to claim 1, characterized in that, The method further includes: Using a preset visualization interface to display the running state of the application system within the preset time period, and sending the predicted abnormal state to preset associated personnel to notify maintenance of the application system.
8. A system monitoring device for implementing the system monitoring method as described in claim 1, characterized in that, The device includes an acquisition module, a processing module, a training module, a prediction module, and an optimization module: The acquisition module is used to acquire the historical operation data and real-time operation data of the application system; The processing module is used to preprocess the historical operation data and divide the preprocessed historical operation data into a training set and a test set; The training module is used to train a long short-term memory (LSTM) neural network based on the training set to obtain an initial model, test the initial model using the test set, and optimize the model parameters of the initial model according to the test results until a target model meeting the preset requirements is obtained; The prediction module is used to input the real-time operation data of the application system into the target model for prediction to obtain the predicted operation data of the application system within a preset time period, where the predicted operation data includes the time and probability of the occurrence of the predicted abnormal state; The optimization module is used to calculate the accuracy of the real-time prediction of the target model, and dynamically optimize the target model according to the accuracy.
9. An electronic device, characterized in that, The electronic device includes a processor and a memory. When the processor executes the computer program stored in the memory, it implements the system monitoring method according to any one of claims 1 to 7.
10. A computer-readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by the processor, it implements the system monitoring method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Training method and device for fault prediction neural network model
CN112115024A