Cloud platform anomaly detection method for label-free operation and maintenance data
By building a noise prediction network based on Gaussian noise diffusion model, the abnormal detection problem of label-free operation and maintenance data of cloud platform is solved, efficient and accurate abnormal identification is achieved, false alarms are reduced, and real-time monitoring of large-scale cloud platforms is suitable for real-time monitoring.
Patent Information
- Application Number
- CN202510705045.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-07-01
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing cloud platform abnormality detection methods rely on manual annotation thresholds, resulting in late reporting, missed reporting and false reporting, and lack effective processing of unlabeled operation and maintenance data, making it difficult to accurately identify abnormalities.
By training the noise prediction network, use the Gaussian noise diffusion model to extract the normal operating data of the cloud platform, build an abnormality detection model, and use prediction errors to judge abnormalities to avoid dependence on label data.
It realizes efficient abnormal detection of label-free operation and maintenance data, reduces false alarms, improves detection accuracy and efficiency, and is suitable for real-time monitoring of large-scale cloud platforms.
Smart Images

Figure CN120234698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud platform anomaly detection, and more specifically to a cloud platform anomaly detection method for unlabeled operation and maintenance data. Background Art
[0002] Current cloud platforms include infrastructure such as virtual machine servers, physical servers, and cloud desktops, and realize software virtualization management of hardware resources through a management platform, and continuously monitor and collect performance index data such as CPU and memory of each virtual machine in real time. However, there are certain limitations in aspects such as real-time monitoring of cloud resources and anomaly detection of nodes. The current automated anomaly detection stops at a simple method based on fixed thresholds of technical personnel's subjective experience. In the actual operation and maintenance process of the cloud platform, this threshold-based alarm mechanism is divorced from the actual usage scenario of cloud resources, and it is easy to have late reports, missed reports, and false alarms of faults, and the significance of alarms is very limited.
[0003] There are the following problems in IaaS cloud operation and maintenance data. First, the data volume is large. Cloud monitoring data is generally time series data. Due to the scale and complexity of cloud systems, in order to reflect the operating state, the data often has a very high refresh rate. Therefore, the number of time series is extremely large. Second, there is a lack of annotation. The annotation of time series data anomaly detection datasets depends on manual operations, and details the start and end times of anomalies. Due to the high cost of experts, the scale of such datasets is usually small and cannot meet the requirements of some anomaly detection algorithms for the scale of training data. Third, feature engineering is complex. Normal but rare behaviors such as software upgrades and virtual machine drift pose challenges to anomaly detection. Traditional threshold-based anomaly detection models are difficult to accurately identify real anomalies, resulting in many false alarms.
[0004] Therefore, those skilled in the art urgently need a cloud platform anomaly detection method for unlabeled operation and maintenance data to solve the above problems. Summary of the Invention
[0005] In view of this, the present invention provides a cloud platform anomaly detection method for unlabeled operation and maintenance data, which determines whether the operation of the cloud platform is abnormal by comparing the prediction errors of the data.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A cloud platform anomaly detection method for unlabeled operation and maintenance data includes the following steps:
[0008] S1: Preprocessing of training data, including:
[0009] Obtain the operation and maintenance data when the cloud platform is running normally, and divide the collected operation and maintenance data with a historical data window and a future data window respectively to generate data pairs composed of historical data and future data;
[0010] S2: Train a model based on the preprocessed training data, including:
[0011] Perform multiple diffusions on the future data. Each time during diffusion, introduce a preset Gaussian noise parameter, and perform prediction by a noise prediction network to obtain predicted noise and diffusion results.
[0012] Calculate the model loss based on the actual noise and the corresponding historical data, and perform parameter optimization;
[0013] S3: Read the detection data online and perform prediction through the trained noise prediction network. When the error exceeds the preset condition, it is determined as abnormal data.
[0014] Preferably, the S1 specifically includes:
[0015] S11: Add a timestamp t to the operation and maintenance data and organize it into time series data.
[0016] S12: Check the data integrity. When there is missing data, perform data filling.
[0017] S13: Perform normalization calculation on the complete time series data to obtain standard time series data.
[0018] S14: Use windows with lengths of and to intercept the standard time series data X into data windows and , and repeatedly intercept the original data with the window, so that the standard time series data X is converted into a series of sliding data windows and .
[0019] Preferably, the S2 specifically includes:
[0020] S21: Set the number of diffusion steps and the Gaussian noise parameter sequence ;
[0021] S22: Calculate the Gaussian noise required for each step of diffusion and calculate the corresponding diffusion results.
[0022] S23: Perform random sampling in each step of diffusion, calculate the model loss and perform gradient descent and weight update of the noise prediction network.
[0023] The loss function is:
[0024]
[0025] Among them, represents the model loss, represents the mathematical expectation of
[0026] Preferably, in S22, the calculation formula for the diffusion result is:
[0027]
[0028] Among them, represents the result after the nth diffusion, represents the diffusion coefficient of the nth diffusion; represents the original data without diffusion; represents the Gaussian noise added in the nth diffusion.
[0029] Preferably, in S3, the detection data is read online and predicted through a trained noise prediction network. The specific steps include:
[0030] Record the current timestamp and read in real-time operation and maintenance data slices based on the current timestamp;
[0031] Divide the real-time operation and maintenance data slices into known data and current operation and maintenance data according to the timestamp;
[0032] Input the known operation and maintenance data W, the diffusion step embedding vector , and the current operation and maintenance data into the noise prediction network to obtain the predicted noise .
[0033] Predict future data based on the predicted noise and the current operation and maintenance data.
[0034] Preferably, the calculation formula for the future data is:
[0035]
[0036] Among them, is the Gaussian noise parameter corresponding to the nth step; is the variance parameter of the nth diffusion step. When , , when , ; when , , when , represents random noise.
[0037] Preferably, in S3, the judgment of the preset conditions specifically includes:
[0038] Calculate the error matrix based on the read detection data and the predicted results output by the model, and calculate the Mahalanobis distance as the anomaly score.
[0039] Set a score threshold, and when the anomaly score is higher than the score threshold, it is determined as abnormal data.
[0040] Preferably, the calculation methods of the error matrix and the anomaly score include:
[0041] Calculate the prediction deviation , 。
[0042] Reshape into a one-dimensional array: 。
[0043] Calculate 's Mahalanobis distance as the anomaly score at timestamp t :
[0044]
[0045] where is the mean of the elements, is 's covariance matrix.
[0046] Compared with the prior art, the advantages of the present invention are:
[0047] (1) In the training process of the present invention, only the operation and maintenance data of the normal operation of the cloud platform needs to be used for feature extraction, effectively avoiding the dependence on data labels in deep learning and reducing the cost of data processing. The present invention uses a prediction model trained to predict future normal operation and maintenance data, and then compares it with the measured data. By reasonable selection of anomaly scores and thresholds, the accuracy of anomaly detection is ensured.
[0048] (2) The present invention adopts a generative deep learning architecture based on a diffusion model to extract features from ultra-long time-series operation and maintenance data, which can well extract the temporal correlation and the correlation between elements of ultra-long time-series data, and can effectively solve the key problem of high-dimensional data feature extraction in the operation and maintenance process of the cloud platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.
[0050] Figure 1 Flowchart of the cloud platform anomaly detection method for large-scale unlabeled operation and maintenance data proposed by the present invention;
[0051] Figure 2 Schematic diagram of data processing in the present invention;
[0052] Figure 3 Schematic diagram of the noise prediction module in the present invention;
[0053] Figure 4 Schematic diagram of the Transformer layer of the noise prediction module in the present invention;
[0054] Figure 5 Schematic diagram of the training process of the noise prediction module in the present invention;
[0055] Figure 6 Schematic diagram of the online prediction process of operation and maintenance data in the present invention. Specific embodiments
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0057] Aiming at the problems such as difficult acquisition of labeled data and inaccurate anomaly detection in cloud platform anomaly detection, the idea of the present invention is that since the prediction model is trained with normal data and can accurately predict normal operation data, when the cloud platform runs abnormally, the gap between the predicted data and the measured data will increase significantly. Specifically, first, a multivariate data prediction model based on the diffusion model is constructed, and the model can extract the time features and inter-element features of operation and maintenance data. Second, the normal operation data is divided into data pairs by historical data windows and future data windows for the training of the prediction model, so that it can predict the operation data under normal operation conditions. Third, the trained prediction model is used for the online prediction of the cloud platform, and the predicted data is compared with the actual operation data to form an error matrix. Fourth, calculate the Mahalanobis distance of the error matrix as the anomaly score of the cloud platform at this moment, and select an appropriate threshold to output the anomaly or normal flag.
[0058] As Figure 1 , an embodiment of the present invention discloses a cloud platform anomaly detection method for tagless operation and maintenance data, including:
[0059] S1: Preprocess the training data, obtain the operation and maintenance data when the cloud platform is running normally, and divide the collected operation and maintenance data with historical data windows and future data windows respectively to generate data pairs composed of historical data and future data.
[0060] In one embodiment, the specific steps of preprocessing include:
[0061] Step 1.1, Add timestamps to the cloud platform probe data , and organize it into time series data , where each data point represents that the data belongs to timestamp t and , m represents the number of data features at each time point. Specifically, when m = 1, X is a univariate time series.
[0062] Step 1.2, Check the integrity of the data, and fill in the missing data with forward, backward or mean filling to ensure the integrity of the data.
[0063] Step 1.3, Use the maximum-minimum normalization function to perform normalization calculations on each element in the time series data to obtain standard data. The calculation method is:
[0064]
[0065] and represent the maximum / minimum vector norm length in X. Obviously, the value range of the standardized data is .
[0066] Step 1.4, As Figure 2 shown, generate historical data window and prediction window sample pairs by setting a sliding window for subsequent model training. Specifically include: Use windows with lengths of and to intercept the standard time series data X into data windows, and , repeatedly intercept the original data with the window so that the standard time series data X is transformed into a series of sliding data windows and , and use them as data sample pairs to train the prediction model.
[0067] S2: Build the noise prediction module in the training model. Perform multiple diffusions on future data. Each time during diffusion, introduce the preset Gaussian noise parameters and have the noise prediction network make a prediction to obtain the predicted noise and the diffusion result. Calculate the model loss based on the actual noise and the corresponding historical data, and optimize the parameters.
[0068] Among them, the diffusion process includes a forward diffusion process and a reverse inverse diffusion process. The forward diffusion process performs noise addition, and the reverse inverse diffusion process performs noise prediction to gradually remove the noise and restore the data. The forward diffusion process mentioned above is a fixed Markov chain, which does not require training and can be directly calculated. Its role is to generate the results after each step of noise addition for training the noise prediction network. As Figure 3 shown, the noise prediction network includes a 1D convolution, a fully connected network, an activation function, and a Transformer network. As Figure 4 shown, the Transformer network includes two attention networks, one for extracting the temporal correlation of the operation and maintenance data and one for extracting the feature correlation of the operation and maintenance data. The noise prediction network predicts the noise added during the forward diffusion process and takes the historical data, the diffusion result of the previous diffusion step, and the diffusion step after embedding transformation as the input of the noise prediction network, and outputs the noise prediction result in this diffusion step. The specific steps are as follows:
[0069] Step 2.1, set the number of diffusion steps n of the diffusion model, representing the number of execution steps of the forward process; set the number of diffusion steps and the Gaussian noise parameter sequence , , which generally increases as n increases. At the end of the forward process when n = N, the value of should be close to 1.
[0070] Step 2.2, calculate the Gaussian noise parameter for the nth diffusion. Additionally, denote as the diffusion coefficient for the nth diffusion.
[0071] Step 2.3, calculate the Gaussian noise added in the nth diffusion, , with the same shape as ; where represents the multivariate standard Gaussian distribution, represents each element in conforming to the standard Gaussian distribution with a mean of 0 and a variance of 1.
[0072] Step 2.4, calculate the result after the nth diffusion.
[0073] Step 3: Train based on the established noise prediction module; input the preprocessed cloud platform operation data, and train the noise prediction network according to the diffusion process shown in Figure 5 to obtain a diffusion model that can accurately predict the cloud platform operation data.
[0074] Step 3.1: Set the number of iterations ; input the training data pair and ; record the iteration number i = 1.
[0075] Step 3.2: Randomly sample the diffusion step number n from .
[0076] Step 3.3: Randomly sample a set of training data and from the training data pair, where the historical data only serves as the known condition of the model in the diffusion model and does not participate in the forward diffusion process and the reverse inverse diffusion process; mark the diffusion result of and after n diffusions with the subscript n, that is, , and record the original without diffusion as being .
[0077] Step 3.4: Record the model parameters in the noise prediction network as , input and the diffusion step number n into the noise prediction network, and output the predicted noise .
[0078] Step 3.5: According to the Gaussian noise in Step 2.3, construct the loss function of the model for gradient descent and weight update of the noise prediction module:
[0079]
[0080] where, represents the model loss, and represents the mathematical expectation of .
[0081] Step 3.6: Update the iteration number , when , repeat Steps 3.2 - 3.6; when , complete the training and save the parameters of the noise prediction module .
[0082] Step 4, Online Data Reading and Preprocessing: Process the real-time operation and maintenance data of the cloud platform using historical data windows and future data windows with the same shape as the training data, specifically including:
[0083] Step 4.1, Process the online data according to the processing methods of the training data in Steps 1.2 and 1.3.
[0084] Step 4.2, Denote the current timestamp in online detection as t, and read in a total of pieces of real-time operation and maintenance data slices, where the operation and maintenance data with timestamps from to are used as known data, denoted as ; the operation and maintenance data with timestamps from to t is used to predict the operation status (normal / abnormal), denoted as .
[0085] Step 5, Input the online operation data and use the noise prediction network to generate predicted future data, as shown in Figure 6 .
[0086] Step 5.1, Read in the trained noise prediction network and the number of diffusion steps N.
[0087] Step 5.2, Initialize n = N and initialize , where represents a multivariate standard Gaussian distribution, and represents that each element in follows a standard Gaussian distribution with a mean of 0 and a variance of 1. When , represents the predicted value of the operation and maintenance data after the (n + 1)-th denoising.
[0088] Step 5.3, Perform an embedding transformation on n to obtain the transformation result .
[0089] Step 5.4, Input , into the noise prediction network to obtain the predicted noise .
[0090] Step 5.5, Calculate the predicted value of the operation and maintenance data for the current step .
[0091]
[0092] Among them, is the Gaussian noise parameter corresponding to the n-th step; is the variance parameter of the n-th diffusion step. When , , when When ; when When , when When represents random noise.
[0093] Step 5.6, from n = N to n = 1, repeat Steps 5.3 - 5.5 until the output .
[0094] Step 6, compare the predicted future data with the read future data, calculate the error matrix, and then calculate the Mahalanobis distance of the error matrix as the anomaly score, specifically including:
[0095] Step 6.1, calculate the prediction deviation , , and reconstruct into a one-dimensional array: .
[0096] Step 6.2, calculate the Mahalanobis distance of as the anomaly score at timestamp t :
[0097]
[0098] where is the mean of the elements, and is the covariance matrix of .
[0099] Step 6.3, determine whether the system is in an abnormal state according to the anomaly score:
[0100]
[0101] There are many methods for choosing the threshold, such as extreme value theory, optimized F1 score, Peak Over Threshold (POT), etc.
[0102] Step 6.4, when , output the anomaly information to the operation and maintenance personnel for processing.
[0103] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0104] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A cloud platform anomaly detection method for unlabeled operation and maintenance data, characterized in that It includes the following steps: S1: Preprocessing of training data, including: Obtain the operation and maintenance data when the cloud platform is running normally, and divide the collected operation and maintenance data with a historical data window and a future data window respectively to generate data pairs composed of historical data and future data; S2: Train the model according to the preprocessed training data, including: Perform multiple diffusions on the future data. Each time diffusion is performed, introduce a preset Gaussian noise parameter, and predict by the noise prediction network to obtain the predicted noise and the diffusion result; Calculate the model loss with the actual noise and the corresponding historical data as conditions, and perform parameter optimization; S3: Read the detection data online, and predict through the trained noise prediction network. When the error exceeds the preset condition, it is judged as abnormal data.
2. The cloud platform anomaly detection method for tagless operation and maintenance data according to claim 1, wherein, The specific content of S1 includes: S11: Add a timestamp t to the operation and maintenance data and organize it into time series data; S12: Check the data integrity, and fill in the missing data when there is missing data; S13: Perform normalization calculation on the complete time series data to obtain standard time series data; S14: Intercept the standard timing data X as a data window using windows of length and , and repeatedly intercept the original data with the window, so that the standard timing data X is converted into a series of sliding data windows and . and .
3. The cloud platform anomaly detection method for tagless operation and maintenance data according to claim 1, characterized in that The specific content of S2 includes: S21: Set the number of diffusion steps N and the Gaussian noise parameter sequence ; S22: Calculate the Gaussian noise required for each step of diffusion, and calculate the corresponding diffusion result; S23: Perform random sampling in each step of diffusion, calculate the model loss and perform gradient descent and weight update of the noise prediction network; The loss function is: ; Among them, represents the model loss, represents the mathematical expectation of.
4. The cloud platform anomaly detection method for unlabeled operation and maintenance data according to claim 3, wherein In S22, the calculation formula of the diffusion result is: ; Among them, represents the result after the nth diffusion, represents the diffusion coefficient of the nth diffusion; represents the original data without diffusion; represents the Gaussian noise added in the nth diffusion.
5. The cloud platform anomaly detection method for tagless operation and maintenance data according to claim 1, wherein In S3, read the detection data online and predict through the trained noise prediction network. The specific steps include: Record the current timestamp, and read in the real-time operation and maintenance data slice based on the current timestamp; Divide the real-time operation and maintenance data slice into known data and current operation and maintenance data according to the timestamp; Input the known operation and maintenance data W and the diffusion step embedding vector into the noise prediction network , the current operation and maintenance data , to obtain the predicted noise ; Calculate the future data prediction value according to the predicted noise and the current operation and maintenance data.
6. The cloud platform anomaly detection method for unlabeled operation and maintenance data according to claim 5, wherein, The future data The calculation formula is as follows: ; Among them, is the Gaussian noise parameter corresponding to the nth step; is the variance parameter of the diffusion step at the nth step. When is the case, ; when is the case, ; when is the case, ; when is the case, represents random noise.
7. A cloud platform anomaly detection method for unlabeled operation and maintenance data according to claim 1 or 5, characterized in that, In S3, the judgment of the preset condition specifically includes: Calculate the error matrix according to the read detection data and the prediction result output by the model, and calculate the Mahalanobis distance as the anomaly score; Set a score threshold. When the anomaly score is higher than the score threshold, it is determined as abnormal data.
8. The cloud platform anomaly detection method for tagless operation and maintenance data according to claim 7, wherein The calculation methods of the error matrix and the anomaly score include: Calculate the prediction deviation , ; Convert to a one-dimensional array: ; Calculate the Mahalanobis distance as the anomaly score for timestamp t : ; Among them, is the mean value of the elements, is the covariance matrix of 9. The cloud platform anomaly detection method for tagless operation and maintenance data according to claim 7, characterized in that, The score threshold is confirmed by the peak over-threshold method.