An unsupervised anomaly detection and analysis solution based on multivariate time series data
The combination of TCN and VAE with online feedback learning and an anomaly reversal mechanism addresses the challenges of real-time anomaly detection in multi-dimensional streaming data, enhancing scalability and adaptability to concept drift.
Patent Information
- Application Number
- CN202111386406.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-22
AI Technical Summary
When processing multidimensional large-scale streaming data, it is difficult to realize real-time anomaly detection and adaptation concept drift in an online environment, and traditional deep learning methods have problems such as high time and space costs and poor scalability.
Unsupervised anomaly detection method based on TCN and VAE is adopted, combined with deep Bayesian network and expansion convolution, implicit relationships in the time direction are captured, and conceptual drift is adapted through an exception inversion mechanism to perform lightweight and fast anomaly analysis.
It realizes fast and robust anomaly detection in large-scale online streaming data, can adapt to long-term non-stationary timing data, solves the concept drift problem, and improves the adaptability and detection efficiency of abnormal changes.
Smart Images

Figure CN114492826B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of machine learning anomaly detection, and specifically is an unsupervised anomaly detection analysis solution based on multivariate time series stream data. Background Art
[0002] Anomalies are individual observations that are significantly different from the observation body. As a direction in data analysis, anomaly detection faces many challenges: the data is unlabeled streaming data, usually generated by multiple devices, and is a multidimensional vector. It is necessary to use unsupervised methods to learn the statistical distribution between normal data dimensions and historical data. There are concept drifts or change points in the time series data of the streaming state, such as adding parts or changing the process flow, and these changes need to be learned in a timely manner. Sometimes we need to find out where the problem is, which indicator has fluctuated. For this reason, we need an anomaly analysis algorithm with dimensional granularity. And we assume that it should be as small as possible and efficient at the same time.
[0003] Traditional deep learning is usually a batch processing method, in which a training data set is given in advance for offline learning to obtain a trained model, and the test set is also fed into the network in batch form for anomaly detection. This method has high time and space costs and poor scalability because the model needs to be retrained from scratch with new training data.
[0004] Compared with batch processing algorithms, online learning data is in a streaming state that arrives in sequence, and the model parameters are dynamically updated during prediction, thus overcoming the shortcomings of batch learning. In real data analysis, it is more efficient and scalable for large amounts of data, and can handle large amounts of data and fast streaming data. Online learning is usually divided into three categories: online supervised learning with complete feedback, online supervised learning with limited feedback, and online unsupervised learning without any feedback. However, online unsupervised learning methods with feedback have yet to be developed.
[0005] There are currently various online detection methods, including tree-based models such as HS-Tree, Isolation Forest (IF), Extended Isolation Forest (Ex.IF), and Random Cut Forest (RCF). The model size of these methods is usually limited by the tree structure and does not capture long-term dependencies. iForestASD and STORM use a sliding window to detect global outliers relative to the current window in the data stream. For density-based methods, DILOF improves LOF and its variants by adopting a new density-based sampling scheme to summarize the data without prior assumptions about the data distribution. Other streaming methods include RS-Hash, which uses subspace grids and random hashing to detect anomalies. LODA and XSTREAM generate many random projections of the data.
[0006] The random projections in LODA aim to preserve the pairwise distances in the original space, so they cannot provide accurate outlier estimates. STORM, HS-Tree, and XSTREAM perform well on high-dimensional data but are prone to overfitting on low-dimensional data. XSTREAM and Ex.IF generally perform well but cannot capture complex concept drift. The random subspace generation in RS-Hash includes many irrelevant features in the subspace while ignoring relevant features in high-dimensional data. iForestASD, RCF, Kitsune, DILOF, and MSTREAm perform differently on different datasets but do not have universality. Summary of the Invention
[0007] The object of the present invention is to provide an unsupervised anomaly detection and analysis solution based on multivariate time-series stream data in view of the deficiencies of the prior art. The technical problem to be solved by the present invention is to identify, detect, and analyze anomalies in multi-dimensional large amounts of data streams. It is necessary to achieve real-time detection and analysis in an online environment, solve the problem of concept drift, and improve its adaptability to abnormal changes.
[0008] The present invention is a lightweight and fast anomaly detection solution applicable to large-scale online stream data. This solution is based on TCN and VAE, has an online feedback learning algorithm, and an intelligent anomaly analysis algorithm with variable-dimensional granularity. At the same time, it disperses the overall complexity to every corner, so the overall calculation is fast and it can adapt to large-scale online stream data. This solution is implemented by the following steps:
[0009] The present invention includes three stages: an offline training stage, an online detection stage, and an intelligent anomaly analysis stage; the model uses a deep Bayesian network to capture the implicit relationships between multiple sequences and dilated convolutions to capture the implicit relationships in the time direction; when streaming data comes in, it automatically calculates thresholds for anomaly classification; and uses an anomaly inversion mechanism to capture concept drift; then, it analyzes the indefinite number of dimensions where anomalies are most likely to occur serially or in parallel, outputs them, and visualizes them.
[0010] I. Offline pre-training stage:
[0011] Step ①: Preprocess the historical normal data, set the window size, and encapsulate it into batches as an unsupervised training set;
[0012] Step ②: Model training; use the training set obtained in step ② to train the entire model to obtain the required parameters W * -s, b * -s, φ-s, θ-s, and the anomaly scores of the training set;
[0013] In the encoder module, use the sliding window mechanism to batch and window the data batches of the entire training set; the input data is represented as {x t-T+i:t+i |b < i < e}, the window length is T + 1, the batch size is e - b + 1, and each window in the batch is concerned; next, the TCN module with several layers captures the time-dependent patterns of the input time series and outputs h t-T:t with the same dimension as the input X t-T:t ; then calculate the mean vector μ t-T:t of h z and the variance vector σ z , sample the latent space vector z0 through the reparameterization technique, and use the PNF mechanism to iteratively calculate z0 through several layers to obtain the prior probability distribution of the non-Gaussian distribution. The entire encoder part formula is as follows:
[0014]
[0015] The first formula in formula (1) shows the capture of time dependence by the TCN module; the second and third formulas calculate the mean vector μ z and the variance vector σ z of the Gaussian distribution according to h φ , where f φ (h) represents the ReLU activation function; the fourth formula samples according to the variance; the mean vector μ z comes from the linear layer, and the variance vector σ z is generated by the Soft-Plus activation function and a small perturbation ∈; the fifth and sixth formulas show the PNF processing of the latent variable z, z = z K ; represents a Gaussian distribution with a mean of μ z and a variance of σ z ; u z , W z , b z and φ represent the model parameters trained in step 2; through the encoder module, the latent space sequence z K t-T:t will finally be obtained;
[0016] The decoder module p θ (x|z) includes a stochastic convolutional neural network layer and a VAE layer, and this process is formulated as:
[0017]
[0018] Among them, the first equation in formula (2) shows the process of the stochastic convolutional neural network module generating the hidden layer sequence ; the second and third equations in formula (2) are similar to the first and second equations in formula (1), and the only difference lies in the process of generating ; the reconstructed sequence is directly generated from the probability distribution probability distribution instead of being generated from the PlanarNF layer; the same parameters and θ represent the model parameters trained in step 2;
[0019] During the offline training process of this model, the network parameters W * -s, b * -s, φ - s, θ - s are optimized through ELBO; in the offline training set, time series data with a window size of T + 1 is taken; according to the Monte Carlo algorithm, the sampling length is set to L, and the l-th sample is denoted as Therefore, a single loss function can be defined:
[0020]
[0021] Among them, h t = TCN(x t-T:t ), TCN is the paradigm of the stochastic convolutional neural network, expressed as x t The prior probability of follows represents the expectation of x t under the condition of satisfying q φ (z t |h t ); the prior probability of x t follows The first term of formula (3) is the standard multivariate Gaussian normal distribution, the second term represents the Kullback-Leibler (KL) loss, and the third term is a non-negative reconstruction error;
[0022] The entire training process depends on a single loss function, expressed as:
[0023]
[0024] where T + 1 represents the window size, represents the loss of a single observation instance, represents the total loss of the observed data under the entire window;
[0025] Step ③: After training, obtain the anomaly score of the training set according to the reconstruction probability of the training set data, and this score will be used as the input of the POT algorithm;
[0026] For a single reconstructed data at time t It is generated by z t from the probability ; If an anomaly occurs at time t, then the reconstructed data will be different from the original data x t Therefore, according to the reconstruction probability of the reconstructed data detect the anomaly score s t of the current original data x t :
[0027]
[0028] where s t represents the anomaly score of each dimension in the original data x t , log p θ represents the reconstruction probability function, is the generation probability of the input sequence x t-T:t , then the anomaly score of the original data x t can be expressed as the mean of the anomaly scores of each dimension, specifically as follows:
[0029]
[0030] where, AS t represents the anomaly score of the observation instance, m is the number of dimensions, is the anomaly score of the i-th dimension at time point t;
[0031] Summarize the anomaly scores of each data point in the training set (i.e., each observation instance) to obtain the anomaly score Ⅰ of the training set. The anomaly score Ⅰ is used as the POT Init Score input for the POT algorithm for subsequent stages.
[0032] II. Online Detection Phase:
[0033] Step (1): Preprocess the online unlabeled streaming data, set the window size, and encapsulate it into a Block as the test set data.
[0034] Step (2): Feed the test set data into the model sequentially to obtain the prediction scores.
[0035] Replace the training data set in step ③ of the offline phase with each Block in the test set data to obtain the anomaly score Ⅱ corresponding to each Block in the test set data. The anomaly score Ⅰ and the anomaly score Ⅱ are used as the input for the POT algorithm to obtain a prediction threshold.
[0036] Step (3): Feed the anomaly score Ⅱ corresponding to each Block and the prediction threshold into the anomaly inversion mechanism algorithm to determine whether it is a long-term anomaly; if not, output the prediction label; if so, perform feedback training to cope with concept drift.
[0037] Step (4): Repeat steps (3) and (4) until all Blocks in the test set data are processed.
[0038] Use the POT algorithm to set a dynamic threshold, which is a method specifically for modeling extreme values; it selects a threshold to isolate the remaining data, and the values exceeding the threshold are distributed similarly to a generalized Pareto distribution with parameters; the POT algorithm outputs a threshold as the standard for dividing anomalies.
[0039] III. Anomaly Analysis Phase:
[0040] Step 1: Feed the dynamically changing prediction threshold output by the POT algorithm and the anomaly score Ⅱ corresponding to each Block predicted by the intermediate variables in the online detection phase into the anomaly analysis algorithm; the anomaly analysis algorithm selects several anomaly dimensions based on the prediction threshold corresponding to the anomaly score Ⅱ; the output of the anomaly dimension tensor is implemented as follows: sequentially judge whether each dimension anomaly score in the anomaly score Ⅱ is less than the prediction threshold. If it is less, mark it as an anomaly dimension, and thus select anomalies with an indefinite number of dimensions.
[0041] Among them, when the anomaly analysis stage and the online detection stage are in series, the prediction threshold used is the prediction threshold of the entire final test set, and the anomaly score II is the anomaly score II of a certain middle segment; if the parallel method is adopted, the prediction threshold used is the value of the dynamically changing prediction threshold at that moment, and the anomaly score II is the anomaly score II corresponding to the data point being processed at that moment.
[0042] Step 2: The output of the anomaly analysis algorithm in Step 1 is an anomaly dimension tensor, and this anomaly dimension tensor can be visually marked on the original data curve.
[0043] The specific implementation of the anomaly reversal mechanism algorithm is as follows:
[0044] The anomaly score II and the prediction threshold corresponding to each Block are input into the anomaly reversal mechanism algorithm to obtain the prediction labels of all data points in each Block; the prediction labels are adjusted through the point adjustment mechanism to obtain the adjusted prediction labels; it is judged whether all the prediction labels in the adjusted Block are anomalies. If all are anomalies, the anomaly situation of this Block is regarded as concept drift, the anomalies are reversed, and it is added to the training set to realize the dynamic expansion of the training set and perform retraining; after the training is completed, re-prediction is performed, and the newly predicted anomaly score II is extended to the end of the original anomaly score I to form a new anomaly score I; the new anomaly score I and the new anomaly score II are used as the input of the POT algorithm, and the dynamically changed prediction threshold is output; the new prediction label is obtained through the changed prediction threshold and the new anomaly score II.
[0045] Advantages of the present invention:
[0046] The online feedback learning algorithm cleverly utilizes the particularity of the anomaly detection task and uses the anomaly reversal mechanism, which can adapt to long-term non-stationary time series data, solves the problems of concept drift and change points, and makes the entire model more robust. The intelligent anomaly analysis mechanism that uses the characteristics of the reference model is used in the anomaly analysis stage of the solution, without making any assumptions, and extracts anomalies with an indefinite number of dimension granularities, with extremely high efficiency. It can better help engineers analyze the anomaly behaviors of various indicators and their correlations, and further solve the anomalies targeted. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 is the overall solution structure diagram of the present invention.
[0048] Figure 2 is the structure diagram of the detection model.
[0049] Figure 3 is the algorithm description for implementing the anomaly reversal mechanism.
[0050] Figure 4It is an algorithm description for realizing intelligent anomaly analysis.
[0051] Figure 5 It is a comparison table of AUC scores with eleven other existing models under seven publicly available datasets.
[0052] Figure 6 It is a comparison graph of AUC results using the anomaly inversion mechanism algorithm and not using this algorithm under seven publicly available datasets.
[0053] Figure 7 It is a schematic diagram of the relationship between retrain times and F1-score for the SMD dataset. Specific implementation manner
[0054] The present invention is further described below in conjunction with the accompanying drawings and specific implementation steps:
[0055] An unsupervised anomaly detection and analysis solution based on multivariate time series data, comprising the following steps:
[0056] I. Offline pre-training stage:
[0057] Step 1: Preprocess the historical normal data, set the window size, and encapsulate it into batches as an unsupervised training set;
[0058] As Figure 1 shown, it shows the overall solution structure of the present invention. The data processing part is responsible for packaging the original data, creating a sliding window, and splitting and packing the data into batches. Its window size is T + 1, and the test data X at time t t depends on the X t-T:t sequence. Where X t-T:t represents the test value sequence from time t - T to time t.
[0059] Step 2: Model training; use the training set obtained in Step 1 to train the entire model to obtain the required parameters W * -s, b * -s, φ - s, θ - s and the anomaly scores of the training set;
[0060] Before deploying and using the model, we need to first design the model and train it using the historical dataset. This process is offline and has low requirements for real-time performance, and requires capturing specific patterns of the data as accurately as possible.
[0061] The basic architecture of UADS-Model is as Figure 2 shown. The details of the encoder module q φ (z∣x) are as Figure 2As shown on the left. In the encoder module, first, the sliding window mechanism is used to batch and window the data of the entire training set; the input data is represented as {x t-T+i:t+i |b < i < e}, where the range of the subscript i is (b, e), the window length is T + 1, the batch size is e - b + 1, and each window in the batch is concerned; next, the TCN module with several layers captures the temporal dependence pattern of the input time series and outputs h t-T:t with the same dimension as the input x t-T:t . This well solves the problem of temporal dependence of multivariate time series, and due to the characteristics of TCN convolution, this step can be carried out very fast with few parameters. Then, the mean vector μ t-T:t and the variance vector σ z of h z are calculated, and the latent space vector z0 is sampled through the reparameterization technique, whose dimension is preset in advance and has the assumption of satisfying the Gaussian distribution. For generalization, we iteratively calculate z0 through several layers using the PNF mechanism to obtain the prior probability distribution of the non-Gaussian distribution. The entire encoder part formula is as follows:
[0062]
[0063] The first equation in formula (1) shows the capture of temporal dependence by the TCN module; the second and third equations calculate the mean vector μ z and the variance vector σ z of the Gaussian distribution according to h, where f φ (h) represents the ReLU activation function; the fourth equation samples according to the variance; the mean vector μ z comes from the linear layer, and the variance vector σ z is generated by the Soft-Plus activation function and a small perturbation ∈; the fifth and sixth equations show the PNF processing of the latent variable z, z = z K ; represents the Gaussian distribution with mean μ z and variance σ z ; u z , W z , b z and φ represent the model parameters trained in step 2; through the encoder module, the latent space sequence z K t-T:t will finally be obtained.
[0064] Figure 2 The right half of Figure 2 represents the details of the decoder module p θ (x|z), including the stochastic convolutional neural network layer and the VAE layer. The process is similar to that of the encoder module, and this process is formulated as:
[0065]
[0066] Among them, the first equation in formula (2) shows the process of generating the hidden layer sequence by the random convolutional neural network module The second and third equations in formula (2) are similar to the first and second equations in formula (1). The only difference lies in the process of generating ; The reconstructed sequence is directly generated from the probability distribution probability distribution instead of being generated from the PlanarNF layer; The same parameters and θ represent the model parameters completed in training in step 2;
[0067] During the offline training process of this model, the network parameters W * -s, b * -s, φ - s, θ - s are optimized by ELBO; In the offline training set, time series data with a window size of T + 1 is taken; According to the Monte Carlo algorithm, the sampling length is set to L, and the l-th sample is denoted as Therefore, a single loss function can be defined:
[0068]
[0069] where h t = TCN(x t-T:t ), TCN is the paradigm of the random convolutional neural network, expressed as x t The prior probability of follows represents the expectation of x t under the condition of satisfying the q φ (z t |h t ) distribution; The prior probability of x t follows The first term of formula (3) is the standard multivariate Gaussian normal distribution, the second term represents the Kullback-Leibler (KL) loss, and the third term is a non-negative reconstruction error;
[0070] The entire training process can depend on a single loss function, expressed as:
[0071]
[0072] where T + 1 represents the window size, Represents the loss of a single observation instance, Represents the total loss of the observed data under the entire window.
[0073] Step 3: After training, obtain the anomaly score of the training set according to the reconstruction probability of the training set data, and this score will be used as the input of the POT algorithm;
[0074] For a single reconstructed data at time t It is generated by z t From the probability If an anomaly occurs at time t, then the reconstructed data Will be different from the original data x t Therefore, according to the reconstructed data Detect the anomaly score s t Of the current original data x t :
[0075]
[0076] Among them, s t Represents the anomaly score of each dimension in the original data x t Log p θ Represents the reconstruction probability function, Is the generation probability of the input sequence x t-T:t Of, Then the anomaly score of the original data x t Can be expressed as the average of the anomaly scores of each dimension, specifically as follows:
[0077]
[0078] Among them, AS t Represents the anomaly score of the observation instance, m is the number of dimensions, Is the anomaly score of the i-th dimension at time point t;
[0079] Aggregate the anomaly scores of each data point (i.e., each observation instance) in the training set to obtain the anomaly score Ⅰ of the training set. The anomaly score Ⅰ is used as the POT Init Score of the POT algorithm for subsequent stages.
[0080] Set the dynamic threshold using the POT algorithm, which is a method specifically for modeling extreme values. It selects a threshold to isolate the remaining data, and the distribution of values exceeding the threshold is similar to the generalized Pareto distribution with parameters. The algorithm outputs a threshold (POT Init Score) as the criterion for us to divide anomalies. The larger the anomaly score, the smaller the possibility that the observed instance is an anomaly, and vice versa. Therefore, if it is less than this threshold, it is treated as an outlier, and vice versa. We use the scores obtained in step 2 as the set of anomaly scores for the training set and use it as the input to the POT algorithm.
[0081] II. Online detection phase:
[0082] Step 1: Preprocess the online unlabeled streaming data, set the window size, and encapsulate it into Blocks as the test set data; similar to step 1 in the offline pre-training phase, we encapsulate the data stream obtained in the online environment into individual Blocks and obtain scores for each Block for use in the subsequent anomaly reversal algorithm
[0083] Step 2: Feed the test set data into the model one by one to obtain the prediction scores;
[0084] Replace the training dataset in step (3) of the offline phase with each Block in the test set data to obtain the corresponding anomaly score II for each Block in the test set data. The anomaly score I and the anomaly score II are used together as the input to the POT algorithm to obtain a prediction threshold;
[0085] Step 3: Feed the anomaly score II corresponding to each Block and the prediction threshold into the anomaly reversal mechanism algorithm to determine whether it is a long-term anomaly; if not, output the prediction label; if so, perform feedback training to cope with concept drift;
[0086] Step 4: Repeat steps 3 and 4 until all Blocks in the test set data are processed;
[0087] Set the dynamic threshold using the POT algorithm, which is a method specifically for modeling extreme values; it selects a threshold to isolate the remaining data, and the distribution of values exceeding the threshold is similar to the generalized Pareto distribution with parameters; the POT algorithm outputs a threshold as the criterion for dividing anomalies;
[0088] In the context of industrial Internet of Things, the processing of streaming data needs to be fast, and as mentioned before, it is necessary to handle the undetected anomalies that occur in the data after online deployment, that is, the anomalies that do not appear in the offline training set, or the concept drift caused by persistent situations that occur during online training. Based on the characteristics of the model, we designed the most efficient and model-fitting algorithm to implement the anomaly reversal mechanism, as shown inFigure 3 The detailed description of the algorithm for this step is as follows:
[0089] The abnormal score Ⅱ corresponding to each Block and the prediction threshold are input into the abnormal inversion mechanism algorithm to obtain the prediction labels of all data points in each Block; the prediction labels are adjusted through the point adjustment mechanism to obtain the adjusted prediction labels; it is judged whether all the prediction labels in the adjusted Block are abnormal. If they are all abnormal, the abnormal condition of this Block is regarded as concept drift, the abnormality is reversed, and it is added to the training set to realize the dynamic expansion of the training set and retraining; after the training is completed, prediction is carried out again, and the newly predicted abnormal score Ⅱ is extended to the end of the original abnormal score Ⅰ to form a new abnormal score Ⅰ; the new abnormal score Ⅰ and the new abnormal score Ⅱ are used as the input of the POT algorithm, and the dynamically changed prediction threshold is output; the new prediction labels are obtained through the changed prediction threshold and the new abnormal score Ⅱ.
[0090] The time complexity of the entire algorithm is low, and the number of retrainings is positively correlated with the number of occurrences of concept drift. The algorithm adds the scores of concept drift points to the training set scores and recalculates the threshold in real time, realizing the dynamic expansion of the training set.
[0091] III. Abnormal analysis stage:
[0092] Step 1: The dynamically changing prediction threshold output by the POT algorithm and the abnormal score Ⅱ corresponding to each Block predicted by the intermediate variable in the online detection stage are sent into the abnormal analysis algorithm; the abnormal analysis algorithm selects several abnormal dimensions based on the prediction threshold corresponding to the abnormal score Ⅱ; the output of the abnormal dimension tensor is realized as follows: successively judge whether the abnormal scores of each dimension in the abnormal score Ⅱ are less than the prediction threshold. If they are less, they are marked as abnormal dimensions, and thus an indefinite number of abnormal dimensions are selected; as Figure 4 shown. Since the input we use has been obtained in the previous two stages, and the implementation of the algorithm fits the model, dispersing the complexity to other stages, only the size needs to be judged, and the actual running time is extremely short.
[0093] Among them, when the abnormal analysis stage is in series with the online detection stage, the prediction threshold used is the prediction threshold of the entire final test set, and the abnormal score Ⅱ is the abnormal score Ⅱ of a certain intermediate segment; if the parallel method is adopted, the prediction threshold used is the value of the dynamically changing prediction threshold at that moment, and the abnormal score Ⅱ is the abnormal score Ⅱ corresponding to the data points processed at that moment;
[0094] Step 2: The output of the abnormal analysis algorithm in Step 1 is the abnormal dimension tensor, and this abnormal dimension tensor can be visually marked on the original data curve.
[0095] Figure 5 The table in shows the AUC scores of OOA-UADS and existing baseline and state-of-the-art online anomaly detection models for streaming data, including the scores of each model on each dataset, the average score (Avg.) of each model, and the rank (Rank) of OOA-UADS on each dataset. Among the seven datasets, OOA-UADS achieved the best results in five datasets and ranked first in terms of the average score. Our model has made a significant improvement in the AUC score compared to the baseline method. On the KDD99 dataset with a large amount of data, our model performed poorly because although the threshold for calculating the score changes dynamically, there is only one threshold for classification.
[0096] The goal of random projection in LODA preserves the pairwise distances in the original space, so it cannot provide accurate outlier estimates. STORM, HS-Tree, and XSTREAM perform well on high-dimensional data but are prone to overfitting on low-dimensional data. XSTREAM and Ex.IF generally perform well but cannot capture complex concept drift. The random subspace generation of RS-Hash incorporates many irrelevant features into the subspace while ignoring relevant features in high-dimensional data. iForestASD, RCF, Kitsune, DILOF, and MSTREAm perform differently on different datasets but do not have universality.
[0097] In summary, our model is applicable to most datasets with different dimensional sizes and anomaly ratios and has strong universality.
[0098] Figure 6 This is our ablation experiment, mainly regarding the retraining step in the online deployment algorithm. The main improvement of our algorithm is to endow the deep learning model with the ability to process streaming data and enable continuous learning. We conducted ablation experiments on all seven online streaming datasets. The results show that for the datasets KDD99, Cardio, Mamm., and Cover, our algorithm successfully improved the AUC value, that is, it performed better. For the datasets Sat., Sat.-2, and Pima, since the retraining conditions were not met during the online detection process and there was no retraining, the results remained unchanged.
[0099] Figure 7 shows the relationship between the retrain times and the F1-score of the SMD dataset. We also plotted the trend line and can see that it is an upward trend. We also counted the relationship between the number of retraining times and the score to show that our retraining is effective.
[0100] The model of the present invention deconstructs and reconstructs multivariate time series streaming data through a temporal convolutional network and a variational autoencoder to learn normal patterns. When performing online streaming data detection, it is encapsulated into blocks and fed into the model to obtain scores and classification labels. According to the proposed "anomaly inversion mechanism", the problem of concept drift in online anomaly detection is solved, and the model is retrained to dynamically update the classification threshold. This mechanism improves the accuracy of handling anomalies in online streaming data. Subsequently, intelligent anomaly analysis at the dimension granularity can be performed in parallel or serially to generate anomaly analysis reports with indefinite dimensions. The present invention can perform intelligent anomaly identification, detection, and analysis on multi-dimensional, complex detection metrics and rapidly growing data streams, with a complete process, providing guarantee for the safe operation of the system.
Claims
1. An unsupervised anomaly detection and analysis solution based on multivariate time-series data, characterized in that The implementation of this method includes three stages: the offline training stage, the online detection stage, and the intelligent anomaly analysis stage; the model uses a deep Bayesian network to capture the implicit relationships between multiple sequences and dilated convolutions to capture the implicit relationships in the time direction; when the streaming data comes, it automatically calculates the threshold for anomaly classification; and uses an anomaly inversion mechanism to capture concept drift; then it analyzes the indefinite number of dimensions most likely to have anomalies serially or in parallel, outputs them, and visualizes them; The specific implementation of this method is as follows: I. Offline pre-training stage: Step 1: Preprocess the historical normal data, set the window size, and encapsulate it into batches as an unsupervised training set; Step 2: Model training; use the training set obtained in Step 1 to train the entire model to obtain the required parameters W * -s, b * -s, φ-s, θ-s, and the anomaly scores of the training set; In the encoder module, the sliding window mechanism is used to batch and window the data of the entire training set; the input data is represented as {x t-T+i:t+i | b < i < e}, the window length is T + 1, the batch size is e - b + 1, and each window in the batch is concerned; next, the TCN module with several layers captures the temporal dependence pattern of the input time series and outputs h t-T:t with the same dimension as the input x t-T:t ; then the mean vector μ t-T:t and variance vector σ z of h z are calculated, and the latent space vector z0 is sampled through the reparameterization technique. z0 is iteratively calculated through several layers using the PNF mechanism to obtain the prior probability distribution of the non-Gaussian distribution. The entire formula of the encoder part is as follows: The first equation in formula (1) shows the capture of time dependence by the TCN module; the second and third equations calculate the mean vector μ of the Gaussian distribution based on h z and the variance vector σ z , where f φ (h) represents the ReLU activation function; the fourth equation samples according to the variance; the mean vector μ z comes from the linear layer, and the variance vector σ z is generated by the Soft-Plus activation function and a small perturbation ∈; the fifth and sixth equations show the PNF processing of the latent variable z, z = z K ; represents a Gaussian distribution with mean μ z and variance σ z ; u z , W z , b z and φ represent the model parameters trained in step 2; through the encoder module, the latent space sequence z K t-T:t ; Decoder module p θ (x|z) includes a random convolutional neural network layer and a VAE layer, and this process is formulated as: Among them, the first equation in formula (2) shows the process of generating the hidden layer sequence by the stochastic convolutional neural network module ; the second and third equations in formula (2) are similar to the first and second equations in formula (1), and the only difference lies in the process of generating ; the reconstructed sequence is directly generated from the probability distribution , rather than from the PlanarNF layer; the same parameters and θ represent the model parameters completed in training in step 2; During the offline training process of the model, the network parameters W * -s, b * -s, φ-s, θ-s are optimized by ELBO; in the offline training set, time series data with a window size of T+1 is taken; according to the Monte Carlo algorithm, the sampling length is set to L, and the first sample is denoted as Therefore, a single loss function can be defined: where h t = TCN(X t-T:t ), TCN is the paradigm of the random convolutional neural network, expressed as x t 's prior probability follows represents the expectation of x t under the condition of satisfying q φ (z t |h t ); x t 's prior probability follows The first term of formula (3) is the standard multivariate Gaussian normal distribution, the second term represents the Kullback-Leibler (KL) loss, and the third term is a non-negative reconstruction error; The entire training process depends on a single loss function, expressed as: where T + 1 represents the window size, represents the loss of a single observation instance, represents the total loss of the observed data under the entire window; Step 3: After training, obtain the anomaly scores of the training set based on the reconstruction probability of the training set data, and this score will be used as the input of the POT algorithm; For a single reconstructed data at time t It is generated by z t from the probability ; if an anomaly occurs at time t, then the reconstructed data will be different from the original data x t Therefore, the anomaly score s of the current original data x is detected according to the reconstruction probability of the reconstructed data t : t Among them, s t represents the anomaly score of each dimension in the original data x t and log p θ represents the reconstruction probability function is the generation probability of the input sequence x t-T:t and then the anomaly score of the original data x t can be expressed as the mean of the anomaly scores of each dimension, as follows: where AS t represents the anomaly score of the observation instance, m is the number of dimensions, is the anomaly score of the i-th dimension at time point t; Summarize the anomaly scores of each data point in the training set to obtain the anomaly score Ⅰ of the training set. The anomaly score Ⅰ is used as the POTInitScore of the POT algorithm for use in subsequent stages; II. Online detection stage: Step 1: Preprocess the online unlabeled streaming data, set the window size, and encapsulate it into Blocks as the test set data; Step 2: Send the test set data into the model one by one to obtain the prediction scores; Replace the training data set in step (3) of the offline stage with each Block in the test set data to obtain the anomaly score Ⅱ corresponding to each Block in the test set data. The anomaly score Ⅰ and the anomaly score Ⅱ are used as the input of the POT algorithm to obtain a prediction threshold; Step 3: Send the anomaly score Ⅱ corresponding to each Block and the prediction threshold into the anomaly inversion mechanism algorithm to determine whether it is a long-term anomaly; if not, output the prediction label; if so, perform feedback training to cope with concept drift; Step 4: Repeat steps 3 and 4 until all Blocks in the test set data are processed; Use the POT algorithm to set a dynamic threshold, which is a method specifically for modeling extreme values; it selects a threshold to isolate the remaining data, and the distribution of values exceeding the threshold is similar to a generalized Pareto distribution with parameters; the POT algorithm outputs a threshold as the criterion for dividing anomalies; III. Anomaly analysis stage: Step 1: Feed the dynamically changing prediction threshold output by the POT algorithm and the anomaly score II corresponding to each Block predicted by the intermediate variable in the online detection stage into the anomaly analysis algorithm; this anomaly analysis algorithm selects a number of anomaly dimensions based on the prediction threshold corresponding to the anomaly score II; the output of the anomaly dimension tensor is implemented as follows: sequentially determine whether each dimension anomaly score in the anomaly score II is less than the prediction threshold. If it is less, mark it as an anomaly dimension, and thus select anomalies of an indefinite number of dimensions; Among them, when the anomaly analysis stage is in series with the online detection stage, the prediction threshold used is the final prediction threshold of the entire test set, and the anomaly score Ⅱ is the anomaly score Ⅱ of a certain middle segment; if it is in parallel, the prediction threshold used is the value of the dynamically changing prediction threshold at that moment, and the anomaly score Ⅱ is the anomaly score Ⅱ corresponding to the data points processed at that moment; Step 2: The output of the anomaly analysis algorithm in step 1 is an anomaly dimension tensor, and this anomaly dimension tensor can be visually marked on the original data curve.
2. The unsupervised anomaly detection analysis and solution method based on multi-source time-series data according to claim 1, characterized in that The specific implementation of the anomaly inversion mechanism algorithm is as follows: The abnormal score II corresponding to each Block and the prediction threshold are input into the abnormal inversion mechanism algorithm to obtain the prediction labels of all data points in each Block; the prediction labels are adjusted through the point adjustment mechanism to obtain the adjusted prediction labels; it is judged whether all the prediction labels in the adjusted Block are abnormal. If all are abnormal, the abnormal condition of this Block is regarded as concept drift, the abnormality is reversed, and it is added to the training set to realize the dynamic expansion of the training set and perform retraining; after the training is completed, a re-prediction is performed, and the newly predicted abnormal score II is extended to the end of the original abnormal score I to form a new abnormal score I; the new abnormal score I and the new abnormal score II are used as the input of the POT algorithm to output the dynamically changed prediction threshold; the new prediction labels are obtained through the changed prediction threshold and the new abnormal score II.
Citation Information
Patent Citations
Large-scale multivariate time series data anomaly detection method oriented to cloud environment
CN112784965A
Lightweight unsupervised anomaly detection method based on multivariate time series data analysis
CN113159163A