A Multivariate Time-Series Data Anomaly Detection Method Based on VAE and Associated Differences

By introducing VAE-based correlation difference mechanism and self-attention mechanism in time series data anomaly detection, combined with the dual indicators of reconstruction error and correlation difference, the problem of insufficient detection accuracy and robustness of traditional methods in high-dimensional and nonlinear data is solved, and more efficient and accurate abnormal detection is achieved.

CN119783010BActive Publication Date: 2025-06-24ZHEJIANG SCI-TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510273266.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-24
Estimated Expiration
2045-03-10

AI Technical Summary

Technical Problem

传统的时间序列数据异常检测方法在面对海量、高维、非线性及动态变化的数据时,检测精度不高、鲁棒性不足以及实时性差。

Method used

采用基于变分自编码器(VAE)与关联差异的多元时序数据异常检测方法,结合自注意力机制和重构误差与关联差异的双重指标,提升检测精度和鲁棒性。

Benefits of technology

显著提高了异常检测的准确性和鲁棒性,避免了传统方法中的梯度爆炸问题和顺序处理的低效率,能够更好地捕捉时间序列中的动态特征和异常模式。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119783010B_ABST
    Figure CN119783010B_ABST
Patent Text Reader

Abstract

The present invention discloses a multivariate time series data anomaly detection method based on VAE and correlation difference. By introducing correlation difference, the model can effectively capture the essence of abnormal changes in time series. The calculation of correlation difference enables the model to quantify the change in the degree of association between variables, thereby more accurately distinguishing normal and abnormal states. At the same time, the present invention uses a variational autoencoder to reconstruct the time series, which can learn the distribution law of the data and find out the factors that best represent the essential characteristics of the data. By further learning the correlation difference of the reconstructed time series, the anomaly discrimination ability of the model of the present invention is enhanced and generalized. In addition, the present invention combines mean squared error, KL divergence and correlation difference as the criteria for judging anomalies, greatly improving the anomaly detection ability of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data anomaly detection, and particularly relates to a multi-temporal data anomaly detection method based on VAE and associated differences. Background Art

[0002] With the rapid development of industrial automation, Internet of Things, big data, and artificial intelligence technologies, time series data has been increasingly widely used in various fields such as industrial equipment monitoring, spacecraft status detection, network security, financial fraud detection, and medical monitoring. These data often reflect the system state in a continuous time series form, and their anomaly detection is of crucial significance for ensuring system security, improving operation efficiency, and reducing maintenance costs. However, traditional anomaly detection methods face many problems such as low detection accuracy, insufficient robustness, and poor real-time performance when dealing with massive, high-dimensional, non-linear, and dynamically changing time series data.

[0003] Traditional statistical methods were the earliest means applied to time series anomaly detection. For example, the autoregressive integrated moving average (ARIMA) model models the autocorrelation, trend, and seasonality in the time series, predicts future data, and detects anomalies in the prediction errors. Although the ARIMA model can capture linear relationships when the data is stationary and has obvious periodicity, it has high requirements for data stationarity, cannot effectively handle non-linear changes and sudden anomalies, and is easily restricted when dealing with long-term dependence relationships. On the other hand, probability distribution-based models such as the Gaussian mixture model (GMM) assume that the data is composed of a mixture of multiple Gaussian distributions, calculates the probability of data points through parameter estimation, and thus judges anomalies. However, in high-dimensional data processing and multi-modal data scenarios, GMM often faces difficulties in determining the number of Gaussian distributions and high computational complexity, which limits its application effect in complex dynamic systems.

[0004] In recent years, with the continuous development of machine learning and deep learning technologies, researchers have begun to try to introduce methods such as support vector machines (SVM), decision trees, and random forests to model time series anomaly detection. SVM distinguishes between abnormal and normal data by constructing an optimal classification hyperplane. Although it performs well in small sample and non-linear scenarios, its performance highly depends on the selection of kernel functions and parameter tuning, and there are computational bottlenecks in large-scale data processing; random forests and other ensemble learning methods improve the detection robustness by integrating multiple decision trees, but are prone to overfitting when the data features are relatively complex or there is a lot of noise, which in turn affects the detection effect. Although these methods have improved the performance of anomaly detection to a certain extent, they are still insufficient for capturing long-term dependence relationships, local dynamic changes, and global distribution features in time series.

[0005] In recent years, deep learning methods have made great progress in the field of anomaly detection. Recurrent neural networks (RNNs) and their improved forms such as long short-term memory networks (LSTMs) and gated recurrent units (GRUs) are widely used in time series data analysis. They can capture long-term dependencies in the sequence through internal memory units and achieve anomaly detection through sequence prediction. However, RNNs and their variants often suffer from the problems of vanishing gradients or exploding gradients when dealing with ultra-long sequences. At the same time, due to their sequential processing characteristics, the training and inference efficiency of the model is relatively low. Convolutional neural networks (CNNs), although outstanding in automatically extracting local features, are insufficient in modeling long-distance dependencies and are difficult to fully reflect the complex dynamic patterns in time series. In addition, some anomaly detection methods based on generative adversarial networks (GANs) have also been proposed, but these methods often focus on generating high-quality reconstructed images or sequences, and there is no unified and robust metric for the discrimination criteria of anomaly detection.

[0006] Against this background, in recent years, some studies have begun to explore anomaly detection methods based on variational autoencoders (VAEs). By introducing latent variables and using the maximum evidence lower bound (ELBO) for parameter estimation, VAEs can learn the latent probability distribution of the input data, thereby generating significant anomaly reconstruction errors different from the original data during the reconstruction process. Although VAEs have great advantages in capturing the global distribution characteristics of data, when solely relying on reconstruction errors for anomaly determination, they often ignore the local temporal dependencies and dynamic change characteristics in time series data, which may lead to less than ideal detection accuracy in cases where data anomalies are relatively concealed or local anomaly fluctuations are small. Summary of the Invention

[0007] In view of the above, the present invention provides a multivariate time series data anomaly detection method based on VAE and correlation difference. This method comprehensively utilizes the global distribution learning ability of VAE, the local dependence capture ability of the self-attention mechanism, and the anomaly determination method based on correlation difference, significantly improving the detection accuracy and robustness.

[0008] A multivariate time series data anomaly detection method based on VAE and correlation difference includes the following steps:

[0009] (1) Obtain multivariate time series data and preprocess it to obtain a clean data set suitable for model training, and then divide the data set into a training set and a test set;

[0010] (2) Establish an anomaly detection model based on VAE and correlation difference mechanism, which includes:

[0011] A correlation difference module, used to capture the time dependencies in multivariate time series data, calculate the correlation difference of multivariate time series data, and serve as part of the anomaly detection index;

[0012] A reconstruction module, which is used to map the feature vectors output by the correlation difference module to the latent space to capture their latent features, and then randomly sample in the latent space to obtain reconstructed time series data;

[0013] A stochastic correlation difference module, which is used to capture the time dependence in the reconstructed time series data, calculate the correlation difference of the reconstructed time series data, so as to further improve the anomaly detection and recognition ability of the model;

[0014] (3) Use the multivariate time series data of the training set to train the above model, and use the reconstruction error and the correlation difference loss as the objectives of model training;

[0015] (4) Input the multivariate time series data of the test set into the trained model to obtain the corresponding correlation difference, and calculate the anomaly score of each time point in the multivariate time series data according to the correlation difference and the reconstruction error;

[0016] (5) Compare the anomaly score of each time point in the multivariate time series data with the set threshold. If it exceeds the threshold, it is determined as an anomaly, so as to identify the anomaly time points in the multivariate time series data.

[0017] Further, the preprocessing of the multivariate time series data in the step (1) includes denoising, normalization and missing value filling.

[0018] Further, the multivariate time series data needs to be positionally encoded before being input into the model, that is, the multivariate time series data is element-wise added to the position encoding vector generated based on sine and cosine functions to encode the order information in the data; this approach ensures that the model can understand the relative positions of the elements in the sequence.

[0019] Further, the correlation difference module is composed of multiple cascaded correlation difference layers. The correlation difference layer first passes the input x through the Anomaly-Attention (anomaly attention mechanism), then adds the output of the Anomaly-Attention to the input x and inputs it to the feed-forward neural network after layer normalization processing, and finally adds the input and output of the feed-forward neural network and outputs it as the output of the correlation difference layer after layer normalization processing.

[0020] Further, the reconstruction module first uses the dimensionality reduction convolutional layer to map the feature vectors output by the correlation difference module to the latent space to fit the mean and variance of the approximate posterior distribution of the latent variable, so as to obtain the approximate posterior distribution of the latent variable, that is, the latent feature. Furthermore, the dimensionality increase convolutional layer is used to randomly sample in the latent space using the reparameterization technique to obtain the final latent variable, which is used as the reconstructed time series data.

[0021] Further, the stochastic correlation difference module is composed of multiple cascaded stochastic correlation difference layers. The stochastic correlation difference layer first passes the input \(x'\) through Anomaly - Attention, then adds the output of Anomaly - Attention to the input \(x'\) and inputs it to a feed - forward neural network after layer normalization processing. Finally, the input and output of the feed - forward neural network are added and input to a fully - connected layer after layer normalization processing, and the output of the fully - connected layer is used as the output of the stochastic correlation difference layer.

[0022] Further, in step (4), the anomaly score at each time point in the multivariate time - series data is calculated by the following formula;

[0023]

[0024] where: \(X\) represents the multivariate time - series data, represents the reconstructed time - series data, is the anomaly score sequence of \(X\), and each element value in the sequence corresponds to the anomaly score at each time point in the multivariate time - series data; \(AssDis(P,S;X)\) is the correlation difference of the multivariate time - series data \(X\), is the correlation difference of the reconstructed time - series data The reconstruction error consists of two parts, namely the mean - square error between \(X\) and and the KL - divergence \(KLD\), represents element - wise multiplication, and \(Softmax(\ )\) represents the Softmax function.

[0025] Further, the expression of the correlation difference \(AssDis(P,S;X)\) is as follows:

[0026]

[0027] where: represents the KL - divergence of relative to , represents the KL - divergence of relative to , \(L\) is the number of layers of the correlation difference module, represents the distribution of the \(i\) - th row in the prior correlation matrix \(P\) l , represents the distribution of the \(i\) - th row in the sequence correlation matrix \(S\) l , and \(N\) is the sequence length of the multivariate time - series data.

[0028] Further, the expressions of the prior correlation matrix \(P\) l and the sequence correlation matrix \(S\) l are as follows:

[0029]

[0030]

[0031] Wherein: Q, K, and σ are the query vector matrix, the key vector matrix, and the scale parameter matrix respectively, , , , X l is the output of the multivariate time series data X input to the l-th layer in the association difference module, , , correspond to the weight parameter matrices of Q, K, and σ in the l-th layer respectively, and σ i is the value of the i-th element in the scale parameter matrix σ, and d model is X l 's feature dimension, T represents transpose, and Rescale( ) represents the rescaling operation.

[0032] The association difference is calculated in the same way as AssDis(P, S; X), that is, replacing the input X with the reconstructed time series data is sufficient.

[0033] Furthermore, the threshold in step (5) is automatically set by using the peaks-over-threshold model based on extreme value theory. This peaks-over-threshold model fits the extreme value probability distribution of the anomaly score sequence, derives the region beyond the threshold, and according to Pickands' theorem, its tail distribution is deduced to approximately converge to the generalized Pareto distribution function, thereby determining the optimal threshold.

[0034] The anomaly detection model of the present invention based on the VAE and association difference mechanism combines the variational autoencoder, uses the maximum evidence lower bound (ELBO) to estimate the probability distribution of the input data; performs random sampling in the latent space using the reparameterization trick to obtain the reconstructed time series data. The difference between the anomaly points and the normal points in the reconstructed time series data will be further enlarged. Learning the association difference of the reconstructed time series data can enhance the ability of the association difference to distinguish anomaly points, and the randomness in the VAE reconstruction process will greatly improve the robustness and performance of the model.

[0035] In the present invention, the calculation of the anomaly score combines the reconstruction error and the correlation difference to identify anomaly points. The reconstruction error includes the mean square error and the KL divergence, which can reflect the reconstruction ability of the model for time series. The correlation difference quantifies the difference between the prior correlation and the sequence correlation, helping to distinguish normal and abnormal data points. The sequence correlation is used to represent the correlation distribution between time points in the time series, providing rich context information for capturing the dynamic characteristics of the sequence (such as periodicity and trend). The prior correlation is based on the prior assumption of the relative positions of time points in the time series, reflecting the strong correlation that may exist between adjacent or nearby time points. Anomaly points usually have a small correlation difference, that is, the gap between their prior correlation and sequence correlation is small.

[0036] Therefore, the present invention has the following beneficial technical effects:

[0037] 1. The present invention effectively captures the key features of the time series through the self-attention mechanism, uses the dual indicators of correlation difference and reconstruction error to more accurately identify anomaly points, and improves the detection accuracy. Compared with existing methods, it performs better on multiple public data sets.

[0038] 2. The present invention avoids the gradient explosion problem of traditional prediction-based methods and the low efficiency of sequential processing, improves the detection speed, and enhances the detection efficiency.

[0039] 3. The present invention makes the model work stably in a complex dynamic environment, reduces misjudgments, and improves the robustness of the model by cleverly using the prior distribution of latent variables introduced by VAE. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 It is a schematic diagram of the working process of the multivariate time series data anomaly detection method based on VAE and correlation difference of the present invention.

[0041] Figure 2 It is a schematic diagram of the model structure based on VAE and correlation difference mechanism of the present invention.

[0042] Figure 3 It is a schematic diagram of the visualization result obtained by detecting the test set data through the model of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0043] In order to describe the present invention more specifically, the technical solutions of the present invention will be described in detail below with reference to the drawings and specific embodiments.

[0044] As shown in Figure 1, the multivariate time series data anomaly detection method based on VAE and correlation difference of the present invention includes the following steps:

[0045] (1)Data preprocessing: Preprocess the multivariate time series data, including denoising, normalization, and missing value filling, to obtain a clean data set suitable for model training. For the collected time series data, there may be noise introduced due to factors such as sensor errors and transmission interference. In this embodiment, the wavelet transform denoising method is adopted. This method is based on the wavelet analysis theory, decomposes the time series data into subsequences of different frequencies; by setting appropriate thresholds, suppresses or removes the noise components in the high-frequency subsequences, and retains the main signal features contained in the low-frequency part, thus achieving the purpose of denoising. In order to eliminate the dimensionality differences between different feature data and enable the model to learn and converge better, this embodiment adopts the min-max normalization method. For the data of each feature dimension x , calculate its minimum value min and maximum value max , and map the data to the [0, 1] interval through the formula 。

[0046] (2)Model construction: The network structure of the anomaly detection model VAE-Anomaly of the present invention is as shown in Figure 2 . For the input of multi-dimensional time series data, first perform element-wise addition of the input vector and the position encoding vector generated based on sine and cosine functions to encode the order information in the data. This approach ensures that the model can understand the relative positions of the elements in the sequence. The final input representation is X, which not only contains the time information of each time point but also encodes its position information in the sequence.

[0047] Assume that the association difference module contains L layers, and the input sequence is , where N is the length of the sequence and D is the feature dimension of each time point. The operation of the L-th layer can be expressed as:

[0048]

[0049]

[0050] where: , l ∈ {1, 2,..., L}, represents the output of the L-th layer, d model is the feature dimension of each layer, X 0 is the input X after position encoding, and is the intermediate state of the L-th layer.

[0051] The Anomaly-Attention module is used to calculate the association difference. In the Anomaly-Attention module of each layer of the association difference layer, four important matrices are defined: the query vector matrix Q, the key vector matrix K, the value vector matrix V, and the scale parameter matrix σ. The calculation methods of these matrices are as follows:​

[0052]

[0053]

[0054]

[0055]

[0056] Wherein: while , is the parameter matrix of the L-th layer and belongs to while .

[0057] For the prior association matrix P of the L-th layer l , its calculation formula is:

[0058]

[0059] Wherein: The prior association matrix is generated based on the learned scale parameter , and the i-th element in the matrix corresponds to the scale value at the i-th time point. By using the Rescale operation, the obtained association weights are divided by the sum of each row, thereby converting them into a normalized discrete distribution P l such that the sum of the elements in each row is 1.

[0060] Meanwhile, the sequence association matrix S l is calculated by taking the inner product of the query vector Q and the key vector K and normalizing it through the Softmax function. The formula is as follows:

[0061]

[0062] Wherein: represents the association matrix of the sequence, and the Softmax operation is normalized in the last dimension. Therefore, each row of S l constitutes a discrete distribution, representing the association weight between each time point and other time points.

[0063] By multiplying S l and the value vector matrix V, the output feature vector e l of this layer can be obtained:

[0064]

[0065] Wherein: It is the feature representation after the deep encoding of the input time series, which contains the context information and long-term dependencies of the time series.

[0066] The Association Discrepancy (AssDis) represents the symmetric KL divergence between the prior association and the sequence association, which measures the information gain between these two distributions. To calculate the overall association discrepancy, it is necessary to average the association discrepancies from multiple layers; specifically, the calculation formula for the association discrepancy is:

[0067]

[0068] where: The KL divergence measures the difference between two discrete distributions. For each layer l, P l and S l are the matrices of the prior association and the sequence association, and represent the distributions of the i-th row respectively.

[0069] Finally, is a vector consisting of N elements, and each element represents the association discrepancy at the corresponding time point in the time series.

[0070] After the association discrepancy module outputs the feature vector e l , the mean μ and variance σ of the approximate posterior distribution of the latent variable Z are fitted through a dimensionality reduction convolutional layer 2 to obtain the approximate posterior distribution of Z . This process can be regarded as finding a regular path in the high-dimensional data space, so that different feature vectors have a structured distribution representation in the latent space. Finally, the dimensionality increase convolutional layer is used to sample from using the reparameterization trick to obtain the final latent variable Z.

[0071] In the random association discrepancy layer, the model can effectively extract the association discrepancies between different time points in the time series by learning the time series data reconstructed by the reconstruction module. Since the reconstruction process can recover the latent features of the original data and introduce a certain degree of randomness to each data point, this makes the difference between abnormal data points and normal data points more obvious in the reconstructed data. By calculating these differences, the association discrepancy mechanism can effectively distinguish between abnormal and normal states within a small error range. In addition, the randomness in the reconstruction process not only enhances the generalization ability of the model to the data, but also through learning diverse samples, the model can show higher flexibility and robustness when dealing with different types of anomalies. This mechanism helps the model better capture the data distribution characteristics in the time series, especially when facing complex and dynamic environments, it can effectively identify potential abnormal patterns and improve the accuracy of anomaly detection.

[0072] (3) Abnormal score calculation.

[0073] The present invention uses the reconstruction error and the associated difference loss as the objectives for model training. The reconstruction error includes the mean squared error and the KL divergence. The final anomaly score is calculated as follows:

[0074]

[0075] Where: denotes element-wise multiplication, and MSE and KLD represent the mean squared error and the KL divergence at that point, respectively.

[0076] (4) Threshold setting.

[0077] After obtaining the anomaly scores at each time point, a threshold needs to be set to determine whether an anomaly occurs. Anomaly scores exceeding this threshold will be classified as anomalies. The present invention uses the peaks-over-threshold (POT) model based on extreme value theory (EVT) to automatically set the optimal threshold. This model fits the extreme value probability distribution of the anomaly score sequence, derives the region beyond the threshold, and according to Pickands' theorem, its tail distribution is approximately convergent to the generalized Pareto distribution function (GPD).

[0078] (5) Anomaly detection.

[0079] After obtaining the final threshold th F the present invention marks anomalies by judging whether the anomaly score is higher than this threshold. Specifically, the determination rules for anomaly points and normal points are as follows:

[0080]

[0081] To evaluate the anomaly detection model (VAE - Anomaly) proposed by the present invention, five representative benchmark datasets were selected from real - world application scenarios for experiments in this embodiment. And the performance of the model of the present invention was compared with the following several advanced benchmark models: LSTM - VAE (Long Short - Term Memory Network Variational Auto - Encoder), AnomalyTransformer (Anomaly Detection Transformer), DCdetector (Dual Attention Contrast Detector), OmniAnomaly (Multivariate Time - Series Omnipotent Anomaly Detector), InterFusion (Multivariate Time - Series Fusion Anomaly Detector), DAGMM (Deep Auto - Encoder Gaussian Mixture Model). The comparison experimental results are shown in Table 1. Generally speaking, the average F1 score of the model of the present invention on the five datasets has increased by 1.076% compared with the baseline method DCdetector. This result fully reflects the good generalization ability and effectiveness of the model of the present invention in complex environments with different anomaly data distributions and large variable differences.

[0082] Table 1

[0083]

[0084] In order to better demonstrate the effectiveness of the association difference and reconstruction error in time series modeling, the present invention selects some test data from the SMD dataset for visual analysis. Figure 3 It shows the fluctuations of different sequences in the test set. The three sequences are the association difference, the reconstructed sequence, and the true original sequence in turn. The abnormal points are marked in the red boxes. It can be clearly seen from the figure that the association difference values corresponding to the abnormal points in the original sequence are significantly lower than those of the normal points, while the reconstruction error shows a significant increase near the abnormal points. This indicates that the association difference mechanism can effectively capture the change in the association degree between the normal and abnormal states in the time series, and the association difference of the abnormal points has significant differential characteristics compared with that of the normal points. At the same time, the reconstruction error fluctuates greatly at the abnormal points, further proving that the reconstruction module can accurately reflect the abnormal characteristics in the data. Combining these two metrics, the model of the present invention can more effectively distinguish abnormal points from normal points, improving the accuracy of anomaly detection.

[0085] Figure 3 The stability of the reconstructed sequence is shown below. Although there are small fluctuations in the original data, the reconstructed sequence can still maintain a relatively stable trend. This characteristic shows that through the learning of the reconstruction module, the model of the present invention can capture the global trend in the time series and avoid misjudgment or overreaction caused by small fluctuations, effectively reducing the false alarm rate. In summary, the present invention combines the dual determination mechanisms of association difference and reconstruction error, enabling the model to show higher sensitivity and robustness when facing local fluctuations and sudden anomalies in time series data, and thus having better performance in complex anomaly detection tasks.

[0086] The above description of the embodiments is for the convenience of those of ordinary skill in the art in the technical field to understand and apply the present invention. It is obvious that those skilled in the art can easily make various modifications to the above embodiments and apply the general principles described herein to other embodiments without creative efforts. Therefore, the present invention is not limited to the above embodiments, and all improvements and modifications made by those skilled in the art based on the disclosure of the present invention should fall within the protection scope of the present invention.

Claims

1. A multivariate time series data anomaly detection method based on VAE and correlation difference, comprising the following steps: (1) obtaining multivariate time series data and preprocessing the data to obtain a clean data set suitable for model training, and then dividing the data set into a training set and a test set; the multivariate time series data is multivariate time series data on soil moisture from the SMAP data set; (2) Establish an anomaly detection model based on VAE and correlation difference mechanism, which includes: The correlation difference module is used to capture the time dependency in multivariate time series data and calculate the correlation difference of multivariate time series data as part of the anomaly detection metric; The reconstruction module is used to map the feature vector output by the correlation difference module to the latent space to capture its potential features, and then randomly sample in the latent space to obtain the reconstructed time series data; The random correlation difference module is used to capture the time dependency in the reconstructed time series data and calculate the correlation difference of the reconstructed time series data to further improve the anomaly detection and recognition ability of the model; The reconstruction module first uses a dimensionality reduction convolution layer to map the feature vector output by the association difference module to the latent space to fit the mean and variance of the approximate posterior distribution of the latent variable, thereby obtaining the approximate posterior distribution of the latent variable, i.e., the latent feature, and then uses a dimensionality increase convolution layer to randomly sample in the latent space using a reparameterization technique to obtain the final latent variable as the reconstructed time series data; (3) The above model is trained using the multivariate time series data of the training set, and the reconstruction error and the associated difference loss are used as the objectives of model training; (4) Input the multivariate time series data of the test set into the trained model to obtain the corresponding association difference. According to the association difference and reconstruction error, the anomaly score of each time point in the multivariate time series data is calculated by the following formula; Where: X represents multivariate time series data, Represents the reconstruction of time series data, is the anomaly score sequence of X, and each element value in the sequence corresponds to the anomaly score of each time point in the multivariate time series data; AssDis(P,S;X) is the association difference of the multivariate time series data X, To reconstruct time series data The reconstruction error consists of two parts, namely X and The mean square error and KL divergence KLD, ⊙ represents element-wise multiplication, Softmax() represents the Softmax function; (5) The abnormal score of each time point in the multivariate time series data is compared with the set threshold. If it exceeds the threshold, it is judged as abnormal, thereby identifying the abnormal time point in the multivariate time series data.

2. According to claim 1, a multivariate time series data anomaly detection method based on VAE and correlation difference is characterized in that: The preprocessing of multivariate time series data in step (1) includes denoising, normalization and missing value filling.

3. According to claim 1, a multivariate time series data anomaly detection method based on VAE and correlation difference is characterized in that: The multivariate time series data needs to be positionally encoded before being input into the model, that is, the multivariate time series data and the position encoding vector generated based on the sine and cosine functions are added element-wise to encode the sequential information in the data.

4. According to claim 1, a multivariate time series data anomaly detection method based on VAE and correlation difference is characterized in that: The association difference module is composed of a plurality of association difference layers in cascade. The association difference layer firstly processes the input x through Anomaly-Attention, then adds the output of Anomaly-Attention to the input x and inputs it into the feedforward neural network after layer normalization. Finally, the input and output of the feedforward neural network are added and processed by layer normalization as the output of the association difference layer.

5. According to claim 1, a multivariate time series data anomaly detection method based on VAE and correlation difference is characterized in that: The random association difference module is composed of a plurality of random association difference layers in cascade. The random association difference layer firstly processes the input x' through Anomaly-Attention, then adds the output of Anomaly-Attention to the input x' and inputs the result into the feedforward neural network after layer normalization. Finally, the input and output of the feedforward neural network are added and input into the fully connected layer after layer normalization. The output of the fully connected layer is used as the output of the random association difference layer.

6. The method for detecting anomalies of multivariate time series data based on VAE and correlation differences according to claim 1 is characterized in that: The expression of the association difference AssDis(P, S; X) is as follows: in: express Relative to The KL divergence of express Relative to The KL divergence of , L is the number of layers of the associated difference module, Represents the prior correlation matrix P l The distribution of the i-th row in , Represents the sequence correlation matrix S l The distribution of the i-th row in , where N is the sequence length of the multivariate time series data.

7. The method for detecting anomalies in multivariate time series data based on VAE and correlation differences according to claim 6 is characterized in that: The prior correlation matrix P l and the sequence correlation matrix S l The expression is as follows: Among them: Q, K, σ are query vector matrix, key vector matrix, and scale parameter matrix respectively. X l is the output of the lth layer in the associated difference module input to the multivariate time series data X, They correspond to the weight parameter matrices of Q, K, and σ at the lth layer, σ i is the value of the i-th element in the scale parameter matrix σ, d model For X l The characteristic dimension of T Represents transposition, and Rescale() represents a rescaling operation.

8. The method for detecting anomalies of multivariate time series data based on VAE and correlation differences according to claim 1 is characterized in that: The threshold in step (5) is automatically set using a super-threshold model based on extreme value theory. The super-threshold model derives the area exceeding the threshold by fitting the extreme value probability distribution of the abnormal score sequence, and derives its tail distribution according to the Pickands theorem to approximately converge to the generalized Pareto distribution function, thereby determining the optimal threshold.