Time sequence anomaly detection method based on adversarial auto-encoder and memory mechanism
By combining adversarial autoencoders with memory mechanisms, utilizing missing value discrimination loss and multiple time series transformations, the problem of poor detection of incomplete time series in existing technologies is solved, and more accurate and stable anomaly detection is achieved.
Patent Information
- Application Number
- CN202510770476.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-19
AI Technical Summary
Existing time series anomaly detection methods ignore the missing label information in incomplete time series, resulting in poor detection results. The interpolation method has problems fitting abnormal samples, resulting in low recognition accuracy.
A method based on adversarial autoencoder and memory mechanism is adopted. Through window replacement, window scaling and smooth transformation of data, combined with missing value discriminator and autoencoder with memory mechanism, an adversarial game is constructed, and the identification loss of missing part is used for anomaly detection.
It significantly improves the anomaly detection accuracy of incomplete data, can better reconstruct incomplete time series data, improves the anomaly recognition ability and stability, and avoids the overfitting of abnormal samples by traditional methods.
Smart Images

Figure CN120671046A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of electrical digital signal processing, and in particular relates to a time series anomaly detection method based on an adversarial autoencoder and a memory mechanism. Background Art
[0002] With the development of information technology, the real-time generation and transmission of large-scale data sets in streaming form has become the norm. Time series data analysis is widely used in many important fields, including medicine, aerospace, ocean exploration, and equipment monitoring. Due to hardware failures and software storage issues in time series data collection, time series often contain missing values. Time series anomaly detection has become increasingly widely used in many important research areas within time series data analysis. Time series anomaly detection methods can be divided into three categories. The first category is prediction models, which predict the sequence based on the observed values in the time series and detect anomalies by comparing the difference between the predicted and true values. The second category is feature models, which represent the data using latent features and use traditional clustering or classification algorithms to identify anomalies in the data's latent feature representation. The third category is reconstruction models, which autoencode the data and use the reconstruction error to detect anomalies.
[0003] Existing anomaly detection methods generally fail to consider the issue of missing data. They only learn latent feature representations from observable data, ignoring the data feature information contained in missing labels. Using only the latent feature representations of observed values makes it difficult to achieve good prediction and reconstruction results for incomplete time series, ultimately leading to suboptimal anomaly detection. Interpolation methods for incomplete time series share similarities with time series anomaly detection methods: both require learning latent feature representations of the data for reconstruction or prediction. Compared to anomaly detection methods, time series interpolation methods introduce missing labels and combine the latent features of observed values with those of missing values to provide a more comprehensive representation of the data. However, the difference between the reconstruction of abnormal and normal samples using interpolation methods is not significant, making the reconstruction of incomplete time series using interpolation methods difficult to use for anomaly detection. Effectively mining the data feature information contained in the missing parts to improve the accuracy of identifying anomalies in incomplete time series remains a challenge.
[0004] Time series anomaly detection has a wide range of applications, including industrial monitoring and medical assistance. Initially, anomaly detection was performed manually by monitors. To reduce the time cost of anomaly detection, researchers have increasingly proposed automated time series anomaly detection methods. Initial time series data anomaly detection methods primarily rely on statistical models and time-frequency analysis. These methods mathematically model data features, predict subsequent data using statistical models, and detect anomalies by comparing the difference between true and predicted values. These methods are characterized by their speed and the lack of numerous hyperparameter adjustments, making them suitable for streaming data processing and real-time anomaly detection. Researchers have proposed the statistically based SPOT method for identifying anomalies. For example, Luo Yonghong's paper, "Research on Missing Value Filling Algorithms for Time Series Data Based on Generative Adversarial Networks," estimates the probability of each data point being an anomaly by assuming that the data values conform to a Gaussian distribution. Fourier transform-based SR methods, such as Huang C et al.'s paper, "Time series anomaly detection for trustworthy services in cloud computing systems," published in EEE Transactions on Big Data, use the Fourier transform to filter time series data, ultimately highlighting anomalies. The aforementioned statistical model-based methods all have strong assumptions and rely on the representation of the dataset itself, resulting in poor generalization. With the development of neural networks, deep learning-based time series anomaly detection methods can achieve better performance than traditional statistical model methods in most scenarios. Deep learning-based time series anomaly detection methods use deep neural networks to learn representations of time series samples and use the learned latent feature representations for anomaly detection. The paper "Unsupervised representation learning by predicting random distances" by Wang H et al. proposes a random distance prediction model (RDP). This model uses the mapping distances obtained from a random mapping function as labels, employs a neural network to fit the sample distribution in a new space, and uses the new mapping vectors obtained by the neural network to cluster the original samples, identifying data in sparse clusters as anomalies. The RDP method fits the distribution representation of the random mapping space, but the underlying principle is not well explained.Pang et al.'s paper "Learning representations of ultrahigh-dimensional data for random distance-based outlier detection" discloses an unsupervised model, REPEN, which jointly trains representation learning and outlier detection. This model learns representations of the data and uses the learned latent feature representations to cluster and identify anomalies. Both RDP and REPEN share a common approach: they use representation learning to learn data feature representations without sample labels, effectively solving the problem of excessive raw sample data making labeling difficult. However, methods that cluster feature representations rely on the spatial distribution of the feature representations and the coordination of the clustering method, making them less robust.
[0005] In addition to the aforementioned deep learning anomaly detection methods based on representational clustering, a more common approach is to use the reconstruction model commonly used in the aforementioned time series data interpolation methods. The reconstruction model maps the data into a latent space and then maps it back to the original space to obtain the reconstructed data. The anomaly of each data point is determined by comparing the reconstruction error between the reconstructed data and the original sample data. For example, the variational autoencoder (VAE) proposed by Kingma et al. learns the probability distribution of the data to reconstruct the data. However, the disadvantage of the VAE model is that it is prone to reconstruction ambiguity in complex data scenarios. Zhang C et al.'s paper "A deep neural network for unsupervised anomaly detection and diagnosis in multivariate time series data" discloses the MSCRED method. This method first maps the time series into a two-dimensional vector space and uses a two-dimensional convolutional neural network to construct an autoencoder to reconstruct the data. This method cleverly applies image autoencoders to time series data, but is limited by the filtering effect of the convolution kernel, making it difficult to detect more detailed anomalies.
[0006] In summary, the existing technology has the following deficiencies: (1) Existing time series anomaly detection methods generally assume the integrity of the data, ignore the missing label information in incomplete time series data, and only learn the latent feature representation through observations. The latent feature representation of observations is difficult to accurately predict and reconstruct incomplete time series data, which leads to the poor performance of existing time series anomaly detection methods in handling incomplete data.
[0007] (2) Reconstructing incomplete data through interpolation methods can mine the potential feature information of missing data labels, but the problem of good fitting of abnormal samples by interpolation methods leads to low accuracy in identifying abnormalities. Summary of the Invention
[0008] The purpose of the present invention is to address the above problems and provide a time series anomaly detection method based on an adversarial autoencoder and memory mechanism, introduce the identification loss of missing parts, better reconstruct incomplete time series data, combine the time series data feature transformation for anomaly detection, and comprehensively analyze the three time series transformations of window replacement, window scaling, and smoothing to detect anomalies based on the reconstruction errors of the three transformed data and the original data, which is conducive to improving the stability of the method for anomaly detection of incomplete time series.
[0009] In order to achieve the above object, the technical solution provided by the present invention is: The time series anomaly detection method based on the adversarial autoencoder and memory mechanism includes the following steps: Step 1: After preprocessing the incomplete time series data, it is used as the first type of time series data. The time series data is transformed using window replacement, window scaling, and smoothing methods to obtain the second, third, and fourth types of time series data; Step 2: Input the first, second, third, and fourth categories of time series data obtained in step 1 into the first, second, third, and fourth autoencoders respectively to obtain the completed samples of the four categories of time series data; The first, second, third and fourth autoencoders are all autoencoders based on memory mechanisms; Step 3: Construct a missing value discriminator, input the completed samples of the four types of time series data obtained in step 2 into the missing value discriminator, and calculate the identification loss result; Step 4: constructing the loss functions of the missing value discriminator and the memory-based autoencoder respectively, and simultaneously training the missing value discriminator and the memory-based autoencoder using the time series data sample set. During the training process, the missing value discriminator and the autoencoder form an adversarial game; Step 5: Calculate the anomaly score of the incomplete time series data of the time series data sample set and determine the anomaly threshold; Step 6: After processing the incomplete time series data to be detected using the methods of steps 1 and 2, calculate the anomaly score of the time series data, and compare the magnitude relationship between the anomaly score and the anomaly threshold to evaluate the degree of anomaly of the incomplete time series data.
[0010] Furthermore, in step 1, the time series data is transformed using window replacement, window scaling, and smoothing methods, specifically including: (1) Window replacement operation: Divide the time series data samples into windows and replace some of the windows to change the time series characteristics of the original data samples. Divide each time series data sample into four windows, numbered [1, 2, 3, 4] in sequence, and then replace the data window with [3, 1, 4, 2] to swap the time series of the data samples. The specific calculation formula for window replacement is: ; Where X 1 i 、X 2 i 、X 3 i 、X 4 i represents the data of the four windows of the i-th time series data sample; X t i Represents the result after the i-th time series data sample window is replaced; (2) Window scaling operation: The time series data samples are divided into windows, and the data in different windows are scaled with different weights to change the shape characteristics of the original time series data samples. Each time series data sample is divided into four windows, and the window scaling weights are [0.5, 2, 0.8, 1.2]. The specific calculation formula for window scaling is: X e i = [0.5, 2, 0.8, 1.2 ]⊙[X 1 i , X 2 i ,X 3 i ,X 4 i ]; Where, X e i Represents the result after scaling the i-th time series data sample window; (3) Smoothing method using sliding window: Smooth the time series data samples and filter the data by sliding window averaging. Set the sliding window width, average the data in the window, and use the average value to replace the original data value.
[0011] Preferably, in step 2, the autoencoder based on the memory mechanism includes an encoder, a memory module and a decoder connected in sequence, the encoder is a BiRNN model, the decoder is a BiLSTM model, and the memory module uses the attention weight W to analyze the sample features output by the encoder. Z i Transform to obtain sample features with attention The decoder pair Decode and obtain the reconstruction result of the sample; ; In the formula Indicates the i time series data samples, Representation sample The missing label of G ( ) represents the mapping function of the autoencoder; For samples The reconstruction result.
[0012] Preferably, in step 3, the missing value discriminator includes a BiLSTM and a fully connected layer, and the input of the missing value discriminator is the completed sample and the prompt matrix R , R = K ⊙ M + 0.5× (1 - K); Where K is a random binary matrix and M is the missing label; The completed sample is a complete sample obtained by filling the reconstruction result to the missing position. ; In the formula express The completion sample of For samples The reconstruction result of Representation sample The missing label.
[0013] Furthermore, in step 3, the specific process of calculating the identification loss result includes: (1) Obtain the hidden layer feature vector through BiLSTM encoding, ; In the formula represents the hidden layer vector of BiLSTM at time t, Represents the mapping function of the hidden layer of BiLSTM; represents the hidden layer vector obtained by the forward LSTM at time t-1, represents the hidden layer vector obtained by reverse LSTM at time t+1, Indicates the completion sample The tth data point; (2) Input the hidden layer feature vector obtained by BiLSTM into the fully connected layer to obtain the identification loss result. ; In the formula 、 are the weight and bias parameters of the fully connected layer respectively, is the activation function, To complete the sample Missing identification results.
[0014] Preferably, in step 4, the objective function of the adversarial game between the missing value discriminator and the autoencoder based on the memory mechanism is: ; ; Where G represents the mapping function of the missing value discriminator; D represents the mapping function of the autoencoder based on the memory mechanism; X represents the time series data sample set; Represents the reconstruction result of X; M represents the missing label set corresponding to X; R is the prompt matrix.
[0015] Preferably, in step 4, the autoencoder loss function based on the memory mechanism consists of two parts, namely the reconstruction loss L pre and the discrimination loss L che , reconstruction loss L pre represents the reconstruction error of the non-missing part, ; ; Where, represents the i-th time series data sample, Representation sample The reconstruction result of Representation sample The missing label, Representation sample The missing identification vector, E is the expected function; The reconstruction loss L pre and the discrimination loss L che Add together to get the total loss function L of the autoencoder based on the memory mechanism G , L G = L pre + αL che ; Where α is the discrimination loss L che The weight hyperparameters.
[0016] Preferably, in step 4, the loss function L of the missing value discriminator is D The calculation formula is, ; Where D represents the identification loss L che The mapping function of represents the completed sample set corresponding to X, M represents the missing label set corresponding to X; R is the prompt matrix.
[0017] Furthermore, in step 5, the calculation formula of the abnormal threshold is: th = µ+ β* σ ; In the formula th represents the abnormal threshold, µ 、 σ are the average values of the anomaly scores of the time series data sample set, β is the scaling factor; The calculation formulas for µ and σ are: ; ; In the formula Indicates the i The anomaly score of a time series data sample; N represents the number of samples in the time series data sample set; ; In the formula represents the jth data point of the i-th time series data sample, express The corresponding reconstruction results; express The corresponding missing labels.
[0018] Preferably, the calculation formula for the anomaly score of the incomplete time series data to be detected is: ; Where λ represents the weight coefficient of the anomaly score of the first type of time series samples, 、 、 、 Represent the abnormal scores of the first to fourth types of time series samples respectively; in 、 、 、 It is calculated using the following formula: ; In the formula represents the anomaly score of the time series data to be detected, represents the jth data point of the incomplete time series to be detected, express The corresponding reconstruction results are, express The corresponding missing labels.
[0019] Compared with the prior art, the present invention has the following beneficial effects: 1) Through the triple innovation of multi-time series transformation integration, memory-enhanced autoencoder, and adversarial learning missing value discriminator, this paper achieves for the first time the efficient utilization of missing label information and the significant amplification of anomaly reconstruction errors, significantly improving the anomaly detection accuracy of incomplete data. The adversarial learning of the missing value discriminator and autoencoder can ensure that the reconstruction of the missing positions of incomplete time series conforms to the true distribution, avoiding the "false completion" of anomaly samples by traditional methods, such as the overfitting of anomaly samples by interpolation methods.
[0020] 2) The anomaly detection method of this invention introduces a discriminative loss for missing parts of incomplete time series. It learns the latent feature representation of the time series from the observations and missing labels, enabling better reconstruction of incomplete time series data. It also introduces a memory mechanism to obtain the latent representation of the data through the normal mode register matrix, suppressing the method's fitting of anomalous samples.
[0021] 3) The present invention uses three time series transformation methods, namely window replacement, window scaling, and smoothing, to map the original time series data, and reconstructs the transformed data and the original time series data through multiple autoencoders. The window replacement operation destroys the original time series structure and exposes time series dependency anomalies; the window scaling operation changes the data scale distribution and strengthens shape anomaly detection; the smoothing operation suppresses noise interference and highlights macroscopic anomaly patterns; finally, the reconstruction errors of the three transformed data and the original time series data are combined to perform anomaly detection on the original time series data, which is beneficial to improving the stability and robustness of anomaly detection for incomplete time series.
[0022] 4) This invention adopts a dual-loss joint optimization mechanism of the autoencoder and the missing value discriminator. The autoencoder based on the memory mechanism reconstructs the loss L pre Constrain the reconstruction accuracy of the observations by the discrimination loss L che Force the missing position to reconstruct the result to "fool" the discriminator; the missing value discriminator passes the loss function L D Distinguish between real / reconstructed missing values; the autoencoder and missing value discriminator improve the completion quality by playing against the target dynamic game. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The present invention will be further described below with reference to the accompanying drawings.
[0024] Figure 1 Schematic diagram of the process of detecting anomalies in time series according to an embodiment of the present invention.
[0025] Figure 2 Schematic diagram of the data flow of the autoencoder based on the memory mechanism according to an embodiment of the present invention.
[0026] Figure 3Schematic diagram of a window replacement, window scaling, and smoothing method for time series data according to an embodiment of the present invention.
[0027] Figure 4 Schematic diagram of the structure of the autoencoder based on the memory mechanism according to an embodiment of the present invention.
[0028] Figure 5 Schematic diagram of the structure of the missing value discriminator according to an embodiment of the present invention.
[0029] Figure 6 This is a code diagram of the training process of the autoencoder and missing value discriminator based on the memory mechanism in an embodiment of the present invention. DETAILED DESCRIPTION
[0030] like Figure 1 and Figure 2 As shown in Figure 2, the time series anomaly detection method based on the adversarial autoencoder and memory mechanism includes: Step 1: After preprocessing the incomplete time series data as the first type of time series data, the time series data are transformed using window replacement, window scaling, and smoothing methods, such as Figure 3 As shown, the second, third and fourth types of time series data are obtained.
[0031] Data normalization: ; where X o i Represents the original data of the i-th time series data sample, Represents X o i The result after normalization of the j-th data point.
[0032] Window replacement operation: Divide the time series data samples into windows and replace some of the windows to change the time series characteristics of the original data samples. Divide each time series data sample into four windows, numbered [1, 2, 3, 4] in sequence, and then replace the data window with [3, 1, 4, 2] to swap the time series of the data samples. The specific calculation formula for window replacement is: ; Where X 1 i 、X 2 i 、X 3 i 、X 4 i Represents the data of the four windows of the i-th time series data sample; X t iRepresents the result after window replacement of the i-th time series data sample.
[0033] Window scaling operation: The time series data samples are divided into windows, and the data in different windows are scaled with different weights to change the shape characteristics of the original time series data samples. Each time series data sample is divided into four windows, and the window scaling weights are [0.5, 2, 0.8, 1.2]. The specific calculation formula for window scaling is: X e i = [0.5, 2, 0.8, 1.2 ]⊙[X 1 i , X 2 i ,X 3 i ,X 4 i ]; Where, X e i Represents the result after scaling the i-th time series data sample window.
[0034] Smoothing method using sliding window: Smooth the time series data samples and filter the data by sliding window averaging. Set the sliding window width to 3 and take the average value of the data in the window to replace the original data value. Smooth the detailed features in the disturbed data and reconstruct the smoothed data samples to facilitate abnormal identification of the overall features of the data samples. The specific calculation formula for smoothing is: X s i,j = (X i,j-1 + X i,j + X i,j+1 ) / 3; where X i,j-1 、X i,j 、X i,j+1 represent the j-1th, jth, and j+1th data points of the i-th time series data sample respectively; X s i,j represents the jth data point after smooth conversion of the i-th time series data sample; Finally, the smoothed result X is obtained s i = {X s i,1 , X s i,2 , ..., X s i,n}.
[0035] Step 2: Input the first, second, third, and fourth categories of time series data obtained in step 1 into the first, second, third, and fourth autoencoders respectively to obtain the completed samples of the four categories of time series data.
[0036] In the embodiment, the first, second, third and fourth autoencoders are all autoencoders based on a memory mechanism, including an encoder, a memory module and a decoder connected in sequence, such as Figure 4 As shown in the figure, the encoder is a BiRNN model, and the decoder is a BiLSTM model. The memory module learns and stores the feature register vectors of the normal pattern. Abnormal samples input to the autoencoder deviate from the normal pattern, resulting in a significant increase in reconstruction error, while normal samples can be accurately reconstructed.
[0037] The memory module uses the attention weight W to analyze the sample features output by the encoder. Z i Transform to obtain the implicit representation of time series data, which is the latent variable of time series features. ; decoder pair Decode and obtain the reconstruction result of the sample.
[0038] The memory module includes the attention weight layer, the addressing hard contraction layer and the storage matrix layer. The attention weight layer converts the latent variables obtained by the encoder into Z i As the query vector, we get the initial weight w of the query register matrix layer i , the addressing hard shrinkage layer is the initial weight w i Perform secondary shrinkage to obtain weights ,pass Add the weights of the register matrix layer to obtain the time series feature latent variables . The register matrix layer is used to store the typical normal patterns of all samples in the time series sample set. The normal feature information in the feature register matrix is used as an addressable register item, and the normal features in the sample are restored by adding the weights of each vector in the register matrix layer. Let the feature register matrix be Y = {y1, y2, y3, . . . , y L}, where y1, y2, y3, . . . y L They represent the 1st, 2nd, 3rd,..., Lth registered vectors respectively, where L is the number of registered vectors.
[0039] By calculating the time series feature latent variables Z i The similarity with each registered vector in the registered matrix layer is used as the attention weight, and w is obtained by the Softmax method. i , ; in d( ) represents the similarity measurement function; exp ( ) represents an exponential function.
[0040] The addressing results obtained by the attention weight layer can reconstruct normal feature data well. However, due to the sparsity of the feature storage matrix, some abnormal samples can still be well reconstructed from some subtle feature representations in the feature storage vector. To alleviate this problem, the addressing hard shrinkage layer is introduced. The continuous ReLU activation function is used to perform a secondary shrinkage on the addressing results of the attention weight layer, eliminating the smaller weight parts of the results obtained from the query vector. ; in is the addressing weight vector of the feature register matrix, max ( ) is the function of ReLU activation unit, λ is the contraction threshold, and its value range is [1 / L , 3 / L]. is a positive scalar that prevents the denominator from being zero.
[0041] Get the weight vector Finally, the eigenvectors in the feature register matrix layer are combined to obtain the implicit features of the time series data. , ; Through the register mechanism of the memory module, the encoder processes the time series data samples. Encoded latent variables Z i It becomes the query vector (Query) of the feature storage matrix in the feature memory module. Each eigenvector of the feature storage matrix Y acts as a key value (Key). The feature memory module fits the entire sample data as a tensor. The size of the feature storage matrix in the memory module has a certain limit, and the number of normal samples in the sample data is much greater than the number of abnormal samples. Therefore, the memory module stores a much higher proportion of normal patterns than abnormal patterns. It is difficult for the query vector (Query) to fit the abnormal samples well. This process avoids the model's good reconstruction of abnormal samples. ; In the formula Indicates the i time series data samples, Representation sample The missing label of G ( ) represents the mapping function of the autoencoder; For samples The reconstruction result.
[0042] Step 3: Construct a missing value discriminator, input the completed samples of the four types of time series data obtained in step 2 into the missing value discriminator, and calculate the identification loss result.
[0043] The missing value discriminator includes BiLSTM and fully connected layers, such as Figure 5 shown.
[0044] The task of the missing value discriminator is to identify the completed samples reconstructed by the autoencoder Among them, which are the data points generated by the autoencoder and which are the original observable data points.
[0045] The purpose of the hint matrix R is to ensure that the missing value discriminator's identification target is unique. If the hint matrix R is not present in the input, or if the hint matrix R and the probability mass distribution of the missing labels M are independent, the missing value discriminator will have multiple identification targets. The maximum-minimum game between the autoencoder and the missing value discriminator will have multiple possible targets, ultimately making it difficult for the autoencoder to accurately mine the hidden features of the missing labels. Therefore, the hint matrix R and the missing labels M cannot be independently distributed, and the hint matrix R needs to contain sufficient dependency information on the missing labels M to ensure the uniqueness of the final game target of the adversarial model composed of the sequence autoencoder and the missing value discriminator.
[0046] The input of the missing value discriminator is the completed sample and the prompt matrix R , R = K ⊙ M + 0.5× (1 - K); Where K is a random binary matrix and M is the missing label.
[0047] The completed sample is a complete sample obtained by filling the reconstruction result to the missing position. ; In the formula express The completion sample of For samples The reconstruction result of Representation sample The missing label.
[0048] The specific process of calculating the identification loss result includes: (1) Obtain the hidden layer feature vector through BiLSTM encoding, ; In the formula represents the hidden layer vector of BiLSTM at time t, Represents the mapping function of the hidden layer of BiLSTM; represents the hidden layer vector obtained by the forward LSTM at time t-1, represents the hidden layer vector obtained by reverse LSTM at time t+1, Indicates the completion sample The tth data point; (2) Input the hidden layer feature vector obtained by BiLSTM into the fully connected layer to obtain the identification loss result. ; In the formula 、 are the weight and bias parameters of the fully connected layer respectively, is the activation function, To complete the sample Missing identification results.
[0049] Step 4: Construct the loss functions of the missing value discriminator and the memory-based autoencoder respectively, and use the time series data sample set to train the missing value discriminator and the memory-based autoencoder simultaneously. During the training process, the missing value discriminator and the autoencoder form an adversarial game.
[0050] The memory-based autoencoder reconstructs the data, while the missing value discriminator identifies missing values in the autoencoder's interpolated data. First, the memory-based autoencoder uses the missing value discriminator's discriminant loss to constrain the missing portion of the reconstruction target. Then, the missing value discriminator uses the autoencoder's interpolated results as samples and introduces missing labels for training. The memory-based autoencoder aims to make the missing value discriminator unable to distinguish the original missing data after interpolation, while the missing value discriminator aims to accurately distinguish the interpolated values reconstructed by the autoencoder, creating a competition between the two.
[0051] The objective function of the adversarial game between the missing value discriminator and the memory-based autoencoder is: ; ; Where G represents the mapping function of the missing value discriminator; D represents the mapping function of the autoencoder based on the memory mechanism; X represents the time series data sample set; Represents the reconstruction result of X; M represents the missing label set corresponding to X; R is the prompt matrix.
[0052] The autoencoder loss function based on the memory mechanism consists of two parts: reconstruction loss L pre and the discrimination loss L che , reconstruction loss L pre represents the reconstruction error of the non-missing part, ; ; Where, represents the i-th time series data sample, Representation sample The reconstruction result of Representation sample The missing label, Representation sample The missing discriminant vector of , E is the expected function.
[0053] The reconstruction loss L pre and the discrimination loss L che Add together to get the total loss function L of the autoencoder based on the memory mechanism G , L G = L pre + αL che ; Where α is the discrimination loss L che The weight hyperparameters.
[0054] The present invention takes the reconstruction loss of the data reconstruction part and the reconstruction identification loss of the missing part as the overall loss function of the autoencoder, constrains the autoencoder from two aspects: the observation value reconstruction loss and the missing data identification loss, so that the model's feature representation of incomplete time series is more stable and comprehensive.
[0055] The loss function L for the missing value discriminator D The calculation formula is, ; Where D represents the identification loss L che The mapping function of represents the completed sample set corresponding to X, M represents the missing label set corresponding to X; R is the prompt matrix.
[0056] In the embodiment, the program code of the missing value discriminator and the training method of the self-encoder based on the memory mechanism is as follows Figure 6 shown.
[0057] Step 5: Calculate the anomaly score of the incomplete time series data in the time series data sample set and determine the anomaly threshold.
[0058] In step 5, the calculation formula for the abnormal threshold is: th = µ+ β* σ ; In the formula th represents the abnormal threshold, µ 、 σ are the average values of the anomaly scores of the time series data sample set, β is the scaling factor.
[0059] The calculation formulas for µ and σ are: ; ; In the formula Indicates the i The anomaly score of a time series data sample; N represents the number of samples in the time series data sample set.
[0060] ; In the formula represents the jth data point of the i-th time series data sample, express The corresponding reconstruction results; express The corresponding missing labels.
[0061] Step 6: After processing the incomplete time series data to be detected using the methods of steps 1 and 2, calculate the anomaly score of the time series data and compare the anomaly score with the anomaly threshold. If the anomaly score of the incomplete time series data is greater than the anomaly threshold, the time series is determined to be abnormal.
[0062] The calculation formula for the anomaly score of the incomplete time series data to be detected is: ; Where λ represents the weight coefficient of the anomaly score of the first type of time series samples, 、 、 、 Represent the abnormal scores of the first to fourth categories of time series samples respectively.
[0063] in 、 、 、 It is calculated using the following formula: ; In the formula represents the anomaly score of the time series data to be detected, represents the jth data point of the incomplete time series to be detected, express The corresponding reconstruction results are, express The corresponding missing labels.
[0064] In an embodiment, incomplete time series data samples are recollected at regular intervals, the time series sample set is dynamically updated, and the first to fourth autoencoders and the missing value detector are retrained using the time series sample set. The anomaly threshold is dynamically adjusted based on the reconstruction results of the time series sample set to adapt to the dynamic changes in the industrial application field and improve the dynamic adaptability of the method of the present invention.
[0065] Corresponding to the above method, the present invention provides an incomplete time series anomaly detection system, comprising the following modules: Data transformation module: used to transform the time series data using window replacement, window scaling, and smoothing methods after preprocessing the time series data to obtain transformed data of the original time series data; The first autoencoder module: uses an autoencoder based on a memory mechanism to reconstruct the original time series data to obtain the completed samples of the original time series data; The second autoencoder module uses an autoencoder based on a memory mechanism to reconstruct the time series data after window replacement to obtain the completed samples of the time series data after window replacement; The third autoencoder module uses an autoencoder based on a memory mechanism to reconstruct the time series data after window scaling to obtain the completed samples of the time series data after window scaling; The fourth autoencoder module uses an autoencoder based on a memory mechanism to reconstruct the time series data after smoothing and filtering to obtain the completed samples of the time series data after smoothing and filtering; Missing value identification module: uses the missing value detector to calculate the identification loss results of the completed samples output by the first to fourth autoencoder modules; Training module: Use the time series sample set to train the autoencoders and missing value detectors of the first to fourth autoencoder modules; Anomaly threshold module: used to determine the anomaly threshold based on the statistical indicators of the anomaly score of samples in the time series sample set; Anomaly score module: used to calculate the anomaly score of samples in a time series sample set or the time series data to be detected; Anomaly determination module: calls the anomaly score module and the anomaly threshold module, compares the anomaly score of the time series to be detected with the anomaly threshold output by the anomaly threshold module, and evaluates the degree of anomaly of the time series data to be detected.
[0066] Implementation results demonstrate that the method provided by this invention significantly improves the accuracy of anomaly detection for incomplete data and enhances the ability to identify complex anomalies. Window replacement, scaling, and smoothing operations enhance the recognition of time series misalignments, amplitude mutations, and trend anomalies, overcoming the sensitivity of single models or methods to local noise. By storing normal pattern prototypes through a memory mechanism, the strong reliance of clustering methods (such as RDP and REPEN) on feature space distribution is avoided. Adversarial training reduces strong assumptions about data distribution and can adapt to complex time series patterns, such as non-stationary data in medical and industrial scenarios. This method exhibits high robustness and generalization, demonstrating groundbreaking detection accuracy and stability advantages in scenarios such as industrial monitoring and medical diagnosis that require processing incomplete time series data. The adversarial training framework and memory mechanism design provide a new paradigm for time series anomaly detection.
Claims
1. A time series anomaly detection method based on adversarial autoencoders and memory mechanisms, characterized by: The following steps are involved: Step 1: After preprocessing the incomplete time series data, it is used as the first type of time series data. The time series data is transformed using window replacement, window scaling, and smoothing methods to obtain the second, third, and fourth types of time series data; Step 2: Input the first, second, third, and fourth categories of time series data obtained in step 1 into the first, second, third, and fourth autoencoders respectively to obtain the completed samples of the four categories of time series data; The first, second, third and fourth autoencoders are all autoencoders based on memory mechanisms; Step 3: Construct a missing value discriminator, input the completed samples of the four types of time series data obtained in step 2 into the missing value discriminator, and calculate the identification loss result; Step 4: constructing the loss functions of the missing value discriminator and the memory-based autoencoder respectively, and simultaneously training the missing value discriminator and the memory-based autoencoder using the time series data sample set. During the training process, the missing value discriminator and the autoencoder form an adversarial game; Step 5: Calculate the anomaly score of the incomplete time series data of the time series data sample set and determine the anomaly threshold; Step 6: After processing the incomplete time series data to be detected using the methods of steps 1 and 2, calculate the anomaly score of the time series data, and compare the magnitude relationship between the anomaly score and the anomaly threshold to evaluate the degree of anomaly of the incomplete time series data.
2. The time series anomaly detection method according to claim 1, characterized in that: In step 1, the time series data is transformed using window replacement, window scaling, and smoothing methods, specifically including: (1) Window replacement operation: Divide the time series data samples into windows and replace some of the windows to change the time series characteristics of the original data samples. Divide each time series data sample into four windows, numbered [1, 2, 3, 4] in sequence, and then replace the data window to [3, 1, 4, 2] to swap the time series of the data samples. The specific calculation formula for window replacement is: ; Where X 1 i 、X 2 i 、X 3 i 、X 4 i represents the data of the four windows of the i-th time series data sample; X t i Represents the result after the i-th time series data sample window is replaced; (2) Window scaling operation: The time series data samples are divided into windows, and the data in different windows are scaled with different weights to change the shape characteristics of the original time series data samples. Each time series data sample is divided into four windows, and the window scaling weights are [0.5, 2, 0.8, 1.2]. The specific calculation formula for window scaling is: X e i = [0.5, 2, 0.8, 1.2 ]⊙[X 1 i , X 2 i ,X 3 i ,X 4 i ]; Where, X e i Represents the result after scaling the i-th time series data sample window; (3) Smoothing method using sliding window: Smooth the time series data samples and filter the data by sliding window averaging. Set the sliding window width, average the data in the window, and use the average value to replace the original data value.
3. The time series anomaly detection method according to claim 2, characterized in that: In step 2, the autoencoder based on the memory mechanism includes an encoder, a memory module and a decoder connected in sequence. The encoder is a BiRNN model, the decoder is a BiLSTM model, and the memory module uses the attention weight W to analyze the sample features output by the encoder. Z i Transform to obtain sample features with attention The decoder pair Decode and obtain the reconstruction result of the sample; ; In the formula Indicates the i time series data samples, Representation sample The missing label of G ( ) represents the mapping function of the autoencoder; For samples The reconstruction result.
4. The time series anomaly detection method according to claim 3, characterized in that: In step 3, the missing value discriminator includes BiLSTM and a fully connected layer. The input of the missing value discriminator is the completed sample and the prompt matrix. R , R = K ⊙ M + 0.5× (1 - K); Where K is a random binary matrix and M is the missing label; The completed sample is a complete sample obtained by filling the reconstruction result to the missing position. ; In the formula express The completion sample of For samples The reconstruction result of Representation sample The missing label.
5. The time series anomaly detection method according to claim 4, characterized in that: In step 3, the specific process of calculating the identification loss result includes: (1) Obtain the hidden layer feature vector through BiLSTM encoding, ; In the formula represents the hidden layer vector of BiLSTM at time t, Represents the mapping function of the hidden layer of BiLSTM; represents the hidden layer vector obtained by the forward LSTM at time t-1, represents the hidden layer vector obtained by reverse LSTM at time t+1, Indicates the completion sample The tth data point; (2) Input the hidden layer feature vector obtained by BiLSTM into the fully connected layer to obtain the identification loss result. ; In the formula 、 are the weight and bias parameters of the fully connected layer respectively, is the activation function, To complete the sample Missing identification results.
6. The time series anomaly detection method according to claim 5, characterized in that: In step 4, the objective function of the adversarial game between the missing value discriminator and the memory-based autoencoder is: ; ; Where G represents the mapping function of the missing value discriminator; D represents the mapping function of the autoencoder based on the memory mechanism; X represents the time series data sample set; Represents the reconstruction result of X; M represents the missing label set corresponding to X; R is the prompt matrix.
7. The time series anomaly detection method according to claim 6, characterized in that: In step 4, the autoencoder loss function based on the memory mechanism consists of two parts: reconstruction loss L pre and the discrimination loss L che , reconstruction loss L pre represents the reconstruction error of the non-missing part, ; ; Where, represents the i-th time series data sample, Representation sample The reconstruction result of Representation sample The missing label, Representation sample The missing identification vector, E is the expected function; The reconstruction loss L pre and the discrimination loss L che Add together to get the total loss function L of the autoencoder based on the memory mechanism G , L G = L pre + αL che ; Where α is the discrimination loss L che The weight hyperparameters.
8. The time series anomaly detection method according to claim 7, characterized in that: In step 4, the loss function L of the missing value discriminator is D The calculation formula is, ; Where D represents the identification loss L che The mapping function of represents the completed sample set corresponding to X, M represents the missing label set corresponding to X; R is the prompt matrix, and E is the expected function.
9. The time series anomaly detection method according to claim 8, characterized in that: In step 5, the calculation formula for the abnormal threshold is: th = µ+ β* σ ; In the formula th represents the abnormal threshold, µ 、 σ are the average values of the anomaly scores of the time series data sample set, β is the scaling factor; The calculation formulas for µ and σ are: ; ; In the formula Indicates the i The anomaly score of a time series data sample; N represents the number of samples in the time series data sample set; ; In the formula represents the jth data point of the i-th time series data sample, express The corresponding reconstruction results; express The corresponding missing labels.
10. The time series anomaly detection method according to claim 9, characterized in that: The calculation formula for the anomaly score of the incomplete time series data to be detected is: ; Where λ represents the weight coefficient of the anomaly score of the first type of time series samples, 、 、 、 Represent the abnormal scores of the first to fourth types of time series samples respectively; in 、 、 、 It is calculated using the following formula: ; In the formula represents the anomaly score of the time series data to be detected, represents the jth data point of the incomplete time series to be detected, express The corresponding reconstruction results are, express The corresponding missing labels.