A Multivariate Time Series Anomaly Detection Method Based on Transformer

Through the multivariate time series anomaly detection method based on the Transformer model, the accuracy and robustness of multivariate time series detection in the prior art are solved, and efficient anomaly detection of time series is achieved.

CN116796272BActive Publication Date: 2025-07-04FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310747003.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2025-07-04
Estimated Expiration
2043-06-21

AI Technical Summary

Technical Problem

When dealing with multivariate time series, existing time series anomaly detection methods have problems such as difficult to cluster assumptions, low accuracy of classification methods on label imbalanced datasets, slow training speed of recurrent neural networks and inability to effectively utilize sequence correlation, and limited information extraction when building graph structures.

Method used

The multivariate time series anomaly detection method based on the Transformer model is adopted. By normalizing the time series data, sliding window segmentation, timestamp and data encoding, combining feature fusion and adversarial training strategies, the encoder, discriminator and decoder of the Transformer model are used for reconstruction error analysis to determine the dynamic anomaly threshold and abnormal score.

Benefits of technology

Accurate and robust anomaly detection of multivariate time series is achieved, the accuracy and efficiency of detection is improved, and the characteristic correlation and long-term dependencies of time series can be effectively utilized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116796272B_ABST
    Figure CN116796272B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for anomaly detection of multivariate time series based on Transformer, comprising: for a multivariate time series data set, normalizing the time series data in the feature dimension; applying a sliding window to the normalized time series data to divide the original time series data into a series of sliding windows, obtaining multiple sliding window time series data; encoding the sliding window time series data through timestamp encoding and data encoding; adopting an anomaly detection method based on the Transformer model to solve the reconstructed values of the time series corresponding to all sliding windows; determining the dynamic anomaly thresholds corresponding to each time series according to the reconstruction error and historical error factors; solving the anomaly scores of each time series according to the corresponding dynamic anomaly thresholds to obtain the anomaly detection results of the time series data. Compared with the prior art, the present invention can accurately and robustly perform anomaly detection on multivariate time series.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of time series anomaly detection, and more particularly to a multi-variable time series anomaly detection method based on Transformer. Background Art

[0002] With the continuous improvement of the industrialization process, a large amount of data will be generated and stored during this process. Among various data types, time series data is one of the important data types. Time series data refers to the sequence data formed by collecting certain indicators at regular intervals over time and arranging the collected results in sequence. Time series data describes the changes of various dimensions of the system over time and is of great significance for capturing the correlation between front and back data and analyzing the development of the system.

[0003] Therefore, time series anomaly detection has always been an important topic in the industrial field. By analyzing continuous time series data, it is determined which instances are different from other instances, that is, the anomalies existing therein are discovered. Time series data has practical applications in many fields, such as financial markets, biological data, user behavior, industrial equipment, etc. Analyzing the anomalies of time series data is of great significance for ensuring the normal operation of the system and making timely predictions to avoid economic losses.

[0004] Traditional anomaly detection tasks were completed by data mining experts who reported errors by manually analyzing data that did not follow the normal trend. However, in a system for monitoring industrial equipment, with the continuous increase in the number of deployed sensors, the complexity of data patterns has been continuously increasing, and the challenges faced by manual fault identification are also becoming greater and greater. With the development of artificial intelligence, the emergence of big data analysis and deep learning can effectively help experts solve this problem. Currently, the methods adopted by time series anomaly detection technology mainly include two types: clustering or classification-based methods and deep learning-based methods. However, these methods still have corresponding problems:

[0005] First, it is difficult to determine the number of clusters K in the clustering-based anomaly detection method. The relatively classic method in the clustering-based approach is the K-Means clustering algorithm, which is also known as subsequence time series clustering. Given a certain time series data, according to the defined sliding window length, the time series data can be converted into a series of subsequences. After defining the number of clusters K, the K-Means clustering algorithm will act on the subsequences until the subsequences can converge to K categories. To perform time series anomaly detection, it is necessary to calculate the distance from each subsequence to the nearest category, usually using the Euclidean distance for calculation. Comparing this distance with a threshold, if it is greater than the threshold, the corresponding time series is abnormal. However, it is difficult to assume the value of K in advance for real system operations.

[0006] Second, the classification-based anomaly detection method regards the anomaly detection problem as a classification problem. First, the model is optimized and trained with the training data set, and then the data in the test set is detected using the model learned during the training process. According to the situation where anomalies are labeled in the training set, the classification-based anomaly detection method can be divided into multi-class anomaly detection and one-class anomaly detection. The original support vector machine algorithm is an one-class anomaly detection method, which is a linear supervised method. By introducing the kernel trick to extend the algorithm, this trick enables the SVM to perform non-linear classification. Subsequently, a new method for detecting anomalies using SVM (Support Vector Machine) was introduced, called One-Class Support Vector Machine (OC-SVM). OC-SVM is a semi-supervised method, where the training set only contains one category: normal data. After fitting the model in the training set, the test data is classified as whether it is similar to the normal data, so that anomalies can be detected. However, in real systems, the labels are usually extremely imbalanced, and the classification method has low accuracy on datasets with extremely imbalanced labels, and the training sets of most real datasets usually also contain abnormal data.

[0007] Third, the anomaly detection method based on recurrent neural network relies on a deep neural network model of recurrent neural network. It uses the input sequence as training data. For each input timestamp, it predicts the data of the next timestamp, and uses the prediction error as the basis for anomaly detection. For example, LSTM (Long Short-Term Memory) is an autoregressive neural network that learns the sequential dependencies in sequential data, where the prediction for each timestamp uses the feedback from the output of the previous timestamp. However, since such models are recurrent models, when the input time series is long, the training speed is slow, especially when the input data has noise, this performance is more obvious. In addition, such systems do not consider the correlation between different time series. Therefore, the effect of such methods in datasets with sequence correlation is not ideal.

[0008] Fourth, the anomaly detection method based on graph neural network is one of the most concerned systematic methods in recent years, but it depends on the construction of graph structure. Graph neural networks mainly focus on data in the form of graphs, mining the relationships between different nodes and edges in the graph. Such methods cannot be directly used for the anomaly detection problem under time series data, and an explicit graph structure needs to be constructed in advance. The anomaly detection system based on graph neural network effectively utilizes the correlation information between time series and improves the accuracy of anomaly detection. However, this method depends on explicitly constructing the graph data structure. When dealing with multivariate time series with fewer dimensions or less close relationships between sequences, the feature relationship graph constructed by graph neural network is too small or too sparse. This limits the information that the graph neural network model can extract from the data, resulting in a performance bottleneck of the system. Summary of the Invention

[0009] The purpose of the present invention is to overcome the above-mentioned defects existing in the prior art and provide a multivariate time series anomaly detection method based on Transformer, which can accurately and robustly perform anomaly detection on multivariate time series.

[0010] The purpose of the present invention can be achieved by the following technical solutions: A multivariate time series anomaly detection method based on Transformer, comprising the following steps:

[0011] S1. For the multivariate time series dataset, perform normalization processing on the time series data in the feature dimension;

[0012] S2. Apply a sliding window to the normalized time series data, divide the original time series data into a series of sliding windows, and obtain multiple sliding window time series data;

[0013] S3. Encode the sliding window time series data through timestamp encoding and data encoding;

[0014] S4. Adopt an anomaly detection method based on the Transformer model to solve the reconstructed values of the time series corresponding to all sliding windows;

[0015] S5. Determine the dynamic anomaly thresholds corresponding to each time series according to the reconstruction error and historical error factors;

[0016] S6. Solve the anomaly scores of each time series according to the corresponding dynamic anomaly thresholds to obtain the anomaly detection results of the time series data.

[0017] Further, the step S3 specifically includes the following steps:

[0018] S31. Decompose the time series information included in the time stamp into corresponding time stamp data and perform time stamp encoding to obtain a time stamp encoding vector;

[0019] S32. Adopt the method of Fourier transform to analyze the periodicity of the time series data and perform periodic encoding on the time stamp data to obtain a periodic encoding vector;

[0020] S33. Embed and project the periodic encoding vector and the time stamp encoding vector to obtain a global time series encoding;

[0021] S34. Perform local time series encoding on the time series data within the sliding window;

[0022] S35. Append the global time series encoding and the local time series encoding to the input time series data to obtain the input data of the Transformer model.

[0023] Further, the step S33 is specifically to perform embedding and projection operations on the periodic encoding vector and the time stamp encoding vector through a learnable embedding layer.

[0024] Further, in the step S4, the Transformer model includes an encoder, a feature fusion module, a first discriminator, a decoder, and a second discriminator. The encoder includes a feature attention module and a time series attention module. The encoder is used to convert the subsequences in the input data into corresponding latent variables;

[0025] The feature fusion module is used to combine multiple latent variables output by the encoder into a vector representation;

[0026] The first discriminator and the encoder form a first adversarial network. By applying a prior distribution to the latent variables, an adversarial training strategy is adopted to guide the prior distribution and the posterior distribution of the latent variables to be approximated;

[0027] The decoder is used to reconstruct the original input to obtain a reconstruction result;

[0028] A second discriminator and the decoder form a second adversarial network, which is used to apply adversarial training between the reconstruction result and the original input.

[0029] Further, the specific process of step S4 is as follows:

[0030] S41. Transmit the original input data to the encoder, apply a multi-dimensional attention mechanism, and respectively run a feature attention module and a temporal attention module to convert the original time series input X into a latent variable Y;

[0031] S42. Repeat the operation of step S41 for T non-overlapping subsequences to obtain corresponding T latent variables Y;

[0032] S43. Through linear interpolation, perform a feature fusion operation on T instances, that is, T latent variables Y, to form a new latent variable encoding Z;

[0033] S44. Apply a prior distribution to the latent variable, and adopt an adversarial training strategy to guide the prior distribution and the posterior distribution of the latent variable to be approximated;

[0034] S45. Input the latent variable into the decoder to reconstruct the original input of the model to obtain a reconstruction result X';

[0035] S46. Apply adversarial training between the reconstruction result X' and the original input X to obtain a trained anomaly detection model, transmit the current input data to the anomaly detection model, and output the reconstruction values of the time series corresponding to all sliding windows.

[0036] Further, in step S41, the feature attention module is used to learn the correlation between different features of the time series, and the temporal attention module is used to learn the long-term and short-term dependencies within the time series;

[0037] The specific process of step S41 is as follows:

[0038] First, transmit the input data to the feature attention module, calculate the feature encoding using the input data of the time series, and after obtaining the feature encoding, combine the original time series input with the feature encoding. At this time, the original input X of the time series is converted into a latent variable Y;

[0039] After that, input the latent variable Y into the temporal attention module to obtain temporal features. After obtaining the temporal features, combine the input Y of the temporal attention module to convert the time series into a latent variable Z:

[0040] Z = k1·O VA +k2·O TA+X + time_enc(X)

[0041] Among them, O VA is the feature encoding, O TA is the temporal feature, k1 is the weighting coefficient of the feature attention module, k2 is the weighting coefficient of the temporal attention module, and time_enc(X) is the time encoding result of the input X.

[0042] Furthermore, in the step S43, the T instances are represented as:

[0043]

[0044] The T instances are merged into a vector by linear interpolation:

[0045]

[0046] Among them, is the feature vector at time t.

[0047] Furthermore, the specific process of the step S44 is:

[0048] Assume that z satisfies the prior distribution p(Z), the original data X satisfies the distribution p(X), q(Z|X) is the encoded distribution, and the encoder generates the posterior distribution about the latent variable Z in the following way:

[0049]

[0050] Among them, is the original data at time t;

[0051] Assume that q(Z|XKt) satisfies the Gaussian distribution, and the reparameterization trick is used for backpropagation through the encoder network. The first discriminator D1 is used to make the posterior distribution q(Z) of the latent variable Z similar to its prior distribution p(Z). The goal of the first discriminator is to amplify the distance between them; on the other hand, the generator G1 composed of the encoder tries to narrow the gap between them in order to deceive the first discriminator D1. This process corresponds to the following optimization objective and uses the min-max strategy for optimization:

[0052]

[0053] Among them, V(D1, G1) is the cross-entropy loss between the encoder and the first discriminator.

[0054] Furthermore, the specific process of applying adversarial training between the reconstruction result X' and the original input X in the step S46 is:

[0055] The decoder G2 generates a reconstruction according to the generation of the encoder G1:

[0056] X′ t = G2(G1(X t K ))

[0057] The goal of the decoder G2 is to confuse the second discriminator D2, and the corresponding optimization function for this process is as follows:

[0058]

[0059] where V(D2, G2) is the cross-entropy loss between the decoder and the second discriminator.

[0060] Furthermore, the specific process of step S5 is as follows:

[0061] S51. Calculate the absolute value of the difference between the true time series data and the reconstructed time series data for all time series data:

[0062]

[0063] S52. Cache the historical reconstruction errors. At time t, after calculating the corresponding reconstruction error, add it to the historical error vector, as shown in the following formula:

[0064]

[0065] S53. Use the moving average model to smooth the reconstruction errors and perform multiple substitutions on the errors. Denote the error sequence after moving average processing as:

[0066]

[0067] S54. Based on the smoothed error sequence, determine the error threshold ε:

[0068]

[0069] where Δμ is the difference in error means, μ is the error mean, Δσ is the difference in error variances, and σ is the error variance. If the reconstruction error corresponding to a time point is greater than the error threshold ε, it indicates that an anomaly occurs at this time point; otherwise, this time point is normal.

[0070] Furthermore, the calculation formula for the anomaly score of the time series in step S6 is:

[0071]

[0072] where, s (i) is the anomaly score corresponding to the i-th subsequence, and e (i) is the error of the i-th subsequence.

[0073] Compared with the prior art, the present invention has the following advantages:

[0074] 1. For multivariate time series, the present invention first normalizes the time series data in the feature dimension and applies a sliding window to split the original time series data into a series of sliding windows; then encodes the time series data within the sliding windows through timestamp encoding and data encoding; adopts an anomaly detection method based on the Transformer model to solve the reconstructed values of the time series corresponding to all sliding windows; and further determines the dynamic anomaly thresholds corresponding to each time series according to the reconstruction error and historical error factors, and calculates the anomaly scores of each time series, thereby measuring the anomaly degree of the time series data. Thus, it can accurately detect anomalies in multivariate time series and ensure the robustness of the detection process.

[0075] 2. In the present invention, the designed Transformer model includes an encoder, a feature fusion module, a first discriminator, a decoder, and a second discriminator. The encoder is used to convert the subsequences in the input data into corresponding latent variables; the feature fusion module is used to combine multiple latent variables output by the encoder into a vector representation; a first adversarial network is formed between the first discriminator and the encoder, and by imposing a prior distribution on the latent variables, an adversarial training strategy is adopted to guide the prior distribution and posterior distribution of the latent variables to be approximately the same; the decoder is used to reconstruct the original input to obtain a reconstruction result; and a second adversarial network is formed between the second discriminator and the decoder to apply adversarial training between the reconstruction result and the original input. Thus, the Transformer model is well applied to time series anomaly detection, and the reconstructed values of the time series corresponding to all sliding windows can be accurately and reliably obtained.

[0076] 3. In the present invention, the encoder in the Transformer model includes a feature attention module and a temporal attention module. The feature attention module is used to learn the correlation between different features of the time series, and the temporal attention module is used to learn the long-term and short-term dependencies within the time series, which can fully and effectively obtain the correlation between different features in the time series and the long-term and short-term dependencies in the time series, thereby ensuring the accuracy of subsequent anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0077] Figure 1 is a schematic flowchart of the method of the present invention;

[0078] Figure 2 is a schematic diagram of time series encoding in the present invention;

[0079] Figure 3 is a schematic diagram of the Transformer model architecture built in the embodiment;

[0080] Figure 4 Schematic diagram of the architecture of the encoder in the Transformer model in the embodiment. Specific implementation mode

[0081] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0082] Embodiment

[0083] As Figure 1 shown, a multi-variable time series anomaly detection method based on Transformer includes the following steps:

[0084] S1. For the multi-variable time series data set, normalize the time series data in the feature dimension;

[0085] S2. Apply a sliding window to the normalized time series data, and split the original time series data into a series of sliding windows to obtain multiple sliding window time series data;

[0086] S3. Encode the sliding window time series data through timestamp encoding and data encoding;

[0087] S4. Adopt an anomaly detection method based on the Transformer model to solve the reconstructed values of the time series corresponding to all sliding windows;

[0088] S5. Determine the dynamic anomaly threshold corresponding to each time series according to the reconstruction error and historical error factors;

[0089] S6. Solve the anomaly score of each time series according to the corresponding dynamic anomaly threshold to obtain the anomaly detection result of the time series data.

[0090] The main contents of this embodiment applying the above technical solution are as follows:

[0091] First, for the multi-variable time series data set composed of multiple time points, normalize the time series data in the feature dimension;

[0092] Second, apply a sliding window to the normalized time series data, and split the original time series data into a series of sliding windows;

[0093] Third, encode the sliding window time series data through timestamp encoding and data encoding;

[0094] Fourth, adopt an anomaly detection method based on the Transformer model to solve the reconstructed values of the time series corresponding to all sliding windows;

[0095] Fifth, determine the dynamic anomaly threshold corresponding to each time series according to the reconstruction error and historical error factors;

[0096] 6. Further solve the anomaly scores of each time series according to the dynamic anomaly threshold, which is used to measure the anomaly degree of time series data.

[0097] In step 3, first determine the timestamps corresponding to the time series data processed by normalization and sliding window; then decompose the timestamp data into corresponding data information, and then encode the data information including year, month, day, hour, minute, etc. with timestamps, and combine periodic encoding for embedding and projection to form local time series encoding and global time series encoding; finally, append the local time series encoding and global time series encoding to the input time series data to form the input data of the Transformer model.

[0098] Specifically, as Figure 2 shown, decompose the time series information contained in the timestamp into corresponding timestamp data;

[0099] Perform timestamp position encoding on the timestamp data including information such as year, month, day, hour, minute, etc.;

[0100] Analyze the periodicity of the time series data using Fourier transform;

[0101] Perform periodic encoding on the timestamp data;

[0102] Through a learnable embedding layer, embed and project the periodic encoding vector and timestamp encoding vector to form global position encoding;

[0103] Perform local position encoding on the time series data within the sliding window;

[0104] Combine the global position encoding and local position encoding and append them to the input data.

[0105] In step 4, the constructed Transformer model architecture is as Figure 3 shown, and the specific process is as follows:

[0106] (4.1) Transmit the input data to the encoder based on the Transformer architecture, apply the multi-dimensional attention mechanism, and run the feature attention module and time series attention module respectively to convert the original time series input X into a latent variable Y, where the encoder architecture is as Figure 4 shown;

[0107] (4.2) Repeat the above operations for T non-overlapping subsequences to form T corresponding latent variables Y;

[0108] (4.3) Use the feature fusion module FFM to perform feature fusion operations on the T instances through linear interpolation to form a new latent variable encoding Z;

[0109] (4.4) Apply a prior distribution to the latent variables, and adopt an adversarial training strategy (a generator G1 and a first discriminator D1 composed of an encoder) to guide the prior distribution and the posterior distribution of the latent variables to be approximated;

[0110] (4.5) Input the latent variables into a decoder based on the Transformer architecture to reconstruct the original input X' of the model;

[0111] (4.6) Apply adversarial training between the reconstruction result and the original input. The second discriminator D2 guides the reconstruction result to be close to the original input data, amplifying the gap between the two. The goal of the generator G2 composed of the decoder is to deceive the second discriminator D2.

[0112] (4.7) Train two generators (i.e., the generator G1 composed of the encoder and the generator G2 composed of the decoder) and two discriminators (i.e., the first discriminator D1 and the second discriminator D2), and apply the trained Transformer model to the dataset to obtain the reconstruction result corresponding to the time series dataset.

[0113] Among them, the specific content of applying the multi-dimensional attention mechanism to the input data in step (4.1) is as follows:

[0114] In order to obtain the correlation between different features in the time series and the long-term and short-term dependencies in the time series, the multi-dimensional attention mechanism includes a feature attention module VA and a temporal attention module TA. First, run the feature attention module, which uses the input data X of the time series to calculate the feature O VA . After obtaining the feature O VA , combined with the original time series input X, the original input X of the time series is converted into a latent variable Y. Then, the temporal attention module TA uses the latent variable Y as the input to calculate the temporal feature O TA . After obtaining the feature O TA , combined with the input Y of the temporal attention module, the time series is converted into a latent variable z, and Z combines both O VA and O TA , which is a multi-dimensional feature, and the calculation process is shown in the following formula:

[0115] Z = k1·O VA + k2·O TA + X + time_enc(X)

[0116] In the encoder, the working process of the feature attention module VA includes:

[0117] In the feature attention module VA, for a certain time point t, the input of this module is the feature sequence X of all variables from T windows. First, it is projected into three feature spaces and transformed into Q(X), K(X), and V(X) as shown in the following formula:

[0118] Q(X) = X · W Q , K(X) = X · W K , V(X) = X · W V

[0119] W Q , W K , W V W, W, and W are the linear transformation matrices corresponding to Q, K, and V respectively. Then, the attention between variables is calculated using the features in the query space and the key space as shown in the following formula:

[0120] S = Q(X) · K(X) T

[0121]

[0122] S is the inner product result of Q and K. α (q,k) is the attention coefficient obtained after normalization, where S (q,k) is the inner product result of the q-th row and the k-th column, and S (q,j) is the inner product result of the q-th row and the j-th column. Finally, the normalized attention matrix is applied to the value space V(X), and a convolution operation is used to calculate the output O VA , as shown in the following formula:

[0123] O VA = α · V(X) · W o

[0124] α is the attention coefficient, and W o is the convolution coefficient.

[0125] The working process of the temporal attention module TA includes:

[0126] For a certain fixed variable, the input of this module consists of the output of the position encoding module and the feature attention module VA. Similar to the feature attention module, the attention matrix S can be calculated, where Si,j represents the attention of the historical memory at time step j to the feature at time step i. When i < j, Si,j is set to zero, that is, the upper right corner of the matrix S is set to 0.

[0127] Step (4.3) performs a feature fusion operation on T instances through linear interpolation. The process includes:

[0128] The feature fusion layer merges T instances into a vector through linear interpolation. The T instances are represented as

[0129]

[0130] Its combined representation is as follows:

[0131]

[0132] The process of using the adversarial training strategy in step (4.4) to guide the prior distribution and posterior distribution of the latent variable to be approximated is as follows:

[0133] The generated latent variables are merged into a vector Z through the feature fusion module FFM. Assume that Z satisfies the prior distribution p(Z), the original data X satisfies the distribution p(X), and q(Z|X) is the encoded distribution. The encoder generates the posterior distribution of the latent variable Z in the following way:

[0134]

[0135] The adversarial network D1 guides the posterior distribution q(Z) to match the prior distribution p(Z). Assume that q(Z|XKt) satisfies the Gaussian distribution, and the reparameterization trick is used for backpropagation through the encoder network. The first discriminator D1 is used to make the posterior distribution q(Z) of the latent variable Z similar to its prior distribution p(Z). The goal of the first discriminator is to amplify the distance between them; on the other hand, the generator G1 composed of the encoder tries to narrow the gap between them in order to deceive the first discriminator D1. This process corresponds to the following optimization objective, which can be optimized using the max-min strategy:

[0136]

[0137] The process of applying adversarial training between the reconstruction result and the original input in step (4.6) is as follows:

[0138] The decoder G2 generates a reconstruction based on the generation of the encoder G1:

[0139]

[0140] The generator G2 composed of the decoder aims to deceive the second discriminator D2. The corresponding optimization function for this process is as follows:

[0141]

[0142] In step five, the process of comprehensively determining the dynamic anomaly thresholds corresponding to each time series according to the reconstruction error and historical error factors is as follows:

[0143] (5.1) Calculate the absolute value of the difference between the real time series data and the reconstructed time series data of all time series data.

[0144]

[0145] (5.2) Cache the historical reconstruction error, determine the abnormality of the current data point, and comprehensively consider the historical error factors. At time t, after calculating the corresponding reconstruction error, add it to the historical error vector as shown in the following formula:

[0146]

[0147] (5.3) Use a smoothing average model to smooth the reconstruction error. Substitute the error multiple times, and denote the error sequence after moving average processing as shown in the following formula:

[0148]

[0149] (5.4) For the error sequence after smoothing, select an error threshold ε. Time points greater than this error threshold are regarded as abnormal, and vice versa. The method of selecting the threshold from the set is defined by the following formula:

[0150]

[0151] (5.5) According to the obtained error threshold, calculate the anomaly score for measuring the anomaly degree of each subsequence in each abnormal time series:

[0152]

[0153] To verify the effectiveness of this technical solution, this embodiment conducts anomaly detection on the existing public dataset SWaT. First, prepare the dataset. In this embodiment, the public dataset SWaT is a safety water treatment dataset collected by the iTrust laboratory of the Singapore University of Technology and Design. The experiment is carried out on a state-of-the-art water treatment test bed. This dataset records sensor values, such as water level, flow rate, etc.; and actuator values, including valves and pumps, etc. This dataset collects 7 days of normal data and 4 days of abnormal operation data, with an anomaly rate of approximately 11.98%. It contains 51 data dimensions, and the division ratio of the training set to the test set is approximately 1:1.

[0154] After that, perform data preprocessing. Normalize the time series included in the SWaT dataset. Specifically, perform normalization operations on 51 dimensions respectively. Select the maximum and minimum values of each dimension feature, subtract the minimum value of the corresponding numerical value and sequence at each moment, and calculate the difference between the maximum value and the minimum value. Divide the two differences as the new numerical value. Apply a sliding window to the normalized time series data. For the SWaT dataset, take the sliding window length as 10. After the time series dataset included in the SWaT dataset is normalized and divided by the sliding window, n - 10 sliding window sequences are formed, where n is the time series length of the training set or the test set.

[0155] Then, the Transformer model is trained. The main hyperparameters in the model are: training environment, training batch size, learning rate, optimizer, sliding window length, and the number of encoder-decoder modules. In this embodiment, the proposed model architecture is implemented using the PyTorch framework, and the hyperparameters of the model are set as follows:

[0156] Training environment - GPU is NVIDIA RTX 3060 graphics card, CPU model is Intel(R) Core(TM) i7-11700, and the memory is 32G;

[0157] The training batch size is set to 50;

[0158] The initial learning rate is set to 0.01;

[0159] The Adam optimizer is used;

[0160] The sliding window length is selected as 10;

[0161] The number of encoder-decoder modules is set to 1 respectively.

[0162] Finally, experimental tests are carried out. In the time series dataset, the number of normal data is usually much larger than that of abnormal data, and the labels of the dataset are extremely unbalanced. In this embodiment, three mainstream metrics are used to evaluate the performance of this method, including accuracy (Pre), recall (Rec), and F1-score (F1). Comparing the proposed method with 6 common baseline methods, the experimental results are shown in Table 1.

[0163] Table 1

[0164] Model Pre Rec F1 PCA 0.2667 0.2325 0.2484 DAGMM 0.2641 0.7182 0.3861 LSTM-NDT 0.7778 0.5109 0.6167 OmniAnomaly 0.9678 0.6869 0.8035 GDN 0.9697 0.6947 0.8094 MTAD-GAT 0.9689 0.6956 0.8098 RTAD (This invention) 0.9764 0.6997 0.8152

[0165] As can be seen from the results in Table 1, the F1-score obtained by this solution on the SWaT dataset is 0.8152, exceeding the compared baseline models and achieving the best result.

[0166] To further verify the effectiveness of each module in the robust Transformer method based on multi-dimensional attention, this embodiment also conducted a comparative experiment on the gain of each module on the final result. The specific settings are as follows: First, without considering the encoder-decoder structure of the transformer, a feed-forward neural network is used instead, and the corresponding model is named w / otransformer; Second, the global position encoding is removed and replaced with the absolute position encoding corresponding to the original Transformer network, and the corresponding model is named w / o ts_enc; Secondly, consider removing the modeling of variable associations using the feature attention module, that is, only considering the temporal attention association, and this model is named w / o var_emb; Finally, consider removing the use of the temporal attention module to model the long-term and short-term temporal dependencies, that is, only considering the associations between variables, and this model is named w / o tim_emb. Table 2 shows the impact of different modules on the model.

[0167] Table 2

[0168] Model Pre Rec F1 Without transformer 0.6696 0.5943 0.6297 Without ts_enc 0.9678 0.6874 0.8038 Without var_emb 0.9346 0.6569 0.7715 Without tim_emb 0.9296 0.6564 0.7694 RTAD 0.9764 0.6997 0.8152

[0169] From the results in Table 2, it can be seen that in the method proposed in this solution, each module has a gain on the final result, and there is no redundant design.

[0170] In summary, this technical solution innovatively applies the Transformer model to the time series anomaly detection technology, and fully considers the correlation between different features in the time series and the long-term and short-term dependencies in the time series, effectively improving the accuracy and robustness of multi-variable time series anomaly detection.

Claims

1. A multi-variable time series anomaly detection method based on Transformer, characterized in that It includes the following steps: S1. For the multi-variable time series dataset, normalize the time series data in the feature dimension; S2. Apply a sliding window to the normalized time series data, split the original time series data into a series of sliding windows, and obtain multiple sliding window time series data; S3. Encode the sliding window time series data through timestamp encoding and data encoding; S4. Adopt an anomaly detection method based on the Transformer model to solve the reconstructed values of the time series corresponding to all sliding windows; S5. Determine the dynamic anomaly thresholds corresponding to each time series according to the reconstruction error and historical error factors; S6. Solve the anomaly scores of each time series according to the corresponding dynamic anomaly thresholds to obtain the anomaly detection results of the time series data; The specific process of step S4 is as follows: S41. Transmit the original input data to the encoder, apply the multi-dimensional attention mechanism, and run the feature attention module and the time series attention module respectively to convert the original time series input X into a latent variable Y; S42. Repeat the operation of step S41 for T non-overlapping subsequences to obtain the corresponding T latent variables Y; S43. Through linear interpolation, perform feature fusion operations on the T instances, that is, the T latent variables Y, to form a new latent variable encoding Z; S44. Apply a prior distribution to the latent variable, and adopt an adversarial training strategy to guide the prior distribution and the posterior distribution of the latent variable to be approximate; S45. Input the latent variable into the decoder to reconstruct the original input of the model and obtain the reconstruction result X'; S46. Apply adversarial training between the reconstruction result X' and the original input X to obtain a trained anomaly detection model, transmit the current input data to the anomaly detection model, and output the reconstructed values of the time series corresponding to all sliding windows; In step S41, the feature attention module is used to learn the correlation between different features of the time series, and the time series attention module is used to learn the long-term and short-term dependencies within the time series; The specific process of step S41 is as follows: First, transmit the input data to the feature attention module, calculate the feature encoding using the input data of the time series, and after obtaining the feature encoding, combine the original time series input with the feature encoding. At this time, the original input X of the time series is converted into a latent variable Y; After that, input the latent variable Y into the time series attention module to obtain the time series features. After obtaining the time series features, combine the input Y of the time series attention module to convert the time series into a latent variable Z: Z = k1O VA + k2·O TA + X + time_enc(X) Among them, O VA is the feature encoding, O TA is the temporal feature, k1 is the weighting coefficient of the feature attention module, k2 is the weighting coefficient of the temporal attention module, and time_enc(X) is the time encoding result of the input X.

2. The multivariate time series anomaly detection method based on Transformer according to claim 1, characterized in that, Step S3 specifically includes the following steps: S31. Decompose the time series information contained in the timestamp into corresponding timestamp data, and perform timestamp encoding to obtain a timestamp encoding vector; S32. Adopt the Fourier transform method to analyze the periodicity of the time series data, and perform periodic encoding on the timestamp data to obtain a periodic encoding vector; S33. Embed and project the periodic encoding vector and the timestamp encoding vector to obtain a global time series encoding; S34. Perform local time series encoding on the time series data within the sliding window; S35. Append the global time series encoding and the local time series encoding to the input time series data to obtain the input data of the Transformer model.

3. A multivariate time series anomaly detection method based on Transformer according to claim 1, characterized in that, In step S4, the Transformer model includes an encoder, a feature fusion module, a first discriminator, a decoder, and a second discriminator. The encoder includes a feature attention module and a time series attention module. The encoder is used to convert the subsequences in the input data into corresponding latent variables. The feature fusion module is used to combine multiple latent variables output by the encoder into a vector representation. A first adversarial network is formed between the first discriminator and the encoder. By imposing a prior distribution on the latent variables, an adversarial training strategy is adopted to guide the prior distribution and the posterior distribution of the latent variables to be approximated. The decoder is used to reconstruct the original input to obtain a reconstruction result. A second adversarial network is formed between the second discriminator and the decoder, which is used to impose adversarial training between the reconstruction result and the original input.

4. A multivariate time series anomaly detection method based on Transformer according to claim 1, characterized in that In step S43, T instances are represented as: Merge the T instances into a vector by linear interpolation: Among them, is the feature vector at time sequence t.

5. A method for multivariate time series anomaly detection based on Transformer according to claim 4, characterized in that, The specific process of step S44 is as follows: Assume that Z follows the prior distribution p(Z), the original data X follows the distribution p(X), and q(Z|X) is the encoded distribution. The encoder generates the posterior distribution of the latent variable Z in the following way: Among them, is the original data at time t; Assume that q(Z|XKt) follows a Gaussian distribution. The reparameterization trick is used to perform backpropagation through the encoder network. The first discriminator D1 is used to guide the posterior distribution q(Z) of the latent variable Z to be similar to its prior distribution p(Z). The goal of the first discriminator is to amplify the distance between them. On the other hand, the generator G1 composed of the encoder tries to narrow the gap between them in order to deceive the first discriminator D1. This process corresponds to the following optimization objective, and the max-min strategy is used for optimization: Among them, V(D1, G1) is the cross-entropy loss between the encoder and the first discriminator.

6. A method for multi-variable time series anomaly detection based on Transformer according to claim 5, characterized in that, The specific process of imposing adversarial training between the reconstruction result X' and the original input X in step S46 is as follows: The decoder G2 generates a reconstruction according to the generation of the encoder G1: The goal of the decoder G2 is to deceive the second discriminator D2. The corresponding optimization function for this process is as follows: Among them, V(D2, G2) is the cross-entropy loss between the decoder and the second discriminator.

7. A method for anomaly detection of multivariate time series based on Transformer according to claim 6, characterized in that, The specific process of step S5 is as follows: S51. Calculate the absolute value of the difference between the true time series data and the reconstructed time series data of all time series data: S52. Cache the historical reconstruction error. At time t, after calculating the corresponding reconstruction error, add it to the historical error vector, as shown in the following formula: S53. Use a smoothed average model to smooth the reconstruction error, substitute the error multiple times, and denote the error sequence after moving average processing as: S54. Based on the error sequence after smoothing, determine the error threshold ε: Among them, Δμ is the difference in the mean error, μ is the mean error, Δσ is the difference in the error variance, σ is the error variance. If the reconstruction error corresponding to a time point is greater than the error threshold ε, it indicates that an anomaly occurs at this time point; otherwise, this time point is normal.

8. A multivariate time series anomaly detection method based on Transformer according to claim 7, characterized in that The calculation formula for the anomaly score of the time series in step S6 is as follows: where s (i) is the anomaly score corresponding to the i-th subsequence, and e (i) is the error of the i-th subsequence.

Citation Information

Patent Citations

  • Multi-dimensional time series data anomaly detection method and system

    CN114065862A

  • Data anomaly detection method and apparatus, and electronic device and storage medium

    WO2021189904A1