A method, apparatus, device and medium for abnormal detection of time-series data

Through the combination of dual-view data embedding and signal decomposition Transformer network, combined with multi-task learning and multi-anomaly measurement, the problem of insufficient accuracy and generalization ability of timing data in the prior art is solved, and more efficient abnormal detection effect is achieved.

CN115758273BActive Publication Date: 2025-05-30PURPLE MOUNTAIN LAB

Patent Information

Application Number
CN202211362344.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-02
Publication Date
2025-05-30
Estimated Expiration
2042-11-02

AI Technical Summary

Technical Problem

The prior art has insufficient accuracy and training inference speed in time-series data anomaly detection, and the generalization ability of the model is poor, making it difficult to detect multiple anomaly types at the same time.

Method used

The two-view data embedding method is adopted to embed time and signal directions, and the time-view embedding features and signal viewing angle embedding features are fused to generate target data embedding features. Then, the input signal is decomposed into the Transformer network, and anomaly determination is made by multi-task learning and multi-anomaly measurement methods.

Benefits of technology

It improves the accuracy of timing data anomaly detection and training and inference speed, enhances the generalization ability of the model, and can more effectively detect data anomalies of multiple signal-related anomalies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115758273B_ABST
    Figure CN115758273B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, equipment and medium for abnormal detection of time-series data, relating to the field of artificial intelligence technology. The method includes: acquiring target time-series data and performing preprocessing on the target time-series data to obtain the preprocessed target time-series data; inputting the preprocessed target time-series data into a trained embedding network for dual-view data embedding in the time direction and the signal direction to obtain a time-view embedding feature and a signal-view embedding feature, and fusing the time-view embedding feature and the signal-view embedding feature to obtain a target data embedding feature; inputting the target data embedding feature into a trained signal decomposition Transformer network to obtain a target feature representation corresponding to the target time-series data; determining a target prediction error, a target reconstruction error and a target distribution distance according to the target feature representation, and performing abnormal determination on the target time-series data according to the target prediction error, the target reconstruction error and the target distribution distance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a method, device, equipment and medium for detecting anomalies in time series data. Background Art

[0002] Anomaly detection in time series data (i.e., Anomaly detection) is one of the most mature applications of time series data analysis at present, and is a process of identifying abnormal events or behaviors from normal time series. Effective anomaly detection is widely used in many fields of the real world, such as quantitative trading, network security detection, autonomous driving vehicles, and daily maintenance of large industrial equipment. Generally speaking, many anomalies can be judged manually. However, when the business combination is complex and the time series scale becomes larger, it becomes difficult to rely on traditional manual and simple absolute value algorithms such as year-on-year and month-on-month comparisons. Therefore, in the face of various industrial scenarios, an automated time series anomaly detection method based on machine learning is particularly important. Based on traditional machine learning methods, statistical models, multivariate normal distribution models, isolation forests and other methods can detect obvious anomaly points to a certain extent, but are sensitive to data noise and only independently model each time series data, making it difficult to solve anomalies caused by the mutual correlation in multi-dimensional time series data. Moreover, the time series data in actual scenarios has characteristics such as large noise, large fluctuations, and being greatly affected by the environment, and traditional machine learning methods are difficult to meet the requirements of complex scenarios.

[0003] In recent years, deep learning-based methods have been gradually applied to the anomaly detection of time series data. The common method is to encode sequences based on recurrent neural networks and then combine methods based on prediction, reconstruction, or data distribution distance to locate anomaly points. The method based on recurrent neural networks and sequence reconstruction uses a type of LSTM in the recurrent neural network for feature encoding and decoding. At the same time, a variational autoencoder (i.e., VAE) is used as the architecture for feature encoding and decoding. This method is a recurrent neural network-based method and uses the reconstruction error as a measure of anomaly. The main reasoning process is as follows: First, perform data preprocessing, input the image into the trained encoder network (i.e., Encoder) for feature extraction and encoding, then input the encoded features into the decoder (i.e., Decoder) for decoding to reconstruct the signal at a certain moment, and finally determine whether it is an anomaly point by means of threshold judgment. However, on the one hand, when using recurrent neural networks for encoding and decoding, since recurrent neural networks can only perform sequential calculations and cannot be parallelized, the training and inference speeds are slow. Slow training leads to an increase in training time and computing power costs, while slow inference speed affects the timeliness of online detection. On the other hand, due to the uncertainty of the occurrence of anomalies and the large number of anomaly types, such as single-point anomalies, context-related anomalies, periodic change anomalies, trend change anomalies, etc. Different anomaly types have different requirements for features and anomaly definitions, and this method does not model the correlation relationship between signals and only uses a single reconstruction error. Therefore, it is difficult to use all anomaly types simultaneously, and the model's adaptability and generalization ability are poor. The method based on convolutional neural networks (i.e., CNN) and attention networks (i.e., Attention Net) has the advantage of parallel computing and can also be used to solve the problem of time series data anomaly detection through specific optimizations and improvements. Although these methods can solve the problem of time series data anomaly detection to a certain extent, there is still room for improvement. In summary, the problem of how to improve the accuracy, training and inference speeds, and the generalization ability of the model when performing time series data anomaly detection remains to be further solved. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a time series data anomaly detection method, device, equipment, and medium, which can improve the accuracy, training and inference speeds, and the generalization ability of the model when performing time series data anomaly detection. The specific scheme is as follows:

[0005] In the first aspect, the present application discloses a time series data anomaly detection method, including:

[0006] Obtain target time series data and perform preprocessing on the target time series data to obtain preprocessed target time series data;

[0007] Input the preprocessed target time-series data into the trained embedding network for dual-perspective data embedding in the time direction and the signal direction to obtain time-perspective embedding features and signal-perspective embedding features, and then fuse the time-perspective embedding features and the signal-perspective embedding features to obtain target data embedding features;

[0008] Input the target data embedding features into the trained signal decomposition Transformer network to obtain the target feature representation corresponding to the target time-series data;

[0009] Determine the target prediction error, the target reconstruction error, and the target distribution distance according to the target feature representation, and perform anomaly determination on the target time-series data according to the target prediction error, the target reconstruction error, and the target distribution distance.

[0010] Optionally, the obtaining of the target time-series data and the preprocessing of the target time-series data to obtain the preprocessed target time-series data includes:

[0011] Obtain the target time-series data and normalize the target time-series data to obtain the normalized target time-series data;

[0012] Perform a one-dimensional convolution operation on the normalized target time-series data in the time direction to obtain the preprocessed target time-series data.

[0013] Optionally, the inputting of the preprocessed target time-series data into the trained embedding network for dual-perspective data embedding in the time direction and the signal direction to obtain time-perspective embedding features and signal-perspective embedding features includes:

[0014] Input the preprocessed target time-series data into the trained embedding network, and perform dilated causal convolution operations on each signal in the time window of the target time-series data to obtain time-perspective embedding features;

[0015] Perform dilated causal convolution operations on the values of all signals at each time point in the time window of the target time-series data to obtain signal-perspective embedding features.

[0016] Optionally, the inputting of the target data embedding features into the trained signal decomposition Transformer network to obtain the target feature representation corresponding to the target time-series data includes:

[0017] Embed the target data embedding feature into the trained signal decomposition Transformer network to obtain the target periodic component feature, the target trend component feature, and the target residual component feature, and add the target periodic component feature, the target trend component feature, and the target residual component feature to obtain the target feature representation corresponding to the target time series data.

[0018] Optionally, the embedding of the target data embedding feature into the trained signal decomposition Transformer network to obtain the target periodic component feature, the target trend component feature, and the target residual component feature includes:

[0019] Embed the target data embedding feature into the trained signal decomposition Transformer network and obtain the periodic component feature through the frequency domain attention module in a preset number of decomposition layers in the signal decomposition Transformer network;

[0020] Remove the periodic component feature from the target data embedding feature to obtain the first residual component feature, and input the first residual component feature into the multi-head attention module in the decomposition layer to obtain the trend component feature;

[0021] Remove the trend component feature from the first residual component feature to obtain the second residual component feature, and input the second residual component feature into the Feed-forward module in the decomposition layer to obtain the residual component feature;

[0022] Use the residual component feature as the input of the next decomposition layer until the target residual component feature is output after being processed by all the decomposition layers, and perform channel fusion on the trend component feature and the residual component feature obtained after being processed by all the decomposition layers respectively to obtain the target periodic component feature and the target trend component feature.

[0023] Optionally, the determining of the target prediction error, the target reconstruction error, and the target distribution distance according to the target feature representation includes:

[0024] Input the target feature representation into the trained prediction network, reconstruction network, and distribution space distance network respectively to obtain the target prediction error, the target reconstruction error, and the target distribution distance.

[0025] Optionally, the abnormal determination of the target time series data according to the target prediction error, the target reconstruction error, and the target distribution distance includes:

[0026] According to the target prediction error, the target reconstruction error, and the target distribution distance, a voting method or a weighted summation method is used to obtain a target error, and the target error is compared with a preset error threshold;

[0027] If the target error is greater than the preset error threshold, it is determined that the target time-series data is abnormal data.

[0028] In a second aspect, the present application discloses a time-series data anomaly detection device, including:

[0029] A preprocessing module, configured to obtain target time-series data and preprocess the target time-series data to obtain preprocessed target time-series data;

[0030] A data embedding module, configured to input the preprocessed target time-series data into a trained embedding network for dual-perspective data embedding in the time direction and the signal direction to obtain a time-perspective embedding feature and a signal-perspective embedding feature, and then fuse the time-perspective embedding feature and the signal-perspective embedding feature to obtain a target data embedding feature;

[0031] A feature determination module, configured to input the target data embedding feature into a trained signal decomposition Transformer network to obtain a target feature representation corresponding to the target time-series data;

[0032] An anomaly determination module, configured to determine a target prediction error, a target reconstruction error, and a target distribution distance according to the target feature representation, and perform anomaly determination on the target time-series data according to the target prediction error, the target reconstruction error, and the target distribution distance.

[0033] In a third aspect, the present application discloses an electronic device, including:

[0034] A memory, configured to store a computer program;

[0035] A processor, configured to execute the computer program to implement the steps of the time-series data anomaly detection method disclosed above.

[0036] In a fourth aspect, the present application discloses a computer-readable storage medium, configured to store a computer program; wherein, when the computer program is executed by a processor, the steps of the time-series data anomaly detection method disclosed above are implemented.

[0037] When performing anomaly detection on time-series data, this application acquires target time-series data and preprocesses the target time-series data to obtain preprocessed target time-series data. The preprocessed target time-series data is input into a trained embedding network for dual-perspective data embedding in the time direction and the signal direction to obtain time-perspective embedding features and signal-perspective embedding features. Then, the time-perspective embedding features and the signal-perspective embedding features are fused to obtain target data embedding features. The target data embedding features are input into a trained signal decomposition Transformer network to obtain a target feature representation corresponding to the target time-series data. Finally, a target prediction error, a target reconstruction error, and a target distribution distance are determined based on the target feature representation, and an anomaly determination is made for the target time-series data according to the target prediction error, the target reconstruction error, and the target distribution distance. It can be seen that when performing anomaly detection on time-series data, this application first collects target time-series data and preprocesses the target time-series data, and performs dual-perspective data embedding on the preprocessed target time-series data in the time direction and the signal direction to obtain time-perspective embedding features and signal-perspective embedding features respectively. After that, the time-perspective embedding features and the signal-perspective embedding features are fused to obtain target data embedding features. Further, a target feature representation is determined through a signal decomposition Transformer network. A target prediction error, a target reconstruction error, and a target distribution distance are determined based on the target feature representation, and further anomaly detection of the target time-series data is performed. Thus, it can be seen that when performing anomaly detection on time-series data, after acquiring target time-series data and preprocessing the target time-series data, a dual-perspective data embedding is performed on the preprocessed target time-series data using an embedding network to obtain time-perspective embedding features and signal-perspective embedding features. Then, the time-perspective embedding features and the signal-perspective embedding features are fused to obtain target data embedding features. Through the dual-perspective data embedding, the correlation features between signals are better learned, and thus data anomaly situations such as multi-signal correlation anomalies can be better detected. On the other hand, when performing anomaly detection on time-series data, a slow recurrent neural network is avoided, so there is an obvious advantage in terms of training and inference speed. The training time is short and the inference speed is fast, which can save computational resource costs and improve the timeliness of feature anomaly detection. Moreover, the target prediction error, the target reconstruction error, and the target distribution distance are jointly used to make an anomaly determination for the target time-series data. Based on the mechanism of multi-task learning and the fusion of multi-anomaly measurement methods, the anomaly detection ability of the model is further improved, thus greatly improving the detection adaptability and generalization ability for different anomalies. In summary, this application can improve the accuracy, training and inference speed, and the generalization ability of the model when performing anomaly detection on time-series data. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0039] Figure 1 Flowchart of a method for detecting abnormal time series data provided by this application;

[0040] Figure 2 Flowchart of a specific method for detecting abnormal time series data provided by this application;

[0041] Figure 3 Network structure diagram of the dilated causal convolution module provided by this application;

[0042] Figure 4 Network structure diagrams of two frequency domain attention modules provided by this application;

[0043] Figure 5 Schematic diagram of the abnormal time series data detection process provided by this application;

[0044] Figure 6 Network structure diagram of the abnormal time series data detection provided by this application;

[0045] Figure 7 Schematic diagram of the training process of the abnormal time series data detection network provided by this application;

[0046] Figure 8 Schematic diagram of the structure of a device for detecting abnormal time series data provided by this application;

[0047] Figure 9 Structure diagram of an electronic device provided by this application. Specific implementation manners

[0048] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0049] In recent years, deep learning-based methods have been gradually applied to the anomaly detection of time series data. The common method is to encode the sequence based on a recurrent neural network, and then combine methods such as prediction, reconstruction, or data distribution distance to locate anomaly points. However, on the one hand, when using a recurrent neural network for encoding and decoding, since the recurrent neural network can only calculate sequentially and cannot be parallelized, the training and inference speeds are slow. The slow training leads to an increase in training time and computing power costs, while the slow inference speed affects the timeliness of online detection. On the other hand, due to the uncertainty of the generation of anomalies and the large number of anomaly types, such as single-point anomalies, context-related anomalies, periodic change anomalies, trend change anomalies, etc. Different anomaly types have different requirements for features and anomaly definitions, and this method does not model the correlation relationship between signals and only uses a single reconstruction error. Therefore, it is difficult to use all anomaly types simultaneously, and the model's adaptability and generalization ability are poor. For this reason, this application provides a method for anomaly detection of time series data, and the problem of improving the accuracy, training and inference speeds, and the generalization ability of the model when performing anomaly detection of time series data remains to be further solved.

[0050] An embodiment of the present invention discloses a method for anomaly detection of time series data. Refer to Figure 1 as shown, the method includes:

[0051] Step S11: Obtain target time series data and preprocess the target time series data to obtain preprocessed target time series data.

[0052] In this embodiment, obtaining target time series data and preprocessing the target time series data to obtain preprocessed target time series data includes: obtaining target time series data and normalizing the target time series data to obtain normalized target time series data; performing a one-dimensional convolution operation on the normalized target time series data in the time direction to obtain preprocessed target time series data. Specifically, the target time series data is normalized, and the normalized target time series data is convolved in the time direction with a convolution step size of 1. In a specific implementation, the convolution kernel size can be 3, 5, or 7. Through the above technical solution, preprocessed target time series data is obtained to facilitate subsequent operations such as data embedding of the preprocessed target time series data.

[0053] Step S12: Input the preprocessed target time series data into a trained embedding network for dual-view data embedding in the time direction and the signal direction to obtain time-view embedding features and signal-view embedding features, and then fuse the time-view embedding features and the signal-view embedding features to obtain target data embedding features.

[0054] In this embodiment, the preprocessed target time-series data is input into the trained embedding network for dual-view data embedding to obtain time-view embedding features and signal-view embedding features, and then the time-view embedding features and the signal-view embedding features are fused to obtain target data embedding features. Among them, the embedding network is the Embedding network, which is specifically a dual-view data embedding network in this embodiment. Further, in a specific implementation manner, the fusion method is to directly add the time-view embedding features and the signal-view embedding features to obtain target data embedding features; in another specific implementation manner, the fusion method is to stack the channels of the time-view embedding features and the signal-view embedding features and then connect a 1×1 convolution to keep the feature channels unchanged after the superposition of the time-view embedding features and the signal-view embedding features. Through the above technical solution, the correlation features between signals are better learned through dual-view data embedding, and thus data anomalies such as multi-signal correlation anomalies can be better detected.

[0055] Step S13: Input the target data embedding features into the trained signal decomposition Transformer network to obtain the target feature representation corresponding to the target time-series data.

[0056] In this embodiment, the target data embedding features are input into the trained signal decomposition Transformer network to obtain the target feature representation corresponding to the target time-series data. Through the above technical solution, the corresponding target feature representation is obtained according to the target data embedding features, so as to facilitate the subsequent determination of the target prediction error, the target reconstruction error, and the target distribution distance based on the target feature representation, and further perform anomaly determination.

[0057] Step S14: Determine the target prediction error, the target reconstruction error, and the target distribution distance according to the target feature representation, and perform anomaly determination on the target time-series data according to the target prediction error, the target reconstruction error, and the target distribution distance.

[0058] In this embodiment, the target prediction error, the target reconstruction error, and the target distribution distance are determined according to the target feature representation, and further, it is jointly determined whether the target time-series data is abnormal according to the target prediction error, the target reconstruction error, and the target distribution distance. Through the above technical solution, the anomaly determination of the target time-series data is jointly performed by using the target prediction error, the target reconstruction error, and the target distribution distance. Based on the mechanism of multi-task learning and the fusion of multi-anomaly measurement methods, the anomaly detection ability of the model is further improved, thereby greatly improving the detection adaptability and generalization ability for different anomalies.

[0059] It can be seen that in this embodiment, when performing abnormal detection of time-series data, first, the target time-series data is collected and preprocessed, and the preprocessed target time-series data is embedded with dual-perspective data in the time direction and the signal direction to obtain the time-perspective embedding features and the signal-perspective embedding features respectively. Then, the time-perspective embedding features and the signal-perspective embedding features are fused to obtain the target data embedding features. Further, the target feature representation is determined through the signal decomposition Transformer network. According to the target feature representation, the target prediction error, the target reconstruction error, and the target distribution distance are determined, and further, the abnormal detection of the target time-series data is performed. Thus, it can be seen that when performing abnormal detection of time-series data in this application, after obtaining the target time-series data and preprocessing the target time-series data, the embedding network is used to perform dual-perspective data embedding on the preprocessed target time-series data to obtain the time-perspective embedding features and the signal-perspective embedding features. Then, the time-perspective embedding features and the signal-perspective embedding features are fused to obtain the target data embedding features. Through the dual-perspective data embedding, the correlation features between signals are better learned, and thus, it is possible to better detect data abnormal situations such as multi-signal correlation anomalies. On the other hand, when performing abnormal detection of time-series data, a relatively slow recurrent neural network is avoided, so there is an obvious advantage in terms of training and inference speed. The training time is short and the inference speed is fast, which can save the computational resource cost and improve the timeliness of feature abnormal detection. Moreover, the target prediction error, the target reconstruction error, and the target distribution distance are jointly used to determine the abnormality of the target time-series data. Based on the mechanism of multi-task learning and the fusion of multi-abnormality measurement methods, the abnormal detection ability of the model is further improved, thereby greatly improving the detection adaptability and generalization ability for different abnormalities. In summary, this application can improve the accuracy, training and inference speed, and the generalization ability of the model when performing abnormal detection of time-series data.

[0060] See Figure 2 As shown, an embodiment of the present invention discloses a specific method for abnormal detection of time-series data. Compared with the previous embodiment, this embodiment further explains and optimizes the technical solution.

[0061] Step S21: Obtain the target time-series data and preprocess the target time-series data to obtain the preprocessed target time-series data.

[0062] Step S22: Input the preprocessed target time series data into the trained embedding network for dual-perspective data embedding in the time direction and the signal direction to obtain time-perspective embedding features and signal-perspective embedding features, and then fuse the time-perspective embedding features and the signal-perspective embedding features to obtain target data embedding features. The embedding network is a dual-perspective data embedding feature extraction network module constructed in the present invention (see the dual-perspective data embedding in FIG. 6). This network module is composed of a time-perspective embedding module, a signal-perspective embedding module, a channel connection layer (concat), and a one-dimensional convolutional layer (conv1d), where the time-perspective embedding module and the signal-perspective embedding module are executed in parallel. The input data is simultaneously input into the time-perspective embedding module and the signal-perspective embedding module to extract time-perspective embedding features and signal-perspective embedding features respectively, then the features are merged through the channel connection layer, and finally, further processing is performed through the one-dimensional convolutional layer to obtain target data embedding features. Both the time-perspective embedding module and the signal-perspective embedding module adopt standard dilated causal convolution (see Figure 3 ) to implement, where the time-perspective embedding module calculates through dilated causal convolution from the time dimension, and the signal-perspective embedding module calculates through dilated causal convolution from the signal maintenance. The number of dilated causal convolution layers adopted in the present invention is set to 4, that is, a total of four layers of causal convolution Dialation (i.e., dilation) are stacked, and the number of convolutional layers is set to 1, 2, 4, 8 from bottom to top in sequence, Figure 3 only 3 layers of causal convolution are drawn in the figure.

[0063] In this embodiment, inputting the preprocessed target time-series data into the trained embedding network for dual-perspective data embedding in the time direction and the signal direction to obtain time-perspective embedding features and signal-perspective embedding features, and then fusing the time-perspective embedding features and the signal-perspective embedding features to obtain target data embedding features includes: inputting the preprocessed target time-series data into the trained embedding network, and performing dilated causal convolution operations on each signal in the time window of the target time-series data to obtain time-perspective embedding features; performing dilated causal convolution operations on the values of all signals at each time point in the time window of the target time-series data to obtain signal-perspective embedding features. Specifically, for the correlation representation between signals, use the trained dual-perspective data embedding network to perform data embedding and preliminary extraction of original features on the input target time-series data, calculate separately from the time dimension and the signal dimension. From the time perspective, perform dilated causal convolution (i.e., dilated causal convolution) operations on each signal in a time window; from the signal perspective, for each time point in a time window, perform dilated causal convolution operations on the values of all signals at this moment to obtain the correlation representation between signals. It should be noted that the network structure diagram of the dilated causal convolution module is as shown in Figure 3 shown. The number of dilated causal convolution layers adopted in the present invention is set to 4, that is, a total of four layers of causal convolution Dialation (i.e., dilation) are stacked, and the number of convolution layers is set to 1, 2, 4, 8 from bottom to top in sequence, Figure 3 and only 3 layers of causal convolution are drawn in the figure. Finally, perform a fusion operation on the results of the two perspectives by directly adding or through channel stacking and convolution to obtain target data embedding features.

[0064] Step S23: Input the target data embedding features into the trained signal decomposition Transformer network to obtain target periodic component features, target trend component features, and target residual component features, and add the target periodic component features, the target trend component features, and the target residual component features to obtain the target feature representation corresponding to the target time-series data.

[0065] In this embodiment, the target data embedding feature is input into the trained Decomposed Transformer network to obtain the target periodic component feature, the target trend component feature, and the target residual component feature, and the target periodic component feature, the target trend component feature, and the target residual component feature are added together to obtain the target feature representation corresponding to the target time series data, including: inputting the target data embedding feature into the trained signal decomposition Transformer network and obtaining the periodic component feature through the frequency domain attention module in a preset number of decomposition layers in the signal decomposition Transformer network; removing the periodic component feature from the target data embedding feature to obtain the first residual component feature, and inputting the first residual component feature into the multi-head attention module in the decomposition layer to obtain the trend component feature; removing the trend component feature from the first residual component feature to obtain the second residual component feature, and inputting the second residual component feature into the Feed-forward module in the decomposition layer to obtain the residual component feature; using the residual component feature as the input of the next decomposition layer until the target residual component feature is output after processing through all the decomposition layers, and performing channel fusion on the trend component feature and the residual component feature obtained after processing by all the decomposition layers respectively to obtain the target periodic component feature and the target trend component feature.

[0066] In this embodiment, the Decomposed Transformer (signal decomposition Transformer) is an improvement based on the standard Transformer network and forms a "signal decomposition Transformer" specifically for time series data anomaly detection. Specifically, the idea of signal decomposition is applied to the Transformer network, and multiple decomposition layers with the same structure and cascaded connections are constructed. The decomposition layer is a signal decomposition network layer redesigned based on the multi-head attention module and the feed-forward module. Among them, the multi-head attention and feed-forward are existing network modules in the Transformer network. The decomposition layer consists of three serial modules, including the frequency domain attention module, the multi-head attention module, and the feed-forward module. Among them, the frequency domain attention module is used to extract the periodic components of the signal, the multi-head attention module is used to extract the trend components of the signal, and the feed-forward module is used to extract the remaining components of the signal. The specific extraction process is as follows: the input data is first input to the frequency domain attention module to calculate the periodic components, and then the input data is subtracted from the calculated periodic component features, that is, the periodic component features are subtracted, as the input of the multi-head attention module, and the trend component features are calculated through the multi-head attention module. Similarly, the input of the multi-head attention module is subtracted from the trend component features, that is, the trend component features are subtracted as the input of the feed-forward module, and the remaining component features are calculated through the feed-forward module. This remaining component will be used as the input of the next decomposition layer for calculation. In a specific implementation, there are n decomposition layers. When all the decomposition layers are calculated, n periodic component features, n trend component features, and one target remaining component feature can be obtained. It can be understood that only the last decomposition layer outputs the target remaining component feature, and the remaining component features output by the previous decomposition layers have all been used as the input of the next decomposition layer.Further, the n periodic components and trend components are respectively subjected to component merging to obtain a final periodic component and a final trend component. The component merging refers to first performing channel fusion (i.e., concat) on the n component feature maps, then processing the features after channel fusion through a one-dimensional convolution operation (i.e., conv1d) to make the size of the feature map the same as that of each component feature map before fusion, and finally further fusing and processing through a Feed-forward module. After merging the periodic component features and trend component features, the final target periodic component features, target trend component features, and target residual component features are obtained.

[0067] It can be understood that since time series data often exhibits periodic characteristics, in order to extract the features of time series data more efficiently and accurately, a framework based on time series signal decomposition is adopted to decompose time series data into three parts: periodic components, trend components, and residual components for feature extraction respectively. This framework is integrated with Transformer and connected in an iterative manner. That is, Transformer is composed of multiple decomposition layers (i.e., decompose layer), and each decomposition layer extracts periodic component features, trend component features, and residual component features. Each decomposition layer is composed of three modules: Frequency Attention, Multi-head Attention, and Feed-Forward. Among them, Frequency Attention is used to calculate periodic component features. The network structure diagrams of the two frequency domain attention modules (Frequency Attention) provided in this embodiment are as Figure 4 shown. Both of the two frequency domain attention modules are implemented by modifying the multi-head attention module (Multi-head Attention). Figure 4 In, the part surrounded by the dashed box is the network structure of the standard multi-head attention module (Multi-headAttention). The main process is that Q, V, and K first perform linear transformations (linear) respectively, then the corresponding outputs are calculated through scaled Dot-product Attention, and finally calculations such as dropout, linear, and layer normalization (layerNorm) are performed to obtain the final result. Here, Q, K, and V respectively refer to query, key, and value, that is, the inputs of the multi-head attention module. Further, the frequency domain attention module provided by the present invention is implemented by embedding a three-step processing process of 1) performing frequency domain transformation (i.e., Fourier transform) on time domain data, 2) selecting the top_K frequency domain components in the frequency domain, and 3) then performing inverse transformation (i.e., inverse Fourier transform) to return to the time domain into the multi-head attention module. The specific embedding position is asFigure 4 as shown, where Figure 4 (a) Embed these three calculation steps at the very beginning position and add two residual connections to prevent overfitting. Figure 4 (b) Insert these three calculation steps into the calculation process of V, and also add a residual connection to prevent overfitting. Multi-head Attention is used to calculate the trend component features; the Feed-Forward component is used to calculate the remaining component features. The Multi-head Attention and Feed Forward modules are defined according to those in Transformer. The output of each decomposition layer is the remaining component features, and the remaining components of the previous layer are used as the input of the current layer. Finally, fuse the features of each layer, the periodic component features, the trend component features, and the remaining component features of the last layer. After channel stacking, connect a 1×1 convolution to maintain the original channel size to obtain the target feature representation. Further, for the calculation method of each decomposition layer, first, use the Fourier transform to extract the frequency-domain features, then take the Top-K frequency-domain components and perform the inverse Fourier transform back to the time domain, and then complete the extraction of the periodic component features through the multi-head attention layer (i.e., Multi-head attention); after removing the periodic component features from the original input, extract the trend component features through the multi-head attention layer, and then calculate the remaining component features through the removal operation and the feature extraction layer (i.e., Feed-forward).

[0068] Step S24: Determine the target prediction error, the target reconstruction error, and the target distribution distance according to the target feature representation, and perform anomaly determination on the target time series data according to the target prediction error, the target reconstruction error, and the target distribution distance.

[0069] In this embodiment, determining the target prediction error, the target reconstruction error, and the target distribution distance according to the target feature representation includes: respectively inputting the target feature representation into the trained prediction network, reconstruction network, and distribution space distance network to obtain the target prediction error, the target reconstruction error, and the target distribution distance. Performing anomaly determination on the target time series data according to the target prediction error, the target reconstruction error, and the target distribution distance includes: using the voting method or the weighted summation method according to the target prediction error, the target reconstruction error, and the target distribution distance to obtain the target error, and comparing the target error with a preset error threshold; if the target error is greater than the preset error threshold, then determine that the target time series data is abnormal data.

[0070] Specifically, the target feature representation is input into the prediction network, the reconstruction network, and the distribution space distance network simultaneously. The prediction network is a fully connected network used to calculate the data result at the next moment; the reconstruction network is a fully connected network used to reconstruct the multi-input data; the distribution space distance network pre-defines a hypersphere center based on the training data and then learns the distance distribution from the current moment data to the hypersphere center. The prediction error, reconstruction error, and distribution distance are calculated respectively according to the above three results, and the voting method or weighted summation method is used to obtain the target error, which is compared with the preset error threshold. If it is greater than the preset error threshold, it is determined as an anomaly. It can be understood that, different from the traditional single anomaly measurement strategy, in order to adapt to different types of anomalies, a method of mixing multiple anomaly measurement strategies is proposed. Through multi-task learning, the prediction, reconstruction, and data distribution space learning of data are realized simultaneously. The multi-task learning method can learn different optimization tasks on the one hand, and on the other hand, it can also improve the feature learning ability conversely. When performing anomaly determination, by integrating the three measurement strategies, various types of anomalies can be detected more effectively.

[0071] In this embodiment, the schematic diagram of the time series data anomaly detection process is as Figure 5 shown. First, the target time series data is collected and preprocessed, and the preprocessed target time series data is embedded with dual-view data in the time direction and the signal direction to obtain the time-view embedding feature and the signal-view embedding feature respectively. Then, the time-view embedding feature and the signal-view embedding feature are fused to obtain the target data embedding feature. Further, the target feature representation is determined through the signal decomposition Transformer network. According to the target feature representation, the target prediction error, the target reconstruction error, and the target distribution distance are determined, and further, the anomaly detection of the target time series data is performed. It should be noted that when performing time series data anomaly detection in this embodiment, the embedding network, the signal decomposition Transformer network, and the prediction network, the reconstruction network, and the distribution space distance network are packaged to obtain the time series data anomaly detection network. The structure diagram of the time series data anomaly detection network is as Figure 6 shown, which involves the full process functional modules and data flow from input to output. Among them, the core part is the dual-view data embedding, the feature representation based on signal decomposition, and the prediction, reconstruction, and distribution network parts for anomaly measurement. The decomposition layer first calculates the periodic component features, then removes them and calculates the trend component features, and then removes them and calculates the remaining features.

[0072] It can be understood that before performing time series data anomaly detection, it also includes training the time series data anomaly detection network. The schematic diagram of the time series data anomaly detection network training process is as Figure 7As shown, read the time-series data in the dataset for training, initialize the weights of each layer of each network, calculate the data embedding features, and further calculate the feature representation based on signal decomposition. Determine the prediction result, reconstruction result, and data distribution result, and further calculate the prediction error, reconstruction error, and data distribution distance. Perform a single optimization operation and update the model parameter weights through backpropagation. The methods that can be used for weight update include, but are not limited to, SGD, RMSProp, Adam, Nesterov Accelerated Gradient, and combinations of the foregoing methods. Continue to optimize until the termination condition is reached and determine whether to terminate the training of this branch. Among them, the termination condition can be setting the total number of optimization times, or the loss value is less than a certain preset value. Finally, save the corresponding network weights after training update and end the model training process. The training adopts an end-to-end unsupervised training method, and different stages and different branches of the model are trained synchronously, updated synchronously, and ended simultaneously.

[0073] In this embodiment, signal decomposition is applied to multi-dimensional time-series data modeling, and an anomaly detection is realized by using multi-view data embedding, frequency-domain attention feature extraction mechanism, and multi-task learning strategy. Based on the signal decomposition-based feature extraction method, it can extract the feature representation of time-series data more reasonably and accurately, with stronger feature expression ability, so it has the advantages of fast speed, high accuracy, and strong adaptability when performing time-series data anomaly detection.

[0074] See Figure 8 As shown, an embodiment of the present application discloses a time-series data anomaly detection device, including:

[0075] A preprocessing module 11, configured to obtain target time-series data and preprocess the target time-series data to obtain preprocessed target time-series data;

[0076] A data embedding module 12, configured to input the preprocessed target time-series data into a trained embedding network for dual-view data embedding in the time direction and the signal direction to obtain a time-view embedding feature and a signal-view embedding feature, and then fuse the time-view embedding feature and the signal-view embedding feature to obtain a target data embedding feature;

[0077] A feature determination module 13, configured to input the target data embedding feature into a trained signal decomposition Transformer network to obtain a target feature representation corresponding to the target time-series data;

[0078] An anomaly determination module 14, configured to determine a target prediction error, a target reconstruction error, and a target distribution distance according to the target feature representation, and perform anomaly determination on the target time-series data according to the target prediction error, the target reconstruction error, and the target distribution distance.

[0079] It can be seen that in this embodiment, when performing time-series data anomaly detection, first, the target time-series data is collected and preprocessed, and the preprocessed target time-series data is embedded with dual-view data in the time direction and the signal direction to obtain the time-view embedding features and the signal-view embedding features respectively. Then, the time-view embedding features and the signal-view embedding features are fused to obtain the target data embedding features. Further, the target feature representation is determined through the signal decomposition Transformer network. The target prediction error, the target reconstruction error, and the target distribution distance are determined according to the target feature representation, and further, the anomaly detection of the target time-series data is performed. Thus, it can be seen that when performing time-series data anomaly detection in this application, after obtaining the target time-series data and preprocessing the target time-series data, the embedding network is used to perform dual-view data embedding on the preprocessed target time-series data to obtain the time-view embedding features and the signal-view embedding features. Then, the time-view embedding features and the signal-view embedding features are fused to obtain the target data embedding features. Through the dual-view data embedding, the correlation features between signals are better learned, and thus, it is possible to better detect data anomaly situations such as multi-signal correlation anomalies. On the other hand, when performing time-series data anomaly detection, a relatively slow recurrent neural network is avoided, so there is an obvious advantage in terms of training and inference speed, with a short training time and a fast inference speed, which can save computational resource costs and improve the timeliness of feature anomaly detection. Moreover, the target prediction error, the target reconstruction error, and the target distribution distance are jointly used to determine the anomaly of the target time-series data. Based on the mechanism of multi-task learning and the fusion of multi-anomaly measurement methods, the anomaly detection ability of the model is further improved, thus greatly enhancing the detection adaptability and generalization ability for different anomalies. In summary, this application can improve the accuracy, training and inference speed, and the generalization ability of the model when performing time-series data anomaly detection.

[0080] In some specific embodiments, the preprocessing module 11 specifically includes:

[0081] The data normalization unit is used to obtain the target time-series data and normalize the target time-series data to obtain the normalized target time-series data;

[0082] The convolution unit is used to perform a one-dimensional convolution operation on the normalized target time-series data in the time direction to obtain the preprocessed target time-series data.

[0083] In some specific embodiments, the data embedding module 12 specifically includes:

[0084] A time - perspective embedding unit for inputting the pre - processed target time - series data into a trained embedding network, performing dilated causal convolution operations on each signal in the time window of the target time - series data to obtain time - perspective embedding features;

[0085] A signal - perspective embedding unit for performing dilated causal convolution operations on the values of all signals at each time point in the time window of the target time - series data to obtain signal - perspective embedding features.

[0086] In some specific embodiments, the feature determination module 13 is specifically configured to: input the target data embedding features into a trained signal decomposition Transformer network to obtain target periodic component features, target trend component features, and target residual component features, and add the target periodic component features, the target trend component features, and the target residual component features to obtain the target feature representation corresponding to the target time - series data.

[0087] In some specific embodiments, the feature determination module 13 specifically includes:

[0088] A periodic component feature determination unit for inputting the target data embedding features into a trained signal decomposition Transformer network and obtaining periodic component features through a frequency - domain attention module in a preset number of decomposition layers in the signal decomposition Transformer network;

[0089] A trend component feature determination unit for removing the periodic component features from the target data embedding features to obtain a first residual component feature, and inputting the first residual component feature into a multi - head attention module in the decomposition layer to obtain trend component features;

[0090] A residual component feature determination unit for removing the trend component features from the first residual component feature to obtain a second residual component feature, and inputting the second residual component feature into a Feed - forward module in the decomposition layer to obtain residual component features;

[0091] A component feature fusion unit for using the residual component features as the input to the next decomposition layer until all decomposition layers are processed to output target residual component features, and respectively performing channel fusion on the trend component features and the residual component features obtained after all decomposition layers are processed to obtain target periodic component features and target trend component features.

[0092] In some specific embodiments, the anomaly determination module 14 specifically includes:

[0093] An error determination unit for inputting the target feature representations into a trained prediction network, a reconstruction network, and a distribution space distance network respectively to obtain a target prediction error, a target reconstruction error, and a target distribution distance.

[0094] In some specific embodiments, the anomaly determination module 14 specifically includes:

[0095] An error comparison unit for obtaining a target error by using a voting method or a weighted summation method based on the target prediction error, the target reconstruction error, and the target distribution distance, and comparing the target error with a preset error threshold;

[0096] An anomaly determination unit for determining that the target time series data is abnormal data if the target error is greater than the preset error threshold.

[0097] Figure 9 Shown is an electronic device 20 provided by an embodiment of the present application. The electronic device 20 may specifically further include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. Among them, the memory 22 is used to store a computer program, and the computer program is loaded and executed by the processor 21 to implement the relevant steps in the time series data anomaly detection method disclosed in any of the foregoing embodiments. Additionally, the electronic device 20 in this embodiment may specifically be an electronic computer.

[0098] In this embodiment, the power supply 23 is used to provide voltage for each hardware device on the electronic device 20; the communication interface 24 can create a data transmission channel between the electronic device 20 and external devices, and the communication protocol it follows is any communication protocol applicable to the technical solution of the present application, and no specific limitation is imposed here; the input / output interface 25 is used to obtain external input data or output data to the outside, and its specific interface type can be selected according to specific application needs, and no specific limitation is made here.

[0099] Additionally, the memory 22, as a carrier for resource storage, can be a read-only memory, a random access memory, a disk, or an optical disc, etc. The resources stored thereon may include an operating system 221, a computer program 222, etc., and the storage method can be temporary storage or permanent storage.

[0100] Among them, the operating system 221 is used to manage and control each hardware device on the electronic device 20 and the computer program 222, which can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to complete the time series data anomaly detection method executed by the electronic device 20 disclosed in any of the foregoing embodiments, the computer program 222 may further include computer programs that can be used to complete other specific tasks.

[0101] Furthermore, the present application also discloses a computer-readable storage medium for storing a computer program; wherein, when the computer program is executed by a processor, the time series data anomaly detection method disclosed above is implemented. For the specific steps of this method, reference may be made to the corresponding content disclosed in the foregoing embodiments, and details will not be repeated here.

[0102] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0103] The above has introduced in detail a time series data anomaly detection method, device, equipment and medium provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for detecting anomalies in time series data, characterized in that, it includes: Obtain target time series data and preprocess the target time series data to obtain preprocessed target time series data; Input the preprocessed target time series data into a trained embedding network for dual-view data embedding in the time direction and the signal direction to obtain time-view embedding features and signal-view embedding features, and then fuse the time-view embedding features and the signal-view embedding features to obtain target data embedding features; Input the target data embedding features into a trained signal decomposition Transformer network to obtain a target feature representation corresponding to the target time series data; Determine a target prediction error, a target reconstruction error, and a target distribution distance according to the target feature representation, and perform anomaly determination on the target time series data according to the target prediction error, the target reconstruction error, and the target distribution distance; Wherein, the inputting the target data embedding features into a trained signal decomposition Transformer network to obtain a target feature representation corresponding to the target time series data includes: Input the target data embedding features into a trained signal decomposition Transformer network to obtain target periodic component features, target trend component features, and target residual component features, and add the target periodic component features, the target trend component features, and the target residual component features to obtain a target feature representation corresponding to the target time series data; The inputting the target data embedding features into a trained signal decomposition Transformer network to obtain target periodic component features, target trend component features, and target residual component features includes: Input the target data embedding features into a trained signal decomposition Transformer network and obtain periodic component features through a frequency domain attention module in a preset number of decomposition layers in the signal decomposition Transformer network; Remove the periodic component features from the target data embedding features to obtain a first residual component feature, and input the first residual component feature into a multi-head attention module in the decomposition layer to obtain trend component features; Remove the trend component features from the first residual component feature to obtain a second residual component feature, and input the second residual component feature into a Feed-forward module in the decomposition layer to obtain residual component features; Use the residual component features as the input of the next decomposition layer until the target residual component features are output after processing by all the decomposition layers, and perform channel fusion on the trend component features and the residual component features obtained after processing by all the decomposition layers to obtain target periodic component features and target trend component features.

2. The method for detecting anomalies in time series data according to claim 1, characterized in that, the obtaining target time series data and preprocessing the target time series data to obtain preprocessed target time series data includes: Obtain target time-series data and normalize the target time-series data to obtain the normalized target time-series data; Perform a one-dimensional convolution operation on the normalized target time-series data in the time direction to obtain the preprocessed target time-series data.

3. The time-series data anomaly detection method according to claim 1, characterized in that, The step of inputting the preprocessed target time-series data into a trained embedding network for dual-view data embedding in the time direction and the signal direction to obtain a time-view embedding feature and a signal-view embedding feature includes: Input the preprocessed target time-series data into a trained embedding network, and perform a dilated causal convolution operation on each signal in the time window of the target time-series data to obtain a time-view embedding feature; Perform a dilated causal convolution operation on the values of all signals at each time point in the time window of the target time-series data to obtain a signal-view embedding feature.

4. The time-series data anomaly detection method according to claim 1, characterized in that, The step of determining a target prediction error, a target reconstruction error, and a target distribution distance according to the target feature representation includes: Input the target feature representation into a trained prediction network, a reconstruction network, and a distribution space distance network respectively to obtain a target prediction error, a target reconstruction error, and a target distribution distance.

5. The time-series data anomaly detection method according to claim 1, characterized in that, The step of performing anomaly determination on the target time-series data according to the target prediction error, the target reconstruction error, and the target distribution distance includes: Adopt a voting method or a weighted summation method according to the target prediction error, the target reconstruction error, and the target distribution distance to obtain a target error, and compare the target error with a preset error threshold; If the target error is greater than the preset error threshold, it is determined that the target time-series data is abnormal data.

6. A time-series data anomaly detection device, characterized in that, comprising: A preprocessing module for obtaining target time-series data and preprocessing the target time-series data to obtain preprocessed target time-series data; A data embedding module for inputting the preprocessed target time-series data into a trained embedding network for dual-view data embedding in the time direction and the signal direction to obtain a time-view embedding feature and a signal-view embedding feature, and then fusing the time-view embedding feature and the signal-view embedding feature to obtain a target data embedding feature; A feature determination module for inputting the target data embedding feature into a trained signal decomposition Transformer network to obtain a target feature representation corresponding to the target time-series data; An anomaly determination module for determining a target prediction error, a target reconstruction error, and a target distribution distance according to the target feature representation, and performing anomaly determination on the target time-series data according to the target prediction error, the target reconstruction error, and the target distribution distance; Among them, the feature determination module is specifically configured to: input the target data embedding feature into the trained signal decomposition Transformer network to obtain a target periodic component feature, a target trend component feature, and a target residual component feature, and add the target periodic component feature, the target trend component feature, and the target residual component feature to obtain a target feature representation corresponding to the target time series data; The feature determination module specifically includes: A periodic component feature determination unit, configured to input the target data embedding feature into the trained signal decomposition Transformer network and obtain a periodic component feature through a frequency domain attention module in a preset number of decomposition layers in the signal decomposition Transformer network; A trend component feature determination unit, configured to remove the periodic component feature from the target data embedding feature to obtain a first residual component feature, and input the first residual component feature into a multi-head attention module in the decomposition layer to obtain a trend component feature; A residual component feature determination unit, configured to remove the trend component feature from the first residual component feature to obtain a second residual component feature, and input the second residual component feature into a Feed-forward module in the decomposition layer to obtain a residual component feature; A component feature fusion unit, configured to use the residual component feature as the input of the next decomposition layer until the target residual component feature is output after being processed by all the decomposition layers, and perform channel fusion on the trend component feature and the residual component feature obtained after being processed by all the decomposition layers respectively to obtain a target periodic component feature and a target trend component feature.

7. An electronic device, Characterized in that, It includes: A memory for storing a computer program; A processor for executing the computer program to implement the steps of the time series data anomaly detection method according to any one of claims 1 to 5.

8. A computer-readable storage medium, Characterized in that, For storing a computer program; wherein, when the computer program is executed by a processor, the steps of the time series data anomaly detection method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Transform-based industrial equipment fault prediction method and device

    CN115186904A

  • Multivariable time sequence anomaly detection method and system, medium, equipment and terminal

    CN115222141A

Cited By

  • A user abnormal behavior detection method and system

    CN122863342A