Generative adversarial network-based anomaly detection method for multivariate time series data

By adopting a GAN-based multivariate time series data anomaly detection method in the network physical system, and using LSTM-RNN to capture the time correlation of the time series, the problem of difficulty in dealing with nonlinear correlation and dynamics in the prior art is solved, and higher anomaly detection accuracy and sensitivity are achieved.

CN119989184APending Publication Date: 2025-05-13BEIJING GUOTENG INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311498277.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively handle the nonlinear correlation and dynamics in multivariate time series data in network physical systems, resulting in a lack of sensitivity and accuracy in anomaly detection.

Method used

The multivariate time series data anomaly detection method based on Generative Adversarial Network (GAN) is adopted, and the long and short-term memory recurrent neural network (LSTM-RNN) is used as a generator and discriminator to capture the time correlation of the time series in the GAN framework, and the exception score is calculated by combining the outputs of the generator and discriminator.

Benefits of technology

This method can better adapt to the characteristics of multivariate time series data in network physical systems, capture abnormal patterns, improve the accuracy and sensitivity of abnormal detection, thereby ensuring the safety and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention provides an unsupervised multivariate time series data anomaly detection method based on a generative adversarial network. The implementation process of the method comprises the steps of collecting and preprocessing data, constructing and training a proposed method model, calculating an anomaly score by using a pre-trained model to carry out anomaly detection, and evaluating the detection performance of the model so as to carry out optimization. The proposed method uses a long short-term memory recurrent neural network (LSTM-RNN) as a base model (i.e., a generator and a discriminator) to capture temporal correlation of time sequence distribution in a GAN framework. In addition, the multivariate time series data anomaly detection method provided by the invention does not independently process each data stream, but considers the whole variable set at the same time to capture potential interaction between variables. Meanwhile, a generator and a discriminator of GAN training are fully utilized, and a novel anomaly score is used to detect anomalies by combining the discriminator and a reconstruction error. By implementing the method provided by the invention, the abnormal behavior in the network physical system can be effectively detected, and timely alarm and response are provided for a system operator, so that the safety and reliability of the system are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] The present invention proposes an unsupervised multivariate time series data anomaly detection method based on a generative adversarial network, which can effectively report anomalies caused by various network intrusions, thereby providing a solid guarantee for the security of cyber-physical systems. Background Art

[0002] Today’s cyber-physical systems (CPS) include complex and large systems such as smart buildings, factories, power plants, and data centers. They are equipped with networked sensors and actuators that generate large amounts of multivariate time series data. This data can be used to continuously monitor the working status of security systems and detect anomalies in a timely manner so that operators can take action to investigate and resolve potential problems. With the popularity of the Internet of Things, the application of networked sensors and actuators in consumer product safety systems and other systems will become more common. This results in multiple systems and devices that can communicate autonomously over the network and perform various tasks. However, since many cyber-physical systems are designed for critical missions, they become a prime target for cyber attacks. In order to prevent intrusion incidents, it is important to closely monitor the behavior of these systems by performing anomaly detection using the multivariate time series data generated by the system. This approach can help detect any abnormal activities in a timely manner and ensure the safe operation of the system.

[0003] The basic task of anomaly detection is to identify points in certain time steps where the system behavior is significantly different from the previous normal state. Traditional statistical process control methods are widely used to monitor the quality of industrial processes in order to detect out-of-range operating conditions. However, these traditional detection techniques are unable to handle the multivariate data streams generated by modern cyber-physical systems, which have an increasingly dynamic and complex nature. To overcome these limitations, researchers have begun to use machine learning techniques to exploit the large amount of data generated by the system for anomaly detection. Due to the lack of labeled data, anomaly detection is usually regarded as an unsupervised machine learning task. However, most of the existing unsupervised methods are built on linear projections and transformations and cannot handle the nonlinear correlations hidden in multivariate time series. In addition, due to the highly dynamic nature of cyber-physical systems, most current techniques detect anomalies by simply comparing the current state with the predicted normal range, which is far from sufficient. Therefore, in order to effectively deal with the anomaly detection problem in cyber-physical systems, new methods need to be developed. These methods should be able to capture the nonlinear correlations in multivariate time series and take into account the dynamic nature of the system. This will help to improve the sensitivity to changes in the behavior of cyber-physical systems and detect potential anomalies in advance to ensure the safe operation of the system.

[0004] The Generative Adversarial Network (GAN) framework builds generative deep learning models through adversarial training. Early studies have proven the effectiveness of GAN in generating time series data. It performs well in generating real complex data sets, and simultaneously trains the generator and discriminator in an adversarial manner, making it outstanding in anomaly detection. Therefore, the present invention proposes an innovative multivariate time series data anomaly detection method based on generative adversarial networks. The method is able to model complex multivariate correlations between multiple data streams and detect anomalies using a GAN-trained generator and discriminator. Unlike traditional classification methods, the GAN-trained discriminator learns how to detect fake data from real data in an unsupervised manner. This novel method can better adapt to the characteristics of multivariate time series data in cyber-physical systems and capture abnormal patterns by learning the distribution of data. By utilizing the adversarial training of the generator and discriminator of GAN, the method can effectively detect abnormal behaviors in cyber-physical systems and provide timely alerts and responses to system operators, thereby ensuring the safety and reliability of the system. Summary of the invention

[0005] In response to the above problems, this patent proposes a multivariate time series data anomaly detection method based on generative adversarial networks. The method uses long short-term memory recurrent neural networks as the basic model (i.e., generator and discriminator) to capture the temporal correlation of time series distribution in the GAN framework. In addition, instead of processing each data stream independently, the proposed method considers the entire set of variables at the same time to capture potential interactions between variables. We also make full use of the generator and discriminator of GAN and use a combined anomaly score to calculate the anomaly score. The overall process of the proposed method is as follows Figure 1 As shown, it mainly includes the following steps:

[0006] S1: Collect training datasets, which are used to train the proposed anomaly detection model. Then preprocess the collected time series datasets, normalize the data to ensure that the data has a similar scale and range, extract key features from the data, and convert the data into vectors to input into the model.

[0007] S2: Construct the proposed anomaly detection model structure, which is based on the GAN framework structure. This framework includes two key components: the generator and the discriminator. The generator and the discriminator are constructed as two long-term and short-term recurrent neural networks. The collaborative work of the generator and the discriminator enables the model to learn the characteristics of abnormal data and perform accurate anomaly detection.

[0008] S3: After the anomaly detection model is built, we will use the preprocessed dataset to train the model. By inputting the preprocessed data into the trained model, we can calculate the anomaly score based on the difference between the model's output and the actual data. This difference value can be used as part of the anomaly score to determine whether the sample is abnormal. In addition to the difference value, we also combine the discriminant error of the discriminator to calculate the final anomaly score. During the training process, the model will gradually learn how to distinguish between normal and abnormal samples, thereby improving the accuracy of anomaly detection.

[0009] S4: After completing the model training, we can preprocess the test data set in the same way and input it into the trained model for anomaly detection. The model will calculate the anomaly score based on the input data and compare it with the preset threshold. If the anomaly score exceeds the threshold, we can determine that the data is abnormal.

[0010] S5: After completing anomaly detection, the performance of the model needs to be evaluated and optimized. This includes calculating the model's accuracy, recall, precision and other indicators to evaluate the overall effect of the model. Based on the evaluation results, adjust the model's parameters, increase the amount of training data or improve the model's architecture to further improve the performance of anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the content of the invention of this patent and the technical solutions in the embodiments, the drawings used will be briefly introduced below.

[0012] Figure 1 It is the overall framework process of the anomaly detection method;

[0013] Figure 2 It is the framework process of the model training phase;

[0014] Figure 3 The framework process for anomaly detection; DETAILED DESCRIPTION

[0015] In order to make the purpose, technical solution and advantages of the embodiments of this patent clearer, the technical solution in the embodiments of this patent will be clearly and completely described below in conjunction with the drawings in the embodiments of this patent. Obviously, the described embodiments are part of the embodiments of this patent, not all of them. Based on the embodiments in this patent, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this patent.

[0016] This patent uses the SWaT and WADI datasets for training. In the SWaT dataset, 51 variables were measured for 11 days. In the original data, 496,800 samples were collected under normal working conditions (data collected in the first 7 days), and 449,919 samples were collected when various cyber attacks were inserted into the system. Similarly, for the WADI dataset, 789,371 samples of 103 variables were collected under normal working conditions in the first 14 days, and 172,801 samples were collected when various cyber attacks were inserted into the system in the last 2 days. For both datasets, we eliminated the first 21,600 samples from the training data (normal data).

[0017] We construct the generator and discriminator of GAN as long short-term recurrent neural network (LSTM-RNN), as follows Figure 2 As shown. Following the typical GAN ​​framework, the generator (G) generates a fake time series with a sequence in a random latent space as input, and passes the generated sequence samples to the discriminator (D), which attempts to distinguish the generated data sequence from the actual normal training data sequence. Instead of processing each data stream independently, the method proposed in this invention considers the entire set of variables simultaneously so that the potential interactions between the variables can be captured into the model. Before discrimination, we use a sliding window to divide the multivariate time series into subsequences. In order to determine the optimal window length for subsequence representation, we use different window sizes to capture the system state at different resolutions, i.e., s w =30×i,i=1,2,...,10. To capture the relevant dynamics of SWaT data, the window is applied to the normal and test datasets with a shift length of 10. We use a LSTM network with a depth of 3 and 100 hidden units as the generator. We use a LSTM network with a depth of 1 and 100 hidden units as the discriminator.

[0018] like Figure 3 As shown, based on the mapping of the generator from the real space to the latent space, we use the generator to calculate the residual between the real test sample and the reconstructed sample, and use the discriminator to classify the time series. The test sample is mapped back to the latent space, and the corresponding reconstruction loss is calculated by the difference between the reconstructed test sample and the actual test sample. At the same time, the test sample is also fed to the trained discriminator to calculate the loss of the discriminator. We combine these two losses to calculate the anomaly score to detect potential anomalies in the data.

[0019] Given that the trained discriminator D can distinguish fake data (i.e., anomalies) from real data with high sensitivity, it can be used as a direct tool for anomaly detection. The trained generator G is able to generate real samples, which is actually a mapping from the latent space to the real data space: G(Z):Z→X, which can be regarded as an implicit system model reflecting the normal data distribution. If the inputs in the latent space are close, the generator outputs similar samples. Therefore, if it is possible to test Find the corresponding Z in the latent space k , then X test With G(Z k The similarity between ) (i.e., the reconstructed test samples) can explain X test To what extent does it follow the distribution reflected by the following formula? In other words, we can also use X test and G(Z k ) to identify anomalies in the test data. To find the best Z corresponding to the test sample k , we first sample a random set Z from the latent space 1 , and obtain the reconstructed original sample G(Z by inputting it into the generator 1 ). Then, we use test and the error function defined by G(Z) The obtained gradient is used to update the sample in the latent space. After enough iterations, the error is small enough, and the sample Z k is recorded as the corresponding mapping of the test sample in the latent space. The residual of the test sample at time t is obtained by Calculate, where is the measured value of n variables at time step t. The anomaly detection loss can be obtained by calculate.

[0020] Based on the above description, the discriminator and generator trained by GAN will output a set of anomaly detection losses for each test data subsequence We calculate the combined discriminative and reconstructed anomaly score by mapping the anomaly detection loss of the subsequence back to the original time series, as follows:

[0021]

[0022] where t∈{1,2,...,N},j∈{1,2,...,n}, and s∈{1,2,...,s w}.

[0023] After calculating the anomaly score, the anomaly score is compared with the selected threshold to detect abnormal data.

[0024] In addition, we use standard metrics, namely precision (Pre), recall (Rec), and F1 score to evaluate the anomaly detection performance of the model, and the corresponding calculation formulas are as follows:

[0025]

[0026]

[0027]

[0028] TP is the number of correctly detected anomalies, FP is the number of incorrectly identified anomalies, and FN is the number of incorrectly identified normals. The performance of the model is measured by calculating the above indicators to continuously optimize the anomaly detection effect of the model.

[0029] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this patent, rather than to limit it. Although this patent has been described in detail with reference to the aforementioned embodiments, ordinary technicians in this field should understand that they can still modify the technical solutions recorded in the aforementioned embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of this patent.

Claims

1. A method for anomaly detection of multivariate time series data based on generative adversarial networks, characterized by: A GAN-based anomaly detection model is constructed, with long short-term memory recurrent neural network as the basic model of the generator and discriminator. The discriminator error is combined with the generator reconstruction error to calculate the anomaly score to detect anomalies.

2. The method according to claim 1, characterized in that: It consists of data collection and preprocessing, construction and training of anomaly detection models, anomaly score calculation for anomaly detection, and model evaluation to optimize model detection results.

3. The data collection and preprocessing module according to claim 2, characterized in that: Obtain time series data generated during daily operation from cyber-physical systems, and subject the systems to a variety of different attack methods and durations to obtain a variety of data.

4. The construction of the anomaly detection model according to claim 2, characterized in that: The generator and discriminator networks in the original GAN ​​framework are replaced with a two-layer long short-term memory recurrent neural network.

5. The calculation of the anomaly score according to claim 2, characterized in that: The error between the generated data of the trained generator and the real data and the discriminant error of the trained discriminator are combined as the total anomaly score of the entire anomaly detection.