Industrial Control Intrusion Detection Method Based on Graph Attention Network and Variational Autoencoder
Through the industrial control intrusion detection method based on graph attention network and variational autoencoder, the problems of insufficient detection accuracy and high false alarm rate in the industrial control system are solved, and efficient and accurate intrusion detection and rapid abnormal positioning are achieved, adapting to complex high-dimensional data and rapidly changing operating modes.
Patent Information
- Application Number
- CN202411889763.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2044-12-20
AI Technical Summary
The existing deep learning models have problems in industrial control systems with insufficient detection accuracy, high false alarm rate, insufficient interpretability of equipment correlation extraction and model abnormality, especially in large and small-to-medium industrial control systems, which are difficult to adapt to complex high-dimensional data and rapidly changing operating modes.
The industrial control intrusion detection method based on graph attention network and variational autoencoder is adopted, and the reconstruction error is enhanced through memory modules, combined with incremental update and gated update mechanisms, and the abnormal score calculation is optimized using feature difference loss function, and the time and spatial characteristics of sensor data are extracted through graph attention network and gated loop unit, and the two-stage discrimination method of threshold and approximate entropy is combined to reduce the false positive rate.
It improves the accuracy and adaptability of industrial-controlled intrusion detection, reduces the false alarm rate, and provides interpretability of equipment abnormalities, and quickly locates the cause of system abnormalities.
Smart Images

Figure CN119728238B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial control system intrusion detection. Specifically, it relates to an industrial control intrusion detection method based on graph attention network and variational autoencoder. Background Art
[0002] The intrusion detection system is one of the effective protection means for industrial control systems against network attacks. The intrusion detection system can monitor system data in real time and identify security threats in the system through effective detection algorithms. With more and more researchers participating in the field of industrial control system intrusion detection, intrusion detection technologies have gradually become diversified, and machine learning and deep learning have become research hotspots in this field in recent years. Machine learning is often used to establish a behavior model of normal operations in the system. When any behavior deviating from the normal model is detected, it may be marked as abnormal, thereby determining the existence of an intrusion behavior. Machine learning algorithms can also be used to identify and extract security-related features in network traffic or system logs to help more accurately identify malicious activities. In addition, with the changes in the network environment and attack strategies, machine learning models can adapt to these changes through continuous learning, thereby maintaining their detection efficiency. However, with the increasing complexity of the network environment, the data in industrial control systems also exhibits the characteristics of high-dimensionality and a large amount of noise, which poses a challenge to the accuracy of intrusion detection. Deep learning has superiority in processing complex high-dimensional data and is therefore widely used in the field of industrial control intrusion detection and shows excellent detection performance.
[0003] Although the deep learning method has shown significant performance advantages in intrusion detection, it still has certain limitations in specific industrial control system application scenarios. This is mainly due to the scale of industrial control systems and the diversity of application scenarios - systems of different scales show obvious differences in response to their unique application requirements. For large-scale industrial control systems, although there are sufficient computing resources to deploy large deep learning models, due to the large number of system devices, their industrial control data has complex high-dimensional spatio-temporal characteristics, and existing deep learning models still have problems with insufficient detection accuracy when capturing this feature; at the same time, the binary anomaly indicators output by most models lack sufficient detailed information and interpretability, making it difficult to identify and locate anomalies, increasing the time cost of fault diagnosis; in addition, there is a large amount of noise in the large-scale industrial control system environment, resulting in a high false alarm rate of the model. For small and medium-sized industrial control systems, although deep learning models can ensure high accuracy, due to the complexity of their models and the large amount of computing resources required for long-term training, it is difficult to deploy these high-accuracy intrusion detection models in the actual working environment; at the same time, due to the diversity of production modes in small and medium-sized industrial control systems, the device operation modes change frequently. When the operation mode of the industrial control system changes, the previously trained model can no longer be effectively used. Summary of the Invention
[0004] The purpose of the present invention is to provide an industrial control intrusion detection method based on a graph attention network and a variational autoencoder to solve the above problems in the prior art.
[0005] The present invention is realized through the following technical solutions:
[0006] An industrial control intrusion detection method based on a graph attention network and a variational autoencoder, comprising:
[0007] Obtain the data of the industrial control system and preprocess the data. The preprocessing includes data cleaning and normalization of the cleaned data; divide the data using a sliding window;
[0008] Build an ME-VAE model. The ME-VAE model includes an encoder, a memory module, and a decoder. Enhance the latent features through the memory module, amplify the reconstruction error of intrusion samples, and incrementally update the memory in the ME-VAE model;
[0009] Judge whether to use the current data to update the memory. If so, update several memory items. If not, do nothing;
[0010] Obtain the difference between the latent features and the features enhanced by the memory module, weight the reconstruction error of the data, and determine intrusion through the weighted reconstruction error, and output the determination result.
[0011] Preferably, building the ME-VAE model includes:
[0012] Given a time input sequence, encode and sample it through an encoder composed of a GRU module and a fully connected layer, and output the latent representation of the intermediate layer;
[0013] Use the latent representation as a query vector, calculate the similarity with each memory item in the memory module, and combine the query vector and the memory item into a new latent feature vector according to the similarity;
[0014] Pass the latent feature vector to the decoder for reconstruction to restore the time input sequence.
[0015] Preferably, enhancing the latent features includes:
[0016] Calculate the distance between the query vector and each unit in the memory module to obtain a weight vector;
[0017] w i = Softmax(d cos (z, m i ))
[0018]
[0019] And the weight vector is corrected through vector sparsification;
[0020]
[0021] In the formula, w i is the first weight vector, d cos (z, m i ) is the cosine similarity between the latent feature vector and the memory item, z is the latent feature vector, and m i is the unit in the memory module. is the corrected weight vector, λ is the sparsity threshold, ε is a positive scalar, and max(., 0) is the ReLU activation function.
[0022] Preferably, the incremental update of the memory includes:
[0023]
[0024] In the formula, v' i is the second weight vector, v i is the matching probability, is the new latent feature vector, v j is the j-th probability among all possible values after softmax calculation, is the cosine similarity between the new latent feature vector and the memory item, and i, j, and L are natural numbers.
[0025] Preferably, the update of the memory item includes:
[0026] Calculate the distance score between the current input data and the memory item, and update the memory module when the distance score is less than the update threshold;
[0027] The calculation of the distance score between the current input data and the memory item includes:
[0028]
[0029] In the formula, D(z, m i ) is the distance between the latent feature vector and the memory item, γ is the exponentially weighted factor, k and i are natural numbers, and Score(z) is the distance score;
[0030] Obtain k items with non-zero first weight vectors after querying the memory module for the latent feature vector, and arrange them in ascending order.
[0031] Preferably, it further includes establishing a loss function to balance the training effect of the ME-VAE model, and the loss function includes:
[0032]
[0033] In the formula, is the mean square error, m is the number of data, y i is the predicted value, y' i is the actual value, is the divergence loss, is the loss function of VAE reconstruction, is the difference loss function, is the Euclidean distance, is the total number of all features, μ is the mean, σ is the variance, m j is the j-th latent feature representation vector, η is the hyperparameter of the weighted loss function of.
[0034] Preferably, it further includes adjusting the reconstruction error by the difference between the input and output features of the memory module:
[0035]
[0036] In the formula, S(t) is the adjusted reconstruction error, ε ′ is the correction parameter, x t is the input data feature, is the data feature after reconstruction, z t is the initial feature, is the feature recombined and generated by the memory module.
[0037] Preferably, it further includes:
[0038] Based on the graph attention network and the gated recurrent unit, a prediction model is established. The prediction model predicts the data of the industrial control system and filters and denoises the predicted data.
[0039] Preferably, the prediction model includes:
[0040]
[0041] In the formula, is the feature of node i after aggregating neighbor nodes, α ij is the attention coefficient of node j to node i in the GAT network, W is the learnable weight matrix, Ψ(i) is the neighbor of node i obtained from the adjacency matrix, is the sliding window input with node i, is the sliding window input with node j, is the concatenation of the sensor embedding and its corresponding sliding window input, v i ′ is the feature embedding vector of node i, e ij The attention coefficient after activation using the non-linear activation function LeakyReLU, eik is the unnormalized value of the attention coefficient between node i and its neighbor node k, a is a learnable parameter vector, z ′(t) is the feature vector of n sensors containing spatio-temporal features in the sliding window, x (t) is the industrial control system data, z (t) is the output of the GAT network.
[0042] Preferably, the filtering and denoising of the predicted data includes:
[0043] Obtain the error sequence, perform phase space reconstruction on the error sequence with a given embedding dimension for the error sequence within the filtering window, and divide the error sequence into several subsequences;
[0044] Calculate the similarity between any two reconstructed vectors. The distance between two reconstructed vectors is defined as the maximum absolute value of the difference between the corresponding elements of the two sequences;
[0045] Calculate the logarithmic probability of all sequences and obtain the average value. Calculate the approximate entropy of the error sequence within the filtering window through the average value;
[0046] Compare the approximate entropy with the entropy threshold, judge and output abnormal data.
[0047] The technical solution of the present invention has at least the following advantages and beneficial effects:
[0048] 1. Aiming at the problems that existing lightweight intrusion detection algorithms cannot balance detection accuracy and efficiency and are difficult to adapt to the changing operation modes of medium and small industrial control systems. Through the method provided by the present invention, the problem that intrusion data similar to normal data is difficult to identify due to small reconstruction errors is solved. This solution uses a memory module to expand the reconstruction error and improve the discrimination ability. Then, a strategy based on incremental update and gated update mechanism is introduced, which is applied to the offline training and online detection stages respectively to enhance the adaptive ability of the algorithm to the system operation mode, and the storage space of the memory module is reduced by combining the feature difference loss function. In addition, by adding weighted reconstruction errors to the feature differences, the abnormal score calculation is optimized, further improving the accuracy of intrusion detection.
[0049] 2. Aiming at the problems that existing intrusion detection algorithms have deficiencies in device correlation extraction and anomaly interpretability of the model and a relatively high false alarm rate, through the method provided by the present invention, the temporal and spatial features of sensor data are extracted by fusing the Graph Attention Network (GAT) and the Gated Recurrent Unit (GRU), making up for the deficiencies of traditional intrusion detection algorithms in capturing the dependencies between sensors. To solve the problem that the original GAT requires a prior graph structure and the graph update process lacks continuity, a graph structure update mechanism based on similarity constraints is proposed. At the same time, to solve the problem of high false alarm rate caused by environmental noise, a two-stage composite discrimination method combining threshold and entropy is proposed. This method first uses the threshold for preliminary discrimination, and then uses approximate entropy to filter out the anomalies caused by noise, thereby reducing the false alarm rate. In addition, in order to enable technicians to quickly locate the cause of system anomalies, an anomaly explanation method based on aggregated neighbor anomaly scores is proposed. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0051] Figure 1 The overall architecture of the ME-VAE model of the present invention;
[0052] Figure 2 The specific structure of the ME-VAE model of the present invention;
[0053] Figure 3 The feature recombination enhancement process of the present invention;
[0054] Figure 4 The latent representation sparsification process of the present invention;
[0055] Figure 5 The effect comparison diagram of the present invention and other models;
[0056] Figure 6 The overall framework of the SC-GRN of the present invention;
[0057] Figure 7 The continuous graph structure update process of the present invention;
[0058] Figure 8 The SC-GRN data prediction model of the present invention;
[0059] Figure 9 The schematic diagram of system anomaly explanation of the present invention. Detailed implementation manners
[0060] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and illustrated herein generally may be arranged and designed in a variety of different configurations.
[0061] Please refer to Figure 1 and Figure 2 , an industrial control intrusion detection method based on a graph attention network and a variational autoencoder, including:
[0062] S101: Obtain the data of the industrial control system and preprocess the data. The preprocessing includes data cleaning and normalizing the cleaned data; dividing the data using a sliding window;
[0063] After obtaining the original sensor data from the industrial control system, it is necessary to preprocess the data to ensure the quality and applicability of the data and provide a solid foundation for subsequent analysis. Specifically, it is divided into three steps: data cleaning, data normalization, and sliding window sampling.
[0064] S102: Establish an ME-VAE model. The ME-VAE model includes an encoder, a memory module, and a decoder. Enhance the latent features through the memory module, amplify the reconstruction error of intrusion samples, and incrementally update the memory in the ME-VAE model;
[0065] This step is the offline training stage of the model. In this stage, the latent features are mainly enhanced through the memory module to make them closer to normal features, thereby amplifying the reconstruction error of intrusion samples. Then, the memory is incrementally updated during model training. When the model is offline trained, since most of the data is normal data, the model pays more attention to the influence of noise data at this time. Therefore, the memory items are directly updated using the incremental update method.
[0066] S103: Determine whether to use the current data to update the memory. If so, update several memory items; if not, do nothing;
[0067] This step is the online testing stage of the model. In this stage, since the operating mode of the system may change, the model pays more attention to the change in data distribution at this time. The gating mechanism is used to determine whether to use the current data to update the memory and update some memory items. Through this update method, the model can have an adaptive ability when the data mode changes, rather than retraining the model.
[0068] S104: Obtain the difference between the potential features and the enhanced features of the memory module, weight the reconstruction error of the data, determine intrusion through the weighted reconstruction error, and output the determination result.
[0069] The present invention first uses a GRU and a fully connected layer to form a VAE network to accelerate the convergence speed of the model while extracting the time series features of the data. Secondly, aiming at the problem of noise in the normal data participating in training, a method of incrementally updating the memory is proposed; aiming at the problem that the operation mode of small and medium-sized systems may change, a gating mechanism is proposed to selectively update the memory to improve the self-adaptability of the model to changing data. Subsequently, the feature difference loss is used to increase the difference of the features in the memory module and reduce the consumption of the memory module. Finally, it is proposed to combine the reconstruction error and the feature error to calculate the anomaly score, further magnify the difference between the normal data and the abnormal data, and improve the intrusion detection accuracy.
[0070] In an exemplary embodiment of the present invention, data cleaning is one of the key steps in data preprocessing, which involves identifying and processing abnormal, missing or inconsistent parts in the data. Since time series data has time dependence, missing values are generally not directly deleted, but filled with the previous data or the next data of the missing data point.
[0071] Data normalization. In industrial multivariate time series analysis, due to the large numerical differences between different sensors. Data with larger values will increase the scale interval, resulting in the model being insensitive to changes in smaller values. Therefore, it is necessary to scale the data so that the data has the same scale. This section uses the maximum-minimum normalization method to alleviate the adverse effects brought by the large numerical differences between variables in time series data and normalize each dimension of the time series data.
[0072] Divide the data using a sliding window. Since sensor data usually contains a large number of time points, it is impossible to input the complete time series into the intrusion detection model at the same time. To capture the time dependence between the data, continuous data segments are obtained using a sliding window with a width of w and a step size of l as the model input.
[0073] is the time step for each movement of the sliding window on the time axis. To capture the dependence of time series data and increase the number of training samples, thereby reducing information loss, the usual practice is to have a certain overlap between adjacent windows. In the test stage, to reduce the computational cost and better simulate the actual situation, the step size l of the training set of this solution is set to 1, and the l of the test set is set to the same as the length of the sliding window w.
[0074] In the intrusion detection of industrial control systems, the system operates in the normal mode most of the time, and real intrusion events are relatively rare. Traditional intrusion detection models identify abnormal behaviors by learning the characteristics of normal operation data. However, when there are only minor differences between intrusion data and normal data, these intrusion activities are difficult to detect. The abnormal data points of these three types of sensors during intrusion are all concentrated around the median line, with a small gap from normal data, and the model often has difficulty distinguishing such intrusion samples.
[0075] To solve this problem, the present invention stores the data characteristics during normal operation through a memory module and uses these characteristics to optimize the latent space of the VAE, making the latent representation closer to the characteristics of normal behaviors, thereby amplifying the reconstruction error of intrusion samples.
[0076] As the basic network structure of the reconstruction model, VAE mainly uses an encoder to compress sample data into the latent space and reconstructs the representation in the latent space into the original data through a decoder. Moreover, VAE has stronger generalization ability than traditional autoencoders and can better perform data reconstruction and generation when facing unseen data.
[0077] The ME module is a key part of the memory-enhanced lightweight variational autoencoder (ME-VAE), which records the latent characteristics of normal data. This module compares the latent variables of the input data with the memory items, then selects similar features for combination, and finally uses this combination as the input of the decoder to reconstruct the data. This process helps to reduce the impact of abnormal data because the continuously updated memory items during the training process will make the reconstruction of intrusion samples more similar to normal samples, resulting in a higher reconstruction error. ME-VAE improves the detection sensitivity to minor anomalies in this way while maintaining the resource efficiency of the system.
[0078] An exemplary implementation of the present invention for establishing an ME-VAE model includes:
[0079] S201: Given a time input sequence, it is encoded and sampled by an encoder composed of a GRU module and a fully connected layer, and the latent representation of the intermediate layer is output;
[0080] S202: Using the latent representation as a query vector, calculate the similarity with each memory item in the memory module, and combine the query vector and the memory item into a new latent feature vector according to the similarity;
[0081] S203: Pass the latent feature vector to the decoder for reconstruction to restore the time input sequence.
[0082] In this embodiment, in order to implement a lightweight intrusion detection model, both the encoder and decoder of the VAE have only one layer. Among them, the encoder consists of a GRU module and two fully connected layers. Since industrial control sensor data has the characteristic of time series, GRU can learn the temporal dependencies hidden in the time window samples.
[0083] Compared with using LSTM as the encoder network in the existing technology TSMAE, GRU has higher computational efficiency and fewer parameters. This is because GRU updates the state only through the update gate and the reset gate, while LSTM has three gates, so GRU further meets the requirements of the lightweight model. After GRU captures the time series features in the data, a fully connected layer is usually added to convert the output of the GRU layer into the mean μ and variance σ of the latent space, and these parameters will be used for sampling the latent variables. Another fully connected layer samples the intermediate layer in a reparameterized manner. Specifically,
[0084] z = μ + σ ⊙ ζ
[0085] In the formula, z is the latent vector finally generated by the encoder, ζ is a random variable, and the fully connected layer uses RELU as the activation function.
[0086]
[0087] In the formula, x is the input parameter. The decoder of the VAE consists of a GRU layer and a fully connected layer. The GRU layer is similar to that in the encoder and is used to convert the latent variable into a time-series output sequence. The final fully connected layer is used to map the latent variable to an output sequence with the same dimension as the input data.
[0088] As Figure 3 shown, an exemplary embodiment of the present invention for enhancing latent features includes:
[0089] Reorganize the latent representation through the attention mechanism so that the reorganized latent representation can be closer to the normal features, thereby amplifying the reconstruction error of abnormal samples and effectively improving the detection effect of the model.
[0090] Calculate the distance between the query vector and each unit in the memory module to obtain the weight vector;
[0091] w i = Softmax(d cos (z, m i ))
[0092]
[0093] Then, the similarity is normalized by the Softmax function. Since memory cells with high similarity represent more likely normal patterns of samples, the similarity is used to weight and synthesize a new latent representation.
[0094]
[0095] As Figure 4 shown, the above process uses a limited number of memory items in the memory module to resynthesize latent features, resulting in a large reconstruction error for intrusion samples and making them easy to detect. However, some abnormal features may also be reconstructed through complex combinations of many memory items. To avoid this problem, vector sparsification is used for correction;
[0096]
[0097] where w i is the first weight vector, d cos (z, m i ) is the cosine similarity between the latent feature vector and the memory item, z is the latent feature vector, m i is the cell in the memory module, is the corrected weight vector, λ is the sparsity threshold, ε is a positive scalar, max(., 0) is the ReLU activation function, M is the dimension of the memory module, is the new latent feature vector.
[0098] At the same time, considering that as a weight factor, their sum must be 1, the weight matrix is normalized. This is because the physical meaning of the weight matrix is to combine different memory items into a new latent feature, and each item in the weight vector represents a composition in a certain proportion. After normalization, finally, through calculation, a new latent representation is obtained. This process realizes the composition of the latent features of the data through the memory items of the memory module. This method can make intrusion samples generate a large reconstruction error, where is the normalized latent feature.
[0099] An exemplary implementation of the present invention for incremental update of memory includes:
[0100] The potential features are recombined through the memory module to make them closer to the normal mode features, thus amplifying the reconstruction error of the model for intrusion samples. However, in real industrial control scenarios, the normal data collected is often not all normal sample data, and there may be noise or other abnormal situations. Since the memory module will regard the noisy data as normal data for learning during training and store it in the memory module, the detection performance of the model will decline. Therefore, this section proposes a method for incrementally updating the memory module. By querying the weighted average of, rather than directly adding them, more attention can be paid to similar terms. Prevent the memory module from being affected by noise during training, thus avoiding learning rare abnormal features. The update method is as shown in the formula as follows:
[0101]
[0102] In the formula, v i ′ is the second weight vector, v i is the matching probability, is the new potential feature vector, v j is the j-th probability among all possible values after softmax calculation, is the cosine similarity between the new potential feature vector and the memory item. i, j, and L are natural numbers.
[0103] In an exemplary embodiment of the present invention, the update of the memory item includes:
[0104] During the training stage of the intrusion detection model, the sensor data when the acquisition system is running normally is used for training, and through the incremental update mechanism, the memory module can accurately remember the normal mode. However, the operating mode of the system may change, which causes the data distribution to change and leads to a decline in the detection performance of the model. For small and medium-sized industrial control systems with flexible positioning, this change is more frequent, which requires the intrusion detection model to adapt to the change of the system operating mode.
[0105] However, the method of offline training the model and then redeploying it will cause unnecessary resource waste. Therefore, this section proposes to update the memory module through a gating update mechanism during the online learning stage, so that the model can quickly adapt to the new normal mode.
[0106] During the intrusion detection stage of the model, the input data not only includes normal data but also intrusion data. Through the incremental update method, abnormal patterns will be introduced into the memory module. A gating mechanism is proposed to judge whether the current input is used to update the memory.
[0107] S301: Calculate the distance score between the current input data and the memory items, and update the memory module when the distance score is less than the update threshold;
[0108] S302: The calculation of the distance score between the current input data and the memory items includes:
[0109] When calculating the distance score, calculate the distance between the new data and the memory items, and then use the Exponential Weighted Average (EWA) function to calculate the weighted distance score. The EWA-weighted distance score pays more attention to similar data, enabling the model to adapt to this change faster and respond more quickly to newly emerging normal patterns. The calculation of the EWA-weighted distance score of the potential features of the current data is as follows:
[0110]
[0111] In the formula, D(z, m i ) is the distance between the potential feature vector and the memory item, γ is the exponential weighting factor, k is a natural number, and Score(z) is the distance score;
[0112] S303: Obtain the k items with non-zero first weight vectors after querying the potential feature vector and the memory module, and sort them in ascending order.
[0113] D(z, m1) ≤... ≤ D(z, m k )
[0114] In the formula, D(z, m1) is the distance between the potential feature vector and memory item 1, and D(z, m k ) is the distance between the potential feature vector and memory item k,
[0115] Compare the distance score with the update threshold. If the score is less than the threshold, update the new feature to the memory module, which ensures that abnormal patterns are not updated to the memory. The updated memory item is selected with the smallest distance score and placed in front of D(z, m1), and z is used to replace this item.
[0116] An exemplary embodiment of the present invention further includes establishing a loss function to balance the training effect of the ME-VAE model.
[0117] This solution uses VAE to reconstruct data, aiming to restore normal samples as much as possible. Based on VAE to reconstruct the input samples, the loss function generally consists of a reconstruction loss and KL divergence. Specifically, the reconstruction loss measures the ability of the model to reconstruct the input samples, usually measured by the distance between the reconstructed samples and the original samples. When reconstructing industrial control time series data or other continuous value data, the commonly used loss function is the mean squared error MSE. This loss function can measure the deviation between the predicted value of the model and the actual value.
[0118]
[0119] MSE is the mean of the squares of the differences between the predicted values and the true values. Optimizing the MSE loss function can make the model pay more attention to large errors because the error squares amplify the larger error values.
[0120] As another metric, the KL divergence loss is used to limit the difference between the latent variable distribution and the standard normal distribution to promote the smoothness and continuity of the latent space and prevent overfitting.
[0121]
[0122] Therefore, the loss function based on VAE reconstruction is as follows:
[0123]
[0124] By introducing a memory module to store the features of normal samples, then performing similarity matching on the latent vectors output by the encoder in the VAE network, and finally generating a vector composed of multiple positive sample features combined according to the similarity weights. However, if the number of combined vectors is too large, abnormal samples may still be reconstructed. Therefore, the weight matrix composed of similarities should be as sparse as possible to reduce the number of normal features used in the memory module.
[0125] In response to the above problems, different solutions have been proposed in the prior art. During the training process, features with small similarities are pruned, and the entropy of the weight matrix is minimized in the loss function to sparsify the weight matrix. Or by adding three memory units to store more normal behavior features. However, in the industrial control intrusion detection scenario, both device resources and computing capabilities are limited. The size of the module for storing normal features should be as small as possible to avoid storing too many similar normal features and wasting storage resources. Therefore, it is necessary to balance the relationship between the size of the storage unit and the feature diversity. To solve this problem, this solution proposes a differential loss function for features, which minimizes the similarity between features through the loss function, so that there are as many different normal behavior features as possible in the memory module, as follows.
[0126]
[0127] During the entire training process of the variational autoencoder, it is necessary to gradually make the distribution of the latent variables tend to a normal distribution. Therefore, in order to make the latent variables composed of normal features in the memory module and the encoded latent variables have similar distributions, this section proposes a feature synthesis loss to avoid excessive distribution differences between the latent variables and the synthesized features, and makes the two vectors as similar as possible by minimizing the Euclidean distance between them, as follows:
[0128]
[0129] Therefore, the proposed ME-VAE model in this solution balances the training effect by combining three loss functions and assigning weights to them respectively.
[0130]
[0131] In the formula, is the mean square error, m is the number of data, y i is the predicted value, y i ′ is the actual value, is the divergence loss, is the loss function of VAE reconstruction, is the difference loss function, is the Euclidean distance, is the total number of all features, μ is the mean, σ is the variance, m j is the j-th latent feature representation vector, η is the hyperparameter for weighing the loss function of.
[0132] An exemplary embodiment of the present invention further includes adjusting the reconstruction error by the difference between the input and output features of the memory module:
[0133] The input sample is reconstructed through the decoding layer of ME-VAE. In the anomaly recognition stage, the main task is to compare the anomaly score with the threshold to determine whether the sample is an intrusion. Existing reconstruction-based models mainly use the reconstruction error as the anomaly determination criterion. However, this method has certain defects. When the intrusion sample is similar to the normal sample, the model may still reconstruct the intrusion sample, and it is difficult for the model to effectively distinguish between normal and intrusion data, thus affecting the performance of the entire model. Therefore, to solve this problem, this section proposes an anomaly score method that adjusts the reconstruction error through feature differences.
[0134] Specifically, using only the reconstruction error as the basis for anomaly recognition may lead to misjudgment problems and increase the difficulty of intrusion recognition. Therefore, in this section, by utilizing the ability of the memory module to capture normal features, the difference between the input and output features of the memory module is used to adjust the reconstruction error.
[0135]
[0136] Wherein, S(t) is the adjusted reconstruction error, and ε ′ is the correction parameter, x t is the input data feature, is the data feature after reconstruction, z t is the initial feature, is the feature generated by recombining the memory module. The reconstruction error and the feature error in the latent space can be combined. The sensitivity of the model to small changes can be enhanced by the method of adjusting the feature difference. When the input sample is abnormal, the feature generated by recombining the memory module t has a large difference from the original z. Even if the decoder can reconstruct data similar to the input, considering the error in the feature space, the final anomaly score can be amplified. Finally, by comparing the overall anomaly score in the sliding window with a pre-set threshold, when the anomaly score exceeds the threshold, it indicates that an intrusion behavior is detected. The setting of the threshold is determined by the maximum anomaly score calculated in the training phase.
[0137] As Figure 5 shown, a comparison chart of the effects of the model ME-VAE of the present invention and other models is provided.
[0138] An exemplary embodiment of the present invention further includes an interpretable industrial control intrusion detection method based on a graph attention network. Specifically,
[0139] A prediction model is established based on a graph attention network and a gated recurrent unit. The prediction model predicts the data of the industrial control system and filters and denoises the predicted data.
[0140] Similarly, data preprocessing is required before this. Since there are large differences between the data value ranges of different feature attributes in the intrusion detection dataset, in order to construct data convenient for subsequent processing by the intrusion detection model, the original data is normalized. At the same time, in order to construct a dataset suitable for spatio-temporal sequence prediction, data sampling is performed using a sliding window, and the training dataset and the test dataset are processed into multivariate time series data.
[0141] Regarding the data prediction based on a graph attention network and a gated recurrent unit. In an industrial control environment, there are often correlations between devices. Graph neural networks are a commonly used method to obtain the correlations between devices. However, existing research based on graph neural networks still has room for improvement. Therefore, in this stage, the limitations of the graph attention network are first described; then how to improve it is described in detail; at the same time, a prediction network combining the improved GAT and a gated recurrent unit is proposed to achieve spatio-temporal feature extraction of sensor data and prediction of sensor values.
[0142] Compound anomaly determination based on threshold and entropy. Since there is noise in the industrial control environment, the noise will also cause anomalies in the device data, resulting in a high false alarm rate of the algorithm. Therefore, in this stage, the necessity of solving the noise problem is first clarified; secondly, a compound discriminator based on threshold and approximate entropy is proposed. Its discrimination strategy is to first screen out the sequences that may be anomalies through the threshold, and then calculate the approximate entropy of these sequences to filter out the anomalies caused by noise; finally, the anomalies are explained by aggregating the neighbor anomaly scores.
[0143] As Figure 6 shown, in order to monitor the operating state of the industrial control system, a large number of sensor devices often need to be deployed. These devices will generate a large amount of high-dimensional and complex data when working, and there is an interaction between each device, making the data potentially correlated. This requires the intrusion detection model to effectively extract the spatio-temporal features of the sensor data to accurately model the operating state of such industrial control systems. However, most of the current intrusion detection models only focus on one type of feature and lack the ability to extract both features simultaneously. Therefore, how to comprehensively extract the time dependence of high-dimensional complex sensor data and the hidden correlation between sensors is an important challenge at present. To further improve the detection efficiency of the model, this section proposes an intrusion detection model of SC-GRN, which extracts the spatio-temporal features of sensor data by combining the improved GAT network and GRU. Specifically, a continuous graph structure learning method based on similarity constraint is proposed to improve the GAT network, learn more continuous and stable spatial features between sensor data, and thus reduce the introduction of noise nodes. The update of the graph structure in each round is carried out under the constraint of the previous round of graph structure, maintaining the continuity of graph structure learning. And by setting a threshold to screen out neighbor nodes with higher correlation, the introduction of redundant neighbor nodes is reduced, and at the same time, the structure of the graph is streamlined, reducing the computational complexity.
[0144] As Figure 7 shown, M represents the global correlation matrix, which is used to retain the graph structure information learned in the previous round. M is generated by random initialization. S represents the similarity relationship between nodes in the current round.
[0145] The vector is randomly initialized and continuously updated during the model training process. The similarity constraint is used to minimize the gap between the two matrices. During the training process, M is dynamically adjusted by minimizing this loss function, so that the finally generated graph structure reaches the optimal.
[0146]
[0147] Always, is the similarity constraint.
[0148] Next, neighbor nodes are selected through a threshold strategy to generate an adjacency matrix A representing the graph structure. Specifically, a threshold is set, and there will be a connection between the adjacency matrices only when the similarity between two nodes is greater than or equal to this threshold.
[0149] By combining local similarity constraints and threshold-based node neighbor selection, noise nodes introduced due to the discontinuity of graph structure learning can be reduced. This method can more accurately reflect the interactions and dependencies between sensors in time series data, thereby improving the performance of intrusion detection.
[0150] As Figure 8 shown, after obtaining the graph structure of the sensor nodes, the improved GAT is used to capture the spatial dependencies between sensors, and then the temporal features of the sensor data are obtained through the GRU cell. The final prediction layer consists of two stacked fully connected layers and makes predictions based on the last layer output by the GRU.
[0151] The prediction model in this embodiment adopts the SC-GRN model. The specific prediction model includes:
[0152] The input of the GAT layer is the sensor embedding vector and the adjacency matrix A representing the graph structure. The weights between nodes are learned through the self-attention mechanism to capture the spatial dependencies between sensors. The current sensor node i obtains the updated representation by aggregating the information of adjacent nodes
[0153]
[0154]
[0155] Through the GAT, the relationships between sensors can be captured and these relationships are transformed into the feature representations of nodes. Finally, n sensor nodes are obtained, and then the temporal features of the sensor data need to be extracted. The GRU is used for temporal feature extraction, which can effectively capture the temporal information of sequence data and learn the dependencies between sequence data.
[0156]
[0157] In the formula, is the feature of node i after aggregating neighbor nodes, α ij is the attention coefficient of node j to node i in the GAT network, W is the learnable weight matrix, Ψ(i) is the neighbor of node i obtained from the adjacency matrix, is the sliding window input for node i, is the sliding window input for node j, is the concatenation of the sensor embedding and its corresponding sliding window input, v i′ is the feature embedding vector of the i-th sensor, e ij is the attention coefficient after activation using the non-linear activation function LeakyReLU, e ik is the unnormalized value of the attention coefficient between node i and its neighbor node k, a is a learnable parameter vector, z ′(t) is the feature vector of n sensors containing spatio-temporal features in the sliding window, x (t) is the industrial control system data, z (t) is the output of the GAT network.
[0158] After the feature vector is learned by the GRU unit, it can help the intrusion detection model identify and extract the time-dependent relationships in the data sequence, improving the intrusion detection accuracy. The prediction layer consists of a fully connected network, and based on the historical input vector representation, the predicted value at the current time t is calculated
[0159] An exemplary embodiment of the present invention for filtering and denoising the predicted data includes:
[0160] Preliminary discrimination of the threshold. The intrusion detection method based on prediction is trained on the normal state of the system, uses the model to predict the next value of the sensor, and monitors the prediction error to detect anomalies. When the system is in the normal mode, the prediction error is usually low, while when the system is in the abnormal state, the prediction error may increase significantly. Therefore, the deviation between the predicted value and the true value can be used as a measure of the anomaly.
[0161] First, a threshold is set to determine whether it is higher than the threshold. When the error exceeds the threshold, it indicates that an abnormal situation possibly caused by an intrusion is detected. This discriminator can quickly and effectively identify those anomalies that deviate significantly from the normal range and is suitable for scenarios with obvious anomalies. Whether there is an abnormal situation is judged based on the comparison between the anomaly score and the preset threshold.
[0162] By setting a parameter as the threshold in the intrusion detection algorithm, if the anomaly score exceeds the threshold, then the time t is put into the filtering window, and then further judgment is made according to the approximate entropy of the sequence in the filtering window.
[0163] S401: Obtain the error sequence, perform phase space reconstruction on the error sequence with a given embedding dimension for the error sequence within the filtering window, and divide the error sequence into several subsequences;
[0164] S402: Calculate the similarity between any two reconstructed vectors. The distance between two reconstructed vectors is defined as the maximum absolute value of the differences between the corresponding elements of the two sequences;
[0165] Specifically, find all similar subsequences. First, set a similarity tolerance threshold and find all sequences with a distance below this threshold. If the distance between the two subsequences with the farthest distance is less than the similarity tolerance threshold, then these two sequences are considered similar under the similarity tolerance threshold.
[0166] Select all cases that meet this condition and record the quantity. Denote the ratio of the approximate quantity to the total quantity as the approximate ratio.
[0167]
[0168] In the formula, is the approximate ratio, is the distance between the two subsequences with the farthest distance, r is the similarity tolerance threshold, and k is the number of subsequences.
[0169] S403: Calculate the logarithmic probability of all sequences and obtain the average value. Calculate the approximate entropy of the error sequences within the filtering window through the average value;
[0170] S404: Compare the approximate entropy with the entropy threshold, judge and output the abnormal data.
[0171] As Figure 9 shown, an exemplary embodiment of the present invention further includes an anomaly explanation based on the aggregated neighbor anomaly score.
[0172] In order to enable engineers to quickly locate and maintain specific abnormal devices when the model detects anomalies, this section locates the devices that may be invaded through the graph structure information learned by GAT.
[0173] Since in the anomaly discrimination stage, the final system anomaly score is provided by the sensor with the highest anomaly score in the current system. However, in a complex system, there are often mutual influences between sensors. The anomaly of a certain sensor may be caused by other sensors. In this case, although the anomaly scores of some sensors are very high, they may not be the root cause of the anomaly, but just a reflection of the anomaly symptoms. In order to further identify the root cause, this section proposes an anomaly explanation method based on the aggregated neighbor anomaly score, and locates the abnormal device through the graph structure learned by GAT and the correlation between neighbor nodes. The specific process is as follows:
[0174] Calculate the anomaly score. First, in the anomaly discrimination stage, calculate the anomaly score of a single sensor and find the sensor with the highest anomaly score.
[0175] Preliminary diagnosis. Conduct a preliminary diagnosis on the sensor i with the highest current anomaly score. If this sensor is not the root cause of the anomaly, further analysis is required.
[0176] Calculate the aggregated neighbor anomaly score. Calculate the aggregated neighbor anomaly scores for all sensors. Specifically, using the adjacency matrix A learned by GAT, find the neighbor sensors of a certain sensor, and further weight and sum the anomaly scores of the direct neighbors of sensor i according to the degree of correlation.
[0177]
[0178] In the formula, R i (t) is the anomaly score after aggregating the one-hop neighbors of sensor i, and err j (t) is the anomaly score of a single sensor.
[0179] Anomaly score sorting. Sort all sensors according to the anomaly scores to obtain an anomaly sorting table. The sensors in the sorting table are arranged according to the magnitude of the anomaly scores, that is, the sensors with higher anomaly scores are at the front. This is because if the anomaly scores of the neighbor nodes of a sensor are higher, then the probability of this sensor having an anomaly is greater.
[0180] Anomaly localization. Provide the sorting table to the system operator for anomaly localization. The sensors earlier in the list are more likely to be the root causes of the anomalies.
[0181] In addition, in each embodiment of the present invention, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0182] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. And the aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, Read—Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.
[0183] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. An industrial control intrusion detection method based on a graph attention network and a variational autoencoder, characterized in that Including: Obtain data of the industrial control system and preprocess the data. The preprocessing includes data cleaning and normalization of the cleaned data; Divide the data using a sliding window; Build an ME-VAE model. The ME-VAE model includes an encoder, a memory module, and a decoder. Enhance the latent features through the memory module, amplify the reconstruction error of intrusion samples, and incrementally update the memory in the ME-VAE model; Determine whether to use the current data to update the memory. If so, update several memory items. If not, do nothing; Obtain the difference between the latent features and the features enhanced by the memory module, weight the reconstruction error of the data, and determine intrusion through the weighted reconstruction error, and output the determination result; The building of the ME-VAE model includes: Given a time input sequence, encode and sample it through an encoder composed of a GRU module and a fully connected layer, and output the latent representation of the intermediate layer; Use the latent representation as a query vector, calculate the similarity with each memory item in the memory module, and combine the query vector and the memory item into a new latent feature vector according to the similarity; Pass the latent feature vector to the decoder for reconstruction to restore the time input sequence; The enhancement of the latent features includes: Calculate the distance between the query vector and each unit in the memory module to obtain a weight vector; w i = Softmax(d cos (z, m i )) And correct the weight vector through vector sparsification; where w i is the first weight vector, d cos (z, m i ) is the cosine similarity between the latent feature vector and the memory item, z is the latent feature vector, and m i is a cell in the memory module, is the modified weight vector, λ is the sparsity threshold, ε is a positive scalar, and max(., 0) is the ReLU activation function; It also includes adjusting the reconstruction error by the difference between the input and output features of the memory module: Where S(t) is the adjusted reconstruction error, ε ′ is the correction parameter, x t is the input data feature, is the reconstructed data feature, z t is the initial feature, is the feature generated by recombining the memory modules.
2. The industrial control intrusion detection method based on graph attention network and variational autoencoder according to claim 1, wherein The incremental update of the memory includes: where, v i ′ is the second weight vector, v i is the matching probability, is the new latent feature vector, v j is the j-th probability among all possible values after softmax calculation, is the cosine similarity between the new latent feature vector and the memory item, and i, j, and L are natural numbers.
3. The industrial control intrusion detection method based on graph attention network and variational autoencoder according to claim 2, wherein The update of the memory item includes: Calculate the distance score between the current input data and the memory item. When the distance score is less than the update threshold, update the memory module; The calculation of the distance score between the current input data and the memory item includes: where D(z, m i ) is the distance between the latent feature vector and the memory item, γ is the exponential weighting factor, k and i are natural numbers, and Score(z) is the distance score; Obtain the k items with non-zero first weight vectors after querying the memory module by the latent feature vector, and arrange them in ascending order.
4. The industrial control intrusion detection method based on graph attention network and variational autoencoder according to claim 3, characterized in that It also includes building a loss function to balance the training effect of the ME-VAE model. The loss function includes: Wherein, is the mean square error, m is the number of data, y i is the predicted value, y i ′ is the actual value, is the divergence loss, is the loss function of VAE reconstruction, is the difference loss function, is the Euclidean distance, is the total number of all features, μ is the mean, σ is the variance, m j is the j-th latent feature representation vector, η is the hyperparameter for weighing the loss function of.
5. The industrial control intrusion detection method based on graph attention network and variational autoencoder according to claim 1, characterized in that, It also includes: Build a prediction model based on a graph attention network and a gated recurrent unit. The prediction model predicts the data of the industrial control system and filters and denoises the predicted data.
6. The industrial control intrusion detection method based on graph attention network and variational autoencoder according to claim 5, wherein The prediction model includes: In the formula, is the feature of node i after aggregating neighboring nodes, α ij is the attention coefficient of node j to node i in the GAT network, W is a learnable weight matrix, Ψ(i) is the neighbor of node i obtained from the adjacency matrix, is the sliding window input with node i, is the sliding window input with node j, is the concatenation of the sensor embedding and its corresponding sliding window input, v u ′ is the feature embedding vector of node i, e ij is the attention coefficient after activation using the non-linear activation function LeakyReLU, e ik is the unnormalized value of the attention coefficient between node i and its neighbor node k, a is a learnable parameter vector, z ′(t) is the feature vector of n sensors containing spatio-temporal features in the sliding window, x (t) is the industrial control system data, z (t) is the output of the GAT network.
7. The industrial control intrusion detection method based on graph attention network and variational autoencoder according to claim 6, wherein The filtering and denoising of the predicted data includes: Obtain an error sequence, perform phase space reconstruction on the error sequence with a given embedding dimension within the filtering window for the error sequence, and divide the error sequence into several subsequences; Calculate the similarity between any two reconstructed vectors. The distance between two reconstructed vectors is defined as the maximum absolute value of the difference between the corresponding elements of the two sequences; Calculate the logarithmic probability of all sequences and obtain the average value, and calculate the approximate entropy of the error sequence within the filtering window through the average value; Compare the approximate entropy with the entropy threshold, judge the abnormal data and output it.
Citation Information
Patent Citations
Deep convolution calculation model supporting increment update
CN108009635A
Attack and defense method and system of variational auto-encoder based on GRU
CN118427704A