Unsupervised Anomaly Detection Method and System for Industrial Control Networks Based on Spatiotemporal Codec
By adopting an unsupervised deep learning method based on spatiotemporal codecs in industrial control networks, combining spatiotemporal feature analysis and self-supervised learning, the problems of low recognition rate and weak generalization of traditional methods are solved, and more efficient and robust anomaly detection effects are achieved.
Patent Information
- Application Number
- CN202510066964.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Traditional industrial control network anomaly detection methods have low recognition rate and weak generalization, making it difficult to effectively deal with complex network attacks and dynamic threats.
Unsupervised deep learning method based on spatiotemporal codec is adopted, and the spatiotemporal self-supervised learning branch works together through spatiotemporal feature analysis branches and spatiotemporal self-supervised learning branches to generate potential representations that understand the global spatiotemporal context, thereby improving the robustness and generalization performance of anomaly detection.
It improves the accuracy and robustness of industrial control network anomaly detection, enhances the generalization ability of the model, and can more effectively identify complex network attacks and dynamic threats.
Smart Images

Figure CN119520165B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network security, and particularly relates to an unsupervised anomaly detection method and system for industrial control networks based on spatio-temporal codecs. Background Art
[0002] Traditional industrial control network anomaly detection methods mainly rely on parsing and identifying abnormal communication behaviors of general Internet protocols, and matching them according to preset rules and eigenvalue to achieve security filtering. However, these methods have defects such as low recognition rate and weak generalization ability, and it is difficult to effectively cope with complex network attacks and dynamic threats. However, in industrial control systems, the system state can be reflected by the readings of internal sensors, and any unauthorized intrusion may cause abnormal changes in these sensor values. Therefore, by analyzing the time-series data and spatial topology of sensors, abnormal behaviors in the production process can be discovered, which is crucial for maintaining the security of industrial control systems.
[0003] In view of the complexity and real-time requirements of industrial control networks, anomaly detection methods based on spatio-temporal feature analysis are particularly important. Spatio-temporal feature analysis can capture the temporal characteristics and spatial correlations in industrial control networks, thereby more accurately identifying abnormal behaviors. For example, by analyzing the periodic patterns and time-orderliness in sensor data, a normal behavior model can be constructed, and deviations from the expected behavior model can be detected based on this model, thereby achieving anomaly detection. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention proposes an unsupervised anomaly detection method and system for industrial control networks based on spatio-temporal codecs, and uses them for anomaly detection of industrial control networks. The present invention adopts two branches, one branch is used for spatio-temporal feature analysis tasks, and the other branch is used for spatio-temporal self-supervised learning tasks. These two branches share the same spatio-temporal encoder, whose function is to generate latent representations that understand the global spatio-temporal context, thereby improving its generalization performance and the robustness of anomaly detection. The original data of engineering control network sensors usually includes network load, temperature, humidity, current, voltage, flow rate, and pressure, all of which can be used as data for anomaly detection. By training and optimizing different time steps and different training parameters, a model with the highest anomaly detection accuracy can be obtained, and then an industrial control network anomaly detection system can be formed, providing certain suggestions for the field of network security.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] An unsupervised anomaly detection method for industrial control networks based on spatio-temporal codecs, comprising the following steps:
[0007] Step 1: Sensor data processing and sample construction; sample and process the historical industrial control sensor raw data and time series to obtain a sample sequence;
[0008] Step 2: Construct and train an unsupervised deep learning model based on a spatio-temporal codec;
[0009] The unsupervised deep learning model based on a spatio-temporal codec includes a spatio-temporal feature analysis module, and the spatio-temporal feature analysis module includes a spatio-temporal feature analysis branch and a spatio-temporal self-supervised learning branch, which are respectively used for spatio-temporal feature analysis and spatio-temporal self-supervised learning;
[0010] Step 3: Spatio-temporal feature analysis of sensor data: Input the data output in Step 1 into the trained unsupervised deep learning model based on a spatio-temporal codec. The spatio-temporal feature analysis branch and the spatio-temporal self-supervised learning branch work together to output the sensor data in a future predetermined time period as the predicted value; Step 4: Abnormal alarm: Determine the confidence space by combining the predicted value and confidence level output in Step 3, and use it as the normal data range for industrial control network detection. When the current actual sensor data exceeds the normal range, trigger a network abnormal alarm.
[0011] Furthermore, both the spatio-temporal feature analysis branch and the spatio-temporal self-supervised learning branch adopt a spatio-temporal encoder-decoder structure. Specifically, the spatio-temporal feature analysis branch adopts a spatio-temporal codec Ⅰ, including a spatio-temporal encoder Ⅰ, a frequency enhancement attention module, and a spatio-temporal prediction decoder; the spatio-temporal self-supervised learning branch adopts a spatio-temporal codec Ⅱ, including a spatio-temporal encoder Ⅱ, a self-supervised learning module, and a spatio-temporal reconstruction decoder; among them, the spatio-temporal codec Ⅰ and the spatio-temporal encoder Ⅱ have the same structure;
[0012] Within the spatio-temporal feature analysis branch, the long-term spatio-temporal dynamic information of the sensor data is encoded and integrated by the spatio-temporal encoder Ⅰ to obtain a latent feature space matrix , and then the spatio-temporal dynamics in the latent feature space matrix are captured by the frequency enhancement attention module and the spatio-temporal prediction decoder ;
[0013] Within the spatio-temporal self-supervised learning branch, a masking algorithm for spatio-temporal data is introduced, and patches are randomly masked by spatio-temporal agnostic sampling for spatio-temporal self-supervised learning of the data. The masked data is complemented and reconstructed by the spatio-temporal reconstruction decoder, and the output , and the reconstructed data is aligned and compared with the calculated by the spatio-temporal encoder Ⅰ in the spatio-temporal feature analysis branch in the latent space, and the result continuously guides the improvement of the spatio-temporal encoder Ⅰ in the spatio-temporal feature analysis branch.
[0014] Further, in step 3, before the sensor data is input into the spatio-temporal feature analysis branch and the spatio-temporal self-supervised learning branch, the following processing is first performed: The sensor data first undergoes preliminary feature extraction and non-linear transformation through a feed-forward neural network to obtain the sensor data signal matrix E. Then, the spatio-temporal feature embedding module integrates spatial and temporal information to generate an embedded representation that fuses spatio-temporal information, and outputs the spatio-temporal embedding matrix E' of the sensor data signal matrix sequence.
[0015] Further, in step 3, the data processing process within the spatio-temporal feature analysis branch is as follows:
[0016] (1) First, the original sensor data matrix of the spatio-temporal feature analysis branch is projected into through a linear transformation, where C is the number of features;
[0017] (2) Then, is input into the spatio-temporal encoder I, which is connected with L stacked spatio-temporal attention modules. Together with E´, they serve as the long-term spatio-temporal dynamic information of the sensor data. After encoding and integration, the latent feature space matrix is obtained;
[0018] (3) Then it enters the frequency enhancement attention module, and the formula is:
[0019] ;
[0020] Among them, is the output of the frequency enhancement attention module, is the Fourier transform formula of the frequency enhancement attention module, is the spatio-temporal embedding matrix of the input sensor data signal matrix sequence, is the spatio-temporal embedding matrix of the output sensor data signal matrix sequence;
[0021] (4) Next, it enters the spatio-temporal prediction decoder, whose input is the output of the frequency enhancement attention module, that is, ; The spatio-temporal prediction decoder stacks spatio-temporal attention modules to capture the spatio-temporal dynamics in the latent space, and finally projects through a linear projection to the prediction target, that is, the future sensor value .
[0022] Further, the spatio-temporal attention module includes a spatial attention module and a temporal attention module. The implementation processes of the spatial attention module and the temporal attention module are the same. The implementation process of the spatial attention module is as follows:
[0023] a. Receive the input sensor data signal tensor , spatio-temporal embedding STE, and mask tensor mask;
[0024] b. Concatenate the input sensor data signal tensor and the spatio-temporal embedding STE along the feature dimension to obtain a new tensor X´´;
[0025] c. Initialize the weight matrices of the queries and keys of the multi-head attention mechanism with learnable parameters;
[0026] d. Perform multi-head attention calculation on the initialized weight matrices G of the queries and keys and the tensor X´´ to obtain the attention-weighted result h;
[0027] e. Perform multi-head attention calculation on the attention-weighted result h and the input tensor to obtain the final result of the spatial attention mechanism.
[0028] Furthermore, perform the following operations on the spatial and temporal attention: Represent the output of the th spatio-temporal attention module as ; Connect it with the spatio-temporal embedding matrix of the input sensor data matrix sequence to obtain , and input into the th spatio-temporal attention module; Use to represent the input regarding node on all time slices, and use to represent the input of all nodes on time slice ;
[0029] In the time dimension, define the time vector of the sensor data signal as the time reference point in the time attention of the th spatio-temporal attention module. The formula for the time attention module is as follows:
[0030] ;
[0031] ;
[0032] where, is the multi-head self-attention formula. The time attention module first transforms into through the attention input . The updated time reference point contains the global sensor data information of the original long-term input , and finally generates the time-coded representation of node ; The time attention module processes the input of each node in parallel. Use Represents the time-coded representation of all nodes;
[0033] Similarly, in the spatial dimension, spatial vector is introduced as the spatial reference point, and then the spatial attention module processes the input on each time slice in parallel. The formula is as follows:
[0034] ;
[0035] ;
[0036] where, is the spatial-coded representation of all nodes at time t, represents the spatial-coded representation on all time slices; finally, the output of the spatial attention module is the sum.
[0037] Furthermore, the frequency enhancement attention module includes an encoder, a decoder, and a Fourier enhancement structure. Both the encoder and the decoder adopt a multi-layer structure. The Fourier enhancement structure is used between the encoder and the decoder. The Fourier enhancement structure uses Fourier transform and inverse Fourier transform, and uses the Fourier transform frequency enhancement attention mechanism to process the data of the encoder. The Fourier enhancement structure converts the time-domain data into frequency-domain data by applying Fourier transform, and executes the Fourier transform frequency enhancement attention mechanism in the frequency domain, and applies inverse Fourier transform to convert the frequency-domain data back into time-domain data.
[0038] Furthermore, the encoder in the spatio-temporal self-supervised learning branch is the same as the encoder used in the spatio-temporal feature analysis branch. In the spatio-temporal encoder of the spatio-temporal self-supervised learning branch, by setting the mask value in the softmax function input to -∞, the masked sensor data does not participate in the calculation. The output of the spatio-temporal encoder in the spatio-temporal self-supervised learning branch is denoted as ;
[0039] The input of the spatio-temporal reconstruction decoder is denoted as , which consists of two parts. One part is the output of the spatio-temporal encoder in the spatio-temporal self-supervised learning branch, and the other part is the mask token vector, which occupies all the masked positions and reveals the absence of the true value;
[0040] The reconstruction decoder stacks spatio-temporal attention modules to recover the sensor data in the latent space, and finally outputs , which matches the calculated by the spatio-temporal encoder in the spatio-temporal feature analysis branch.
[0041] Further, in the spatio-temporal self-supervised learning branch, based on the masking algorithm of spatio-temporal feature data, self-supervised learning is performed on sensor data; first, the sensor data signal matrix is divided into small blocks, that is, a continuous data signal segment with a length of ; then, randomly mask the patches with spatio-temporal agnostic sampling;
[0042] The steps of self-supervised learning of sensor data are as follows:
[0043] (1) Input the original sensor data matrix , where P is the length of the input sensor data matrix sequence, N is the number of sensors, and C is the number of features;
[0044] (2) Generate a mask tensor , and set the elements of the tensor to 1;
[0045] (3) Divide the sensor data signal matrix into patches, that is , ;
[0046] (4) Calculate the number of masked patches ;
[0047] (5) Randomly draw from patches, and record the corresponding patch indices as ;
[0048] (6) Loop through ;
[0049] (7) Set all elements of the -th patch in the tensor to 0;
[0050] (8) Output the mask tensor .
[0051] The present invention also provides an unsupervised anomaly detection system for industrial control networks based on a spatio-temporal codec, which is used to implement the unsupervised anomaly detection method for industrial control networks based on a spatio-temporal codec as described above. The system includes a data processing module, a sample construction module, an unsupervised deep learning model based on a spatio-temporal codec, and an anomaly warning module;
[0052] The data processing module performs data cleaning and standardization operations on the original sensor data;
[0053] The sample construction module is used to collect historical sensor data to obtain a sensor data sequence, obtain multiple sensor data subsequences through a sliding window form, and use the normalized sensor data subsequences as sample sequences;
[0054] The unsupervised deep learning model based on the spatio-temporal codec includes a spatio-temporal feature analysis module, which includes a spatio-temporal feature analysis branch and a self-supervised learning branch. For each sensor data sequence to be measured, sensor data prediction is performed, and the sensor data for a future predetermined period is output as the prediction result.
[0055] The anomaly warning module is used to determine the confidence space as the normal range of the sensor value, and trigger an industrial control network anomaly warning when the actual sensor value at the current moment is not within the normal range.
[0056] Compared with the prior art, the advantages of the present invention are as follows:
[0057] (1) The present invention combines the self-supervised learning technology and the spatio-temporal attention mechanism, and proposes a method applicable to anomaly detection in industrial control networks, an anomaly detection method based on spatio-temporal feature analysis. By combining the advantages of both, the data utilization efficiency is improved, and the anomaly detection model has good generalization ability.
[0058] (2) Using a multi-task learning framework, multiple tasks are learned simultaneously; through training and tuning of different time steps and different training parameters, the model with the highest anomaly detection accuracy can be obtained, and an anomaly detection system is formed, improving the ability of feature extraction for long time slices, enhancing the ability to capture long-term dynamic information and global correlation information, improving the accuracy of anomaly detection in the field of industrial control networks, and providing certain suggestions for the field of network security. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0060] Figure 1 It is a schematic flow chart of the method of the present invention;
[0061] Figure 2 It is a schematic structural diagram of the anomaly detection module of the present invention;
[0062] Figure 3 It is a schematic structural diagram of the Fourier enhancement;
[0063] Figure 4 It is a schematic structural diagram of the spatio-temporal feature analysis branch;
[0064] Figure 5 It is a schematic structural diagram of the unsupervised anomaly detection system for industrial control networks based on the spatio-temporal codec of the present invention;
[0065] Figure 6 This is the prediction result graph obtained in the embodiment of the present invention;
[0066] Figure 7 This is the comparison graph of the effect differences between the embodiment of the present invention and other methods on the Abilene dataset;
[0067] Figure 8 This is the comparison graph of the prediction result and the true value obtained in the embodiment of the present invention. Detailed implementation manners
[0068] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments.
[0069] Based on self-supervised learning and frequency-enhanced attention mechanism, the present invention integrates temporal and spatial features, proposes a new anomaly detection method applicable to industrial control networks, an anomaly detection method based on spatio-temporal feature analysis, and forms a system. The process of the present invention is as Figure 1 shown.
[0070] As Figure 2 and Figure 5 shown, the present invention designs an anomaly detection module, including an unsupervised deep learning model based on a spatio-temporal codec and an anomaly warning module. The model adopts a multi-task learning framework, and the overall architecture of the model has two branches. One branch is used for the spatio-temporal feature analysis task of sensor data, and the other branch is used for the spatio-temporal self-supervised learning task of sensor data signals. These two branches share the same spatio-temporal encoder, whose function is to generate a latent representation of the input that understands the global spatio-temporal context, thereby improving its generalization performance and the robustness of prediction.
[0071] The design process of the present invention is: sensor data processing and sample construction; constructing an unsupervised deep learning model based on a spatio-temporal codec (including building an embedding module that fuses spatio-temporal features; constructing a spatio-temporal attention module; constructing a frequency-enhanced attention module; building a spatio-temporal feature analysis branch; building a spatio-temporal self-supervised learning branch; defining a loss function and evaluation metrics); constructing an anomaly warning module; forming an anomaly detection module, and further forming an anomaly detection system; method effect evaluation. Each part will be introduced in detail below.
[0072] Combined with Figures 1-4 shown, the unsupervised anomaly detection method for industrial control networks based on a spatio-temporal codec in this embodiment includes the following steps: Step 1, sensor data processing and sample construction; Step 2, constructing an unsupervised deep learning model based on a spatio-temporal codec and training it; Step 3, spatio-temporal feature analysis of sensor data; Step 4, anomaly alarm. Each step will be introduced in detail below.
[0073] Step 1, Sensor data processing and sample construction: Sample and process the historical raw sensor data and time series to obtain a sample sequence.
[0074] This step is to perform an historical data extraction operation to collect the raw data generated by sensors in the industrial control network (system) within a specified time range.
[0075] It should be noted that the raw data of industrial control network sensors includes network load, temperature, humidity, current, voltage, flow rate, and pressure, and all these data can be used as data for anomaly detection. In this embodiment, the network load parameter is taken as an example for illustration, which will be described in detail later.
[0076] First, extract the sensor data related to the time series from the raw sensor data to form a sensor data signal matrix. This sensor data network is denoted as G, and the sensor data signal matrix represented at time slice t is , where the network G consists of a group of sensors, denoted as the node set V, and , is the feature vector of node v observed at time t, and C is the number of features.
[0077] Sample the historical sensor data and time series in the form of a sliding window to obtain multiple sensor data subsequences, and perform normalization processing on each sensor data subsequence to map the original sequence data to the [1, 0] interval. The specific method is: first calculate the maximum and minimum values of the sequence data, denoted as and , for each data in the sequence data, perform calculation, and use the normalized sensor data subsequence as the sample sequence.
[0078] Regarding the characteristics of the sensor data in the time series, mine more time-related features. Use time-of-day to model the behavior patterns and change rules at different times of the day, and capture the periodicity and trend of the data changing over time. Use day-of-week to model the behavior patterns and change rules on different days of the week, and capture the periodicity and seasonality of the data according to the day of the week.
[0079] Finally, perform standardization processing on the sensor data so that it is distributed within the same mean and standard deviation range to improve the effect of model training and anomaly detection.
[0080] Step 2, Construct and train an unsupervised deep learning model based on a spatio-temporal autoencoder;
[0081] The unsupervised deep learning model based on the spatio-temporal codec includes a spatio-temporal feature analysis module, which includes a spatio-temporal feature analysis branch and a spatio-temporal self-supervised learning branch, for spatio-temporal feature analysis and spatio-temporal self-supervised learning respectively.
[0082] Both the spatio-temporal feature analysis branch and the spatio-temporal self-supervised learning branch adopt a spatio-temporal encoder-decoder structure. Specifically, the spatio-temporal feature analysis branch adopts a spatio-temporal codec Ⅰ, which includes a spatio-temporal encoder Ⅰ, a frequency enhancement attention module, and a spatio-temporal prediction decoder; the spatio-temporal self-supervised learning branch adopts a spatio-temporal codec Ⅱ, which includes a spatio-temporal encoder Ⅱ, a self-supervised learning module, and a spatio-temporal reconstruction decoder; among them, the spatio-temporal codec Ⅰ and the spatio-temporal encoder Ⅱ have the same structure, that is, the spatio-temporal feature analysis branch and the spatio-temporal self-supervised learning branch share the spatio-temporal encoder.
[0083] Step 3: Spatio-temporal feature analysis of sensor data: Input the data output in Step 1 into the trained unsupervised deep learning model based on the spatio-temporal codec. The spatio-temporal feature analysis branch and the spatio-temporal self-supervised learning branch work together to output the sensor data in a future predetermined time period as the predicted value.
[0084] Within the spatio-temporal feature analysis branch, the long-term spatio-temporal dynamic information of the sensor data is embedded into (the long-term spatio-temporal dynamic information of the sensor data is encoded and integrated by the spatio-temporal encoder Ⅰ to obtain the latent feature space matrix ), and then the spatio-temporal dynamics in the latent space are captured through the frequency enhancement attention module and the spatio-temporal prediction decoder.
[0085] Within the spatio-temporal self-supervised learning branch, a masking algorithm for spatio-temporal data is introduced. Random patches are masked with spatio-temporal agnostic sampling to perform spatio-temporal self-supervised learning on the sensor data. The masked data is complemented and reconstructed through the spatio-temporal reconstruction decoder Ⅰ, and is output. The reconstructed data is aligned and compared with calculated by the spatio-temporal encoder Ⅰ in the spatio-temporal feature analysis branch in the latent space. The result continuously guides and improves the spatio-temporal encoder Ⅰ of the spatio-temporal feature analysis branch. The two branches work together to finally output the sensor data in a future predetermined time period as the predicted value.
[0086] It should be noted that in step 3, before the sensor data is input into the spatio-temporal feature analysis branch and the spatio-temporal self-supervised learning branch, the following processing is first performed: The sensor data first undergoes preliminary feature extraction and non-linear transformation through a feedforward neural network to obtain the sensor data signal matrix E. Then, the spatio-temporal feature embedding module integrates spatial and temporal information to generate an embedded representation that fuses spatio-temporal information, and outputs the spatio-temporal embedding matrix E' of the sensor data signal matrix sequence. Such processing can effectively provide high-quality input data for the subsequent spatio-temporal feature analysis branch and spatio-temporal self-supervised learning branch.
[0087] As an example, first, the sensor data is fed into a feedforward neural network, which consists of an input layer, multiple hidden layers, and an output layer. The sensor data is transmitted unidirectionally in the network, and non-linear transformation is performed on the output of the attention head to obtain the sensor data signal matrix sequence , where P is the length of the input sensor data matrix sequence, N is the number of sensors, and d is a parameter; then, a spatio-temporal embedding module that fuses spatio-temporal features is built, which combines spatial embedding (SE) and temporal embedding (TE) to generate an embedded representation that fuses spatio-temporal information, and outputs the spatio-temporal embedding matrix of the sensor data signal matrix sequence , where Q is the length of the output sensor data matrix sequence.
[0088] Note: The spatio-temporal embedding module operates on the input data, aiming to extract spatio-temporal features of the sensor data, before the internal encoders of the two branches. The position is inside the "branch" of Figure 1 , and before the position of "before inputting to the encoder" in Figure 2 , the embedded representation of spatio-temporal information is first performed.
[0089] The spatio-temporal feature analysis branch inputs the original data (processed sample data). In the spatio-temporal self-supervised learning branch, some of the original data (processed sample data) is randomly masked, and the input is the damaged data. Therefore, the inputs of the two branches are not the same data, and the spatio-temporal self-supervised learning branch has one more mask (masking) process than the feature analysis branch.
[0090] In step 2, the spatio-temporal feature analysis branch includes a spatio-temporal encoder, a frequency enhancement attention module, and a spatio-temporal prediction decoder, as Figure 4 shown.
[0091] S201. Construct a spatio-temporal attention module:
[0092] The spatio-temporal attention module includes a spatial attention module and a temporal attention module. The implementation processes of the spatial attention module and the temporal attention module are the same. The spatial and temporal attention mechanisms are used to perform weighted aggregation on the input sensor data in the spatial and temporal dimensions to capture the importance of the sensor data in space and time. The implementation process of the spatial attention module is as follows:
[0093] a. Receive the input sensor data signal tensor , spatio-temporal embedding STE, and mask tensor mask;
[0094] b. Concatenate the input sensor data signal tensor and the spatio-temporal embedding STE along the feature dimension to obtain a new tensor X´´;
[0095] c. Initialize the weight matrices of the query and key of the multi-head attention mechanism with learnable parameters I;
[0096] d. Perform multi-head attention calculation on the initialized query and key weight matrices G and the tensor X´´ to obtain the attention-weighted result h;
[0097] e. Perform multi-head attention calculation on the attention-weighted result h and the input tensor to obtain the final result of the spatial attention mechanism.
[0098] To better encode the generalized global pattern of the sensor data, the following operations are performed on the spatial and temporal attention: Represent the output of the th spatio-temporal attention module as ; Connect it with the spatio-temporal embedding matrix of the input sensor data matrix sequence to obtain , and input into the th spatio-temporal attention module; Use to represent the input of all time slices regarding node , and use to represent the input of all nodes on time slice .
[0099] In the temporal dimension, define the temporal vector of the sensor data signal as the temporal reference point in the temporal attention of the th spatio-temporal attention module. The calculation formula in the temporal attention module is as follows:
[0100] ;
[0101] ;
[0102] Among them, is the multi-head self-attention formula. The temporal attention module first passes the attention input and converts it to . The updated temporal reference point contains the global sensor data information of the original long-term input and finally generates the temporal encoding representation of the node . The temporal attention module processes the input of each node in parallel, and uses to represent the temporal encoding representations of all nodes.
[0103] Similarly, in the spatial dimension, a spatial vector is introduced as the spatial reference point. Then, the spatial attention module processes the input on each time slice in parallel, and the formula is as follows:
[0104] ;
[0105] ;
[0106] Among them, is the spatial encoding representation of all nodes at time t, represents the spatial encoding representations on all time slices; finally, the output of the spatial attention module is the sum of .
[0107] S202. Construct a frequency enhancement attention module:
[0108] The frequency enhancement attention module includes an encoder, a decoder, and a Fourier enhancement structure. Both the encoder and the decoder adopt a multi-layer structure. The Fourier enhancement structure is used in the middle of the encoder and the decoder. The Fourier enhancement structure uses Fourier transform and inverse Fourier transform, and uses the Fourier transform frequency enhancement attention mechanism to process the data of the encoder. The Fourier enhancement structure converts the time-domain data into frequency-domain data by applying Fourier transform, and executes the Fourier transform frequency enhancement attention mechanism in the frequency domain, and applies inverse Fourier transform to convert the frequency-domain data into time-domain data. The frequency enhancement attention module improves the performance of the model by combining the encoder, the decoder, and the Fourier enhancement structure, and enhances the model's ability to extract frequency-domain features by adaptively modeling the frequency information of different channels of the input sequence.
[0109] As an example, the multi-layer structure adopted by the encoder is: where k represents the number of layers of the encoder, , respectively represent the sensor data sequences of the k-th layer and the (k - 1)-th layer of the encoder, is the embedded historical sensor data sequence, is:
[0110] ;
[0111] ;
[0112] ;
[0113] wherein, respectively represent the seasonal component of the sensor data after the j-th decomposition block in the k-th layer, is the mixture of experts decomposition block; is the frequency enhancement attention module, implemented through the discrete Fourier transform (DFT) mechanism.
[0114] The decoder also adopts a multi-layer structure, , wherein represents the number of decoder layers, respectively represent the seasonal component and the trend component of the sensor data in the p-th layer of the decoder, respectively represent the seasonal component and the trend component of the sensor data in the (p - 1)-th layer of the decoder, is formalized as:
[0115] ;
[0116] ;
[0117] ;
[0118] ;
[0119] ;
[0120] wherein, respectively represent the seasonal component and the trend component of the sensor data after the -th decomposition block in the p-th layer, represents the projection of the -th extracted trend, is the mixture of experts decomposition block, is the frequency enhancement attention module, is the frequency enhancement attention mechanism, is the feed-forward layer; the final prediction is the sum of these two refined decomposition components, wherein is the season component after depth conversion projected onto the target dimension.
[0121] The Fourier enhancement structure uses the discrete Fourier transform DFT, and the DFT is defined as , where \(i\) is the imaginary unit, , is a complex sequence in the frequency domain. Similarly, the inverse DFT is defined as .
[0122] Combined with Figure 3 shown, the input of the Fourier enhancement structure ( ) is first linearly projected using i.e., . Then \(q\) is transformed from the time domain to the frequency domain, and the Fourier transform of \(q\) is denoted as . Among them, \(D\) is the parameter of the model hidden layer. In the frequency domain, only randomly selected \(M\) modes are retained, and the selection operator ,
[0123] ;
[0124] Among them, is the Fourier transform, and \(M\) is much smaller than \(N\);
[0125] ;
[0126] Among them, is the Fourier transform formula of the frequency enhancement attention module, is the inverse Fourier transform, is a randomly initialized parameterized kernel; let and ,
[0127] is namely: , among which, is the input channel, i.e., and is the output channel; Padding(·) is to zero-pad the result of to , and then perform the inverse Fourier transform back to the time domain.
[0128] The frequency enhancement attention module uses the Fourier transform frequency enhancement attention mechanism. The input is the query, key, and value. In cross-attention, the query comes from the decoder, and the key and value come from the encoder, and are obtained through the following three formulas:
[0129] ;
[0130] ;
[0131] ;
[0132] Among them, , is the weight matrix. The attention formula is:
[0133] ;
[0134] In FEA-f, the query, key, and value are transformed by Fourier transform, and by randomly selecting M modes, a similar attention mechanism is performed in the frequency domain. Denote the Fourier-transformed representation as , and the operation formula in FEA-f is as follows,
[0135] ;
[0136] ;
[0137] ;
[0138] ;
[0139] Among them, is the activation function. Let and , and zero-padding is performed before inverse Fourier transform.
[0140] S203. Construct a spatio-temporal feature analysis branch:
[0141] The structure of this branch is as Figure 4 shown. The data processing process within the spatio-temporal feature analysis branch is as follows:
[0142] (1) First, project the original sensor data matrix of the spatio-temporal feature analysis branch into by linear transformation, where P is the length of the input sensor data matrix sequence, N is the number of sensors, C is the number of features, and d is a parameter;
[0143] (2) Then input into the spatio-temporal encoder. This encoder is connected with L stacked spatio-temporal attention modules and are jointly used as the long-term spatio-temporal dynamic information of the sensor data and embedded into (the long-term spatio-temporal dynamic information is encoded and integrated to obtain the latent feature space matrix ), and ;
[0144] (3) Then enter the frequency enhancement attention module, and the formula is:
[0145] ;
[0146] Among them, is the output of the frequency enhancement attention module, is the Fourier transform formula of the frequency enhancement attention module, is the spatio-temporal embedding matrix of the input sensor data signal matrix sequence, is the spatio-temporal embedding matrix of the output sensor data signal matrix sequence, is the matrix output in step (2).
[0147] (4) Next, enter the spatio-temporal prediction decoder, whose input is the output of the frequency enhancement attention module, that is, ; The spatio-temporal prediction decoder stacks spatio-temporal attention modules to capture spatio-temporal dynamics in the latent space, and finally projects to the prediction target, that is, future sensor values, through linear projection . Among them, Figure 4 y in is a predicted value of the output, that is, the value of a sensor at a certain moment, is a matrix containing the predicted values of all sensors, and y is an element of.
[0148] S204. Construct a spatio-temporal self-supervised learning branch:
[0149] There is redundancy in sensor data in the time dimension, so it is necessary to mask as much as possible the continuous part of the sensor data sequence to construct a self-supervised learning task to prevent it from recovering a discrete missing value from adjacent time windows. Moreover, masking should be performed along both the time and space dimensions to capture spatio-temporal dependencies.
[0150] The present invention introduces a masking algorithm for spatio-temporal data for spatio-temporal self-supervised learning of sensor data.
[0151] First, divide the sensor data signal matrix into small pieces, that is, a continuous data signal segment of length . Then, randomly mask patches with spatio-temporally agnostic sampling. Each step will be introduced in detail later.
[0152] The spatio-temporal self-supervised learning branch includes a spatio-temporal encoder II, a self-supervised learning module, and a spatio-temporal reconstruction decoder.
[0153] As a preferred implementation, in the spatio-temporal self-supervised learning branch, based on the masking algorithm for spatio-temporal feature data, self-supervised learning is performed on sensor data; first, divide the sensor data signal matrix into small pieces, that is, a continuous data signal segment of length ; then, randomly mask patches with spatio-temporally agnostic sampling.
[0154] Among them, the steps of self-supervised learning of sensor data are as follows:
[0155] (1) Input the original sensor data matrix , where P is the length of the input sensor data matrix sequence, N is the amount of sensor data, and C is the number of features;
[0156] (2) Generate a mask tensor , and set the elements of the tensor to 1;
[0157] (3) Divide the sensor data signal matrix into patches, i.e., , ;
[0158] (4) Calculate the number of masked patches (round down);
[0159] (5) Randomly draw from patches, and record the corresponding patch indices as ;
[0160] (6) Loop through ;
[0161] (7) Set all elements of the -th patch in the tensor to 0;
[0162] (8) Output the mask tensor .
[0163] The encoder in the spatio-temporal self-supervised learning branch is the same as the encoder used in the spatio-temporal feature analysis branch. In the spatio-temporal encoder of the spatio-temporal self-supervised learning branch, by setting the masked value in the input of the softmax function to -∞, the masked sensor data does not participate in the calculation, and the output of the spatio-temporal encoder in the spatio-temporal self-supervised learning branch is denoted as .
[0164] The input of the spatio-temporal reconstruction decoder is denoted as , which consists of two parts. One part is the output of the spatio-temporal encoder in the spatio-temporal self-supervised learning branch , and the other part is the mask token vector, which occupies all the masked positions and reveals the absence of the true value; the mask token vector is shared and trained with the entire model, indicating the existence of masked patches that need to be restored.
[0165] The reconstruction decoder stacks spatio-temporal attention modules to restore the sensor data in the latent space, and finally outputs , which matches calculated by the spatio-temporal encoder in the spatio-temporal feature analysis branch.
[0166] S205. Define the loss function and evaluation metrics.
[0167] Three commonly used evaluation metrics are selected as the prediction accuracy metrics for the model, namely Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE). MAE is selected as the loss function, and the true sensor data and the predicted values of the sensor data are used to calculate the evaluation metrics. The specific formulas can refer to the existing technology and will not be elaborated here. For Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE), the closer the value is to 0, the more accurate the predicted data is.
[0168] Step 4, Anomaly warning: Determine the confidence interval by combining the predicted value and the confidence level output in Step 3, and use it as the normal data range for industrial control network detection. When the current actual sensor data exceeds the normal range, trigger a network anomaly alarm.
[0169] Specifically, the predicted value of the future sensor value , and the 95% confidence interval is . is the lower limit value of the corresponding confidence interval, and is
[0170] As a preferred embodiment, in an industrial control system, network load is a key parameter, and their abnormal fluctuations may indicate equipment failures or cyberattacks.
[0171] As an application example, an unsupervised anomaly detection method based on a spatio-temporal autoencoder proposed in this embodiment is specifically used for anomaly detection of network load (bandwidth, latency, throughput, network traffic, etc.).
[0172] First, collect the historical sensor data of the network load in the industrial control system, clean the data to remove noise and outliers, and standardize the data to make it have a unified dimension and range. Use the sliding window technique to extract multiple subsequences from the historical data to form sample sequences, and each sample sequence contains the network load data within a certain time window.
[0173] Next, construct an unsupervised deep learning model based on a spatio-temporal autoencoder, and the spatio-temporal feature analysis branch is responsible for analyzing the spatio-temporal features of the network load data.
[0174] The spatio-temporal self-supervised learning branch performs self-supervised learning on the data through a masking algorithm to improve the generalization ability of the model. Through a feed-forward neural network, preliminary feature extraction and non-linear transformation are performed on the sample sequence to integrate spatial and temporal information, generating an embedded representation that fuses spatio-temporal information. The embedded representation that fuses spatio-temporal information is input into the spatio-temporal encoder-decoder, and the long-term spatio-temporal dynamic information of the network load data is captured through the spatio-temporal attention module. In self-supervised learning, the network load data is masked, and part of the data is randomly masked, forcing the model to learn from the remaining data and predict the masked data. The masked data is complemented and reconstructed through the spatio-temporal reconstruction decoder, and compared with the original data to optimize the model performance.
[0175] Finally, the real-time network load data is input into the trained model to predict the sensor data (network load data) in the future time period.
[0176] The confidence space is determined based on the predicted value and confidence level output by the model as the normal data range. When the real-time data exceeds the normal range, an anomaly alarm is triggered.
[0177] By comparing the network load data predicted by the model with the actual data, the accuracy and robustness of the model are evaluated. In practical applications, this method can effectively identify abnormal fluctuations in network load, issue early warnings in a timely manner, and reduce potential equipment failures and safety risks. This method not only improves the accuracy of anomaly detection, but also enhances the generalization ability and robustness of the model, providing a strong guarantee for the stable operation of industrial control systems.
[0178] To verify the effectiveness and advancement of the model, this embodiment selects the real datasets Abilene and GEANT datasets to carry out anomaly detection work on industrial control networks. The Abilene dataset collects the traffic data statistics of 24 weeks with 12 nodes in a certain year. A sample is generated every 5 minutes, totaling 48,384 samples. Each sample consists of a 12×12 data matrix. The network composition of GEANT includes 23 nodes and 38 links. The complete dataset consists of 4 months of statistical data, with a collection interval of 15 minutes. There are 10,773 samples in total, and each sample consists of 23×23 data. The dataset is divided into a training set, a validation set, and a test set according to the ratio of 6:2:2 in terms of time. The model receives the data of the past 4 hours to predict the network load in the next 4 hours.
[0179] The experiment was carried out on a computer with an NVIDIA GeForce RTX 3090 graphics card. The experimental environment is based on Python language version 3.12, and the deep learning framework PyTorch version 2.5.1 is used to build and implement the studied model.
[0180] Use the trained model for prediction. In this embodiment, network traffic is taken as an example for prediction, and the prediction results are as follows Figure 6 as shown.
[0181] To prove the effectiveness of the proposed scheme, ARIMA, SVR, GRU, GCN, and T-GCN are also selected in this scheme for comparative experiments with the method proposed in the present invention. The detailed information of the experimental effects is shown in Table 1 below, which are the comparisons regarding MAE, MSE, and RMSE respectively.
[0182] Table 1 Comparison of MAE, MSE, and RMSE between the present invention and existing methods
[0183]
[0184] From the comparison results, it can be seen that the unsupervised deep learning model based on the spatio-temporal codec has significant advantages in prediction accuracy. Taking the Abilene dataset as an example, compared with the ARIMA model, this model has made significant progress in the accuracy of network traffic prediction. This model has achieved a 4.65% reduction in MAE and a 12.63% reduction in RMSE. Compared with the SVR model, this model has reduced by 3.78% in MAE and 9.93% in RMSE. Compared with the T-GCN model, this model has achieved a 1.67% reduction in MAE and a 3.39% reduction in RMSE. The effect differences between this model and other methods on the Abilene dataset are as follows Figure 7 as shown. The parameters of some predicted network traffic values output are shown in Table 2 below.
[0185] Table 2 Parameters of traffic true values and predicted values
[0186]
[0187] The comparison between the prediction results and the true values is as follows Figure 8 as shown. It can be seen that the predicted values and the true values are within a very small gap. Next, determine the confidence space for anomaly detection based on the prediction results. Using a 95% confidence space, the results are shown in Table 3.
[0188] Table 3 Confidence space for traffic anomaly detection
[0189]
[0190] The confidence interval is the normal range. If the actual value exceeds this normal range, an anomaly alarm will be triggered.
[0191] As another embodiment of the present invention, there is provided an unsupervised anomaly detection system for industrial control networks based on a spatio-temporal codec, which is used to implement the unsupervised anomaly detection method for industrial control networks based on a spatio-temporal codec as described above. For example, Figure 1 and Figure 5 As shown, the system includes a data processing module, a sample construction module, an unsupervised deep learning model based on a spatio-temporal codec, and an anomaly warning module.
[0192] That is to say, finally, through the various modules constructed above, the present invention forms an industrial control network anomaly detection system. This anomaly detection system includes a (sensor data) data processing module, a (sensor data) sample construction module, a model construction module, a (sensor data) spatio-temporal feature analysis module, and an anomaly warning module. This system executes the above method to perform industrial control network anomaly detection.
[0193] The data processing module performs data cleaning and standardization operations on the original sensor data, and according to the characteristics of time-series sensor data, mines more time-related features. This sensor data network is denoted as G, and the sensor data signal matrix at time slice t is , where the sensor data network G consists of a group of sensors, denoted as the node set V, and , is the feature vector of node v observed at time t, and C is the number of features.
[0194] The sample construction module is used to collect historical sensor data to obtain a sensor data sequence, obtain multiple sensor data subsequences through a sliding window form, and use the normalized sensor data subsequences as sample sequences. The specific method of normalization is: first calculate the maximum and minimum values of the sequence data, denoted as and respectively, and perform calculation on each data in the sequence data, and use the normalized sensor data subsequences as sample sequences.
[0195] The model construction module uses each sensor data subsequence to train an unsupervised deep learning model based on a spatio-temporal codec, and places the spatio-temporal embedding module, spatio-temporal attention module, and frequency enhancement attention module for sensor data that fuse spatio-temporal features in the above method to construct an industrial control network anomaly detection model, and performs model training based on the corresponding sample sequences to obtain a trained industrial control network anomaly detection model.
[0196] The spatio-temporal feature analysis module, including a spatio-temporal feature analysis branch and a self-supervised learning branch, inputs each sensor data sequence to be measured into the corresponding self-supervised spatio-temporal frequency enhanced attention model for sensor data prediction, and outputs the sensor data in a future predetermined time period as the prediction result. This module mainly utilizes the spatio-temporal feature analysis branch and the self-supervised learning branch in the above method, and the specific data processing and function implementation process will not be elaborated here.
[0197] The anomaly warning module judges by comparing the real-time actual sensor value with the confidence space. If it falls within the normal range, it works normally. When the actual sensor value at the current moment is not within the normal range, it triggers an industrial control network anomaly warning to notify relevant personnel or automatically activate preset response measures.
[0198] In summary, the present invention combines the self-supervised learning technology and the spatio-temporal frequency enhanced attention mechanism, combines the advantages of both, improves the data utilization efficiency, and enables the anomaly detection model to have good generalization ability. By fusing time and space features, an anomaly detection method suitable for industrial control networks, namely an unsupervised anomaly detection method for industrial control networks based on spatio-temporal codecs, is proposed and formed into a system. Using a multi-task learning framework, multiple tasks are learned simultaneously, improving the ability to extract features in long time slices, breaking through the technical bottlenecks of the current model's poor ability to extract long-term global spatio-temporal features and high time complexity, enhancing the ability to capture long-term dynamic information and global correlation information, and improving the accuracy of anomaly detection in the field of industrial control.
[0199] Certainly, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Any changes, modifications, additions, or substitutions made by those of ordinary skill in the art within the scope of the essence of the present invention shall fall within the protection scope of the present invention.
Claims
1. An unsupervised anomaly detection method for industrial control networks based on spatiotemporal codec, characterized in that: The following steps are involved: Step 1: Sensor data processing and sample construction: Sampling and data processing of historical industrial control sensor raw data and time series to obtain sample sequences; Step 2: Build and train an unsupervised deep learning model based on spatiotemporal codec; The unsupervised deep learning model based on the spatiotemporal codec includes a spatiotemporal feature analysis module, which includes a spatiotemporal feature analysis branch and a spatiotemporal self-supervised learning branch, which are used for spatiotemporal feature analysis and spatiotemporal self-supervised learning, respectively; the spatiotemporal feature analysis branch and the spatiotemporal self-supervised learning branch both adopt a spatiotemporal encoder-decoder structure, specifically, the spatiotemporal feature analysis branch adopts a spatiotemporal codec I, including a spatiotemporal encoder I, a frequency enhancement attention module and a spatiotemporal prediction decoder, and the spatiotemporal encoder I is superimposed with L spatiotemporal attention modules; the spatiotemporal self-supervised learning branch adopts a spatiotemporal codec II, including a spatiotemporal encoder II, a self-supervised learning module, and a spatiotemporal reconstruction decoder; wherein the spatiotemporal codec I has the same structure as the spatiotemporal encoder II; In the spatiotemporal feature analysis branch, the long-term spatiotemporal dynamic information of the sensor data is encoded and integrated through the spatiotemporal encoder I to obtain the potential feature space matrix S (L) , and then the latent feature space matrix S is captured by the frequency enhancement attention module and the spatiotemporal prediction decoder (L) The spatiotemporal dynamics of In the spatiotemporal self-supervised learning branch, a masking algorithm for spatiotemporal data is introduced, and patches are randomly masked using spatiotemporal agnostic sampling. The masked data is completed and reconstructed through a spatiotemporal reconstruction decoder, and the output is The reconstructed data is compared with the S calculated by the spatiotemporal encoder I in the spatiotemporal feature analysis branch. (L) Place them in the latent space for alignment and comparison, and the results continuously guide the improvement of the spatiotemporal encoder I of the spatiotemporal feature analysis branch; Step 3: Spatiotemporal feature analysis of sensor data: The data output from step 1 is input into the trained unsupervised deep learning model based on spatiotemporal codec. The spatiotemporal feature analysis branch and the spatiotemporal self-supervised learning branch work together to output the sensor data of the future predetermined period as the predicted value. The following operations are performed on spatial and temporal attention: the output of the (l-1)th spatiotemporal attention module is represented as S (l-1) ∈R P×N×d , where P is the length of the input sensor data matrix sequence and N is the number of sensors; connect it with the spatiotemporal embedding matrix E of the input sensor data matrix sequence to obtain Z (l-1) ∈R P×N×2d , and Z (l-1) Input the lth spatiotemporal attention module; use Represents the input of node v on all time slices, using Represents the input of all nodes on time slice t; In the time dimension, the time vector T′ of the sensor data signal is It is defined as the time reference point in the temporal attention of the lth spatiotemporal attention module. The calculation formula in the temporal attention module is as follows: Among them, MHSA(·) is the multi-head self-attention formula. The temporal attention module first pays attention to the input Will Convert to Yt′ r , the updated time reference point Yt′ r Contains the original long input The global sensor data information finally generates the time encoding representation of node v The temporal attention module processes the input of each node in parallel, using T l ∈R P ×N×2d represents the time-coded representation of all nodes; Similarly, in the spatial dimension, the spatial O′ vector is introduced As a spatial reference point, the spatial attention module then processes the input on each time slice in parallel, with the following formula: in, is the spatial encoding representation of all nodes at time t, S l ∈R P×N×2d represents the spatial encoding representation of all time slices; finally, the output of the spatial attention module is S (l) of and; Step 4: Abnormal alarm: Determine the confidence space by combining the predicted value and confidence level output in step 3, and use it as the normal data range for industrial control network detection. When the actual sensor data exceeds the normal range, a network abnormal alarm is triggered.
2. The unsupervised anomaly detection method for industrial control networks based on spatiotemporal codec according to claim 1 is characterized in that: In step 3, before the sensor data is input into the spatiotemporal feature analysis branch and the spatiotemporal self-supervised learning branch, the following processing is performed: the sensor data is first subjected to preliminary feature extraction and nonlinear transformation through a feedforward neural network to obtain the sensor data signal matrix E, and then the spatial and temporal information is integrated through the spatiotemporal feature embedding module to generate an embedded representation of the fused spatiotemporal information, and the spatiotemporal embedding matrix E′ of the sensor data signal matrix sequence is output.
3. The unsupervised anomaly detection method for industrial control networks based on spatiotemporal codec according to claim 2 is characterized in that: In step 3, the data processing process in the spatiotemporal feature analysis branch is as follows: (1) First, the original sensor data matrix X of the spatiotemporal feature analysis branch is linearly transformed to obtain S (0) ; (2) Then S (0) Input spatiotemporal encoder I, which is connected to superimpose L spatiotemporal attention modules, S (0) Together with E′, it is the long-term spatiotemporal dynamic information of sensor data. After encoding and integration, the potential feature space matrix S is obtained. (L) ; (3) Then enter the frequency enhancement attention module, the formula is: S′ (0) =FEA-f(E′,E,S (L) )∈R Q×d ; Among them, S′ (0) is the output of the frequency enhancement attention module, FEA-f(·) is the Fourier transform formula of the frequency enhancement attention module, L is the number of spatiotemporal attention modules, Q is the length of the output sensor data matrix sequence, d is a parameter, E is the spatiotemporal embedding matrix of the input sensor data signal matrix sequence, and E′ is the spatiotemporal embedding matrix of the output sensor data signal matrix sequence; (4) Next, we enter the spatiotemporal prediction decoder, whose input is the output of the frequency enhancement attention module, i.e., S′ (0) ; The spatiotemporal prediction decoder stacks L' spatiotemporal attention modules to capture the spatiotemporal dynamics in the latent space and finally maps it to the prediction space through linear projection. The target is the sensor value in the future 4. The unsupervised anomaly detection method for industrial control networks based on spatiotemporal codec according to claim 3 is characterized in that: The spatiotemporal attention module includes a spatial attention module and a temporal attention module. The implementation process of the spatial attention module and the temporal attention module is the same. The implementation process of the spatial attention module is as follows: a. Receive the input sensor data signal tensor X′, spatiotemporal embedding STE and mask tensor mask; b. Concatenate the input sensor data signal tensor X′ and the spatiotemporal embedding STE along the feature dimension to obtain a new tensor X”; c. Initialize the query and key weight matrices of the multi-head attention mechanism using learnable parameters; d. Perform multi-head attention calculation on the initialized query and key weight matrix G and tensor X' to obtain the attention weighted result h; e. Perform multi-head attention calculation on the weighted attention result h and the input tensor X′ to obtain the final result of the spatial attention mechanism.
5. The unsupervised anomaly detection method for industrial control networks based on spatiotemporal codec according to claim 1 is characterized in that: The frequency enhancement attention module includes an encoder, a decoder and a Fourier enhancement structure, wherein the encoder and the decoder both adopt a multi-layer structure, and a Fourier enhancement structure is used between the encoder and the decoder. The Fourier enhancement structure adopts Fourier transform and inverse Fourier transform, and uses the Fourier transform frequency enhancement attention mechanism to process the encoder data. The Fourier enhancement structure converts time domain data into frequency domain data by applying Fourier transform, and executes the Fourier transform frequency enhancement attention mechanism in the frequency domain, and applies inverse Fourier transform to convert frequency domain data into time domain data.
6. The unsupervised anomaly detection method for industrial control networks based on spatiotemporal codec according to claim 1 is characterized in that: The encoder in the spatiotemporal self-supervised learning branch is the same as the encoder used in the spatiotemporal feature analysis branch. In the spatiotemporal encoder of the spatiotemporal self-supervised learning branch, the mask value in the softmax function input is set to -∞, so that the masked sensor data does not participate in the calculation. The output of the spatiotemporal encoder of the spatiotemporal self-supervised learning branch is recorded as The input of the spatiotemporal reconstruction decoder is recorded as It consists of two parts. One part is the spatiotemporal encoder output of the spatiotemporal self-supervised learning branch. The other part is the mask marker vector, which occupies all mask positions and reveals the absence of true values; The reconstruction decoder stacks L″ spatiotemporal attention modules to recover the sensor data in the latent space and finally outputs The S calculated by the spatiotemporal encoder in the spatiotemporal feature analysis branch (L) Match.
7. The unsupervised anomaly detection method for industrial control networks based on spatiotemporal codec according to claim 6 is characterized in that: In the spatiotemporal self-supervised learning branch, a masking algorithm based on spatiotemporal feature data is used to perform self-supervised learning on sensor data. First, the sensor data signal matrix is divided into small blocks, i.e., a block with a length of l. m Then, the patch is randomly masked using spatiotemporal agnostic sampling; The steps for self-supervised learning of sensor data are as follows: (1) Input raw sensor data matrix X∈R P×N×C , P is the length of the input sensor data matrix sequence, N is the number of sensors, and C is the number of features; (2) Generate mask tensor u∈R P×N×C , the elements of the tensor are set to 1; (3) Divide the sensor data signal matrix into patches, namely {1,2,...,N patch }, (4) Calculate the number of mask patches n m =α m ×N patch ; (5) Randomly select {1,2,...,N patch } extract n m The corresponding patch index is denoted as P sample ={p1, p2, ..., p m }; (6) Loop through P sample ; (7) The pth i All elements of the patch are set to 0; (8) Output mask tensor u.
8. Unsupervised anomaly detection system for industrial control networks based on spatiotemporal codec, characterized in that: Used to implement the unsupervised anomaly detection method for industrial control networks based on a spatiotemporal codec as described in any one of claims 1 to 7, the system includes a data processing module, a sample construction module, an unsupervised deep learning model based on a spatiotemporal codec, and an anomaly warning module; The data processing module performs data cleaning and standardization operations on the raw sensor data; The sample construction module is used to collect historical sensor data to obtain a sensor data sequence, obtain multiple sensor data subsequences in a sliding window form, and use the normalized data subsequences as sample sequences; The unsupervised deep learning model based on the spatiotemporal codec includes a spatiotemporal feature analysis module, which includes a spatiotemporal feature analysis branch and a self-supervised learning branch, performs sensor value prediction for each sensor data sequence to be tested, and outputs the sensor value of a predetermined period in the future as a predicted value; The abnormal warning module is used to determine the confidence space as the normal range of the sensor value, and trigger the industrial control network abnormal warning when the actual sensor value at the current moment is not within the normal range.
Citation Information
Patent Citations
Monitoring video anomaly detection method and system based on dynamic self-supervised network
CN116883896A