Industrial control network spatial-temporal characteristic lightweight anomaly detection method based on knowledge distillation
By applying knowledge distillation technology in industrial control network abnormal detection, the spatial and temporal characteristics are extracted and fused, the spatial and temporal characteristics fusion problems and high computational complexity in the existing methods are solved, and efficient and reliable abnormality detection is achieved.
Patent Information
- Application Number
- CN202510368328.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-07-01
AI Technical Summary
The existing industrial control network abnormality detection methods are difficult to effectively integrate spatial and temporal characteristics, and the calculation complexity is high, making it difficult to meet the real-time requirements of the industrial control environment.
Using a knowledge distillation-based method, the spatial and temporal characteristics are extracted and fused through multi-level knowledge distillation technology, the complexity of the model calculation is reduced, and abnormal detection is achieved through dynamic weight adjustment and multi-dimensional scoring mechanisms.
It realizes the accuracy of detection while significantly reducing the complexity of model calculation, making the detection method more suitable for the actual deployment requirements of industrial control network environments and has efficient and reliable abnormal detection capabilities.
Smart Images

Figure CN120238341A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial control network security, and specifically relates to a lightweight anomaly detection method for industrial control network spatiotemporal features based on knowledge distillation. Background Art
[0002] As the nervous system of industrial production, the security of industrial control networks is directly related to the stable operation of the national economy and social production. Traditional industrial control network security protection mainly relies on boundary protection methods such as firewalls and intrusion detection, but in the face of increasingly complex network attacks, these methods can no longer meet actual needs.
[0003] At present, there are several main problems in the detection of industrial control network traffic anomalies: First, existing detection methods often separate time features and spatial features and process them separately, ignoring the inherent correlation between the two, resulting in limited detection accuracy. Among them, time features mainly reflect the dynamic behavior characteristics of device communication, such as periodicity and timing, while spatial features reflect static structural characteristics such as topological relationships and communication modes between devices. Secondly, mainstream deep learning models (such as CNN, RNN, etc.) generally have high computational complexity and high resource consumption in industrial control network anomaly detection, which makes it difficult to meet the strict real-time requirements of industrial control environments.
[0004] In recent years, researchers have proposed a variety of improvement schemes. Some scholars use LSTM networks to extract temporal features, but it is difficult to handle long-term dependencies; some studies use graph neural networks to model spatial relationships between devices, but the computational overhead is large. In terms of model lightweighting, knowledge distillation technology shows good application prospects. However, since industrial control networks contain both spatial and temporal features, their feature representation is relatively complex, and traditional knowledge distillation is difficult to simultaneously maintain the structural information of these two features. In addition, since industrial control systems have special expertise and special conditions, ordinary knowledge distillation methods cannot effectively transfer these domain-specific knowledge. Therefore, there is currently a lack of research on knowledge distillation in the field of industrial control network anomaly detection. How to effectively integrate spatiotemporal features and maintain the lightweight characteristics of the detection model is still a key technical problem that needs to be solved in this field.
[0005] Therefore, an industrial control network anomaly detection method with high computational efficiency and good generalization ability has important theoretical value and practical significance. Summary of the invention
[0006] Aiming at the deficiencies of the existing technology, the technical problem to be solved by the present invention is to provide a lightweight anomaly detection method for spatio-temporal features of industrial control networks based on knowledge distillation. This detection method extracts and fuses spatial and temporal features respectively, and uses multi-level knowledge distillation technology to significantly reduce the model calculation complexity while ensuring the detection accuracy, making it more suitable for the actual deployment requirements of industrial control network environments.
[0007] The technical solution adopted by the present invention to solve the above technical problem is as follows:
[0008] A lightweight anomaly detection method for spatio-temporal features of industrial control networks based on knowledge distillation, the method includes the following:
[0009] Obtain industrial control network traffic data and perform preprocessing to obtain preprocessed network traffic features, including preprocessed frequency domain features F freq , preprocessed statistical features F stat and preprocessed time series features F temp ;
[0010] Construct a lightweight anomaly detection model for spatio-temporal features, including a spatial feature extraction module, a time series feature extraction module, a fusion module, a knowledge distillation module, and an anomaly detection module;
[0011] The time series feature extraction module is used to process the preprocessed time series features F temp and statistical features F stat to obtain a time series feature representation F t ;
[0012] The spatial feature extraction module is used to process the preprocessed frequency domain features F freq and statistical features F stat to obtain a spatial feature representation F s ;
[0013] The fusion module is used to fuse the time series feature representation F t and the spatial feature representation F s to obtain a normalized fused feature F norm ;
[0014] The knowledge distillation module includes a teacher model and a student model. The teacher model uses the normalized fused feature F norm as input to train the student model. During training, three levels of loss constraints, namely spatial knowledge distillation, time series knowledge distillation, and task-specific knowledge distillation, are set. The comprehensive loss function L total is;
[0015] L total =α s (L feature +Lrelation +L topo )+α t (L hidden +L temporal +L periodic )+α a (L anomaly +L boundary )
[0016] Among them, α s , α t , α a are the weight coefficients of the losses at each level, and the three are uniformly represented by α m . m ∈ {s, t, a}, and according to α m =softmax(v m ·performance m ), it is dynamically adjusted through the validation set. performance m represents the effect measure of knowledge transfer at each level, and v m is the weight vector parameter used to adjust the importance of knowledge transfer at different levels; L feature is the distillation loss at the feature level; L relation is the relational consistency loss; L topo is the topology preservation loss; L hidden is the hidden layer state distillation loss; L temporal is the temporal relationship loss; L periodic is the periodicity preservation loss; L anomaly is the anomaly discrimination knowledge distillation loss; L boundary is the boundary awareness loss;
[0017] Using the trained student model as the lightweight student model of the anomaly detection module, the lightweight student model is used to perform anomaly detection on real-time industrial control network traffic data.
[0018] Furthermore, spatial knowledge distillation mainly focuses on the transfer of interaction relationships and topological structure information between devices, including the distillation loss at the feature level, the relational consistency loss, and the topology preservation loss. The distillation loss at the feature level is:
[0019]
[0020] Among them, T s and S s respectively represent the spatial feature representations of the teacher model and the student model;
[0021] The relational consistency loss is:
[0022] L relation =∥G t (T s )-Gs (S s )∥1
[0023] Among them, G t (·) and G s (·) represent the construction functions of the feature relationship diagrams of the teacher model and the student model, which are obtained by calculating the similarity matrix between feature vectors;
[0024] The topological preservation loss is:
[0025] L topo = KL(P t ∥P s ) + λ t ·R topo (S s )
[0026] Among them, P t and P s are the topological prediction distributions of the teacher model and the student model respectively, R topo is the topological regularization term, KL is the Kullback-Leibler divergence, and λ t is the weight coefficient;
[0027] Temporal knowledge distillation focuses on transferring time series patterns and dynamic behavior features, including hidden layer state distillation loss, temporal relationship loss, and periodicity preservation loss.
[0028] The hidden layer state distillation loss is:
[0029] L hidden = MSE(h t , h s ) + λ h ·KL(P(h t ) ∥ P(h s ))
[0030] Among them, h t and h s are the hidden layer states of the teacher model and the student model respectively, P(·) represents the state distribution, MSE is the mean square error, and λ h is the weight coefficient;
[0031] The temporal relationship loss is:
[0032] L temporal = ∑|r t (i, j) - r s (i, j)|
[0033] Among them, r t (i, j) and r s(i,j) represents the relationship strength between the time steps i and j of the teacher model and the student model;
[0034] The periodicity preservation loss is:
[0035]
[0036] where T seq and S seq represent the temporal features of the teacher and student models, and FFT(·) represents the fast Fourier transform, which is used to capture the periodic features of the signal;
[0037] The task-specific knowledge distillation includes the anomaly discrimination knowledge distillation loss and the boundary awareness loss.
[0038] The anomaly discrimination knowledge distillation loss is:
[0039] L anomaly = CE(y t , y s ) + λ a ·KL(q t ∥q s )
[0040] where y t , y s are the anomaly discrimination results, q t , q s are the anomaly probability distributions, CE is the cross entropy, and λ a is the weight coefficient;
[0041] The boundary awareness loss is:
[0042]
[0043] where w i is the sample weight, which assigns a higher weight to the boundary samples, and f t (.) and f s (.) represent the feature extraction functions of the specific tasks of the teacher model and the student model respectively, and f t (x i ) and f s (x i ) are the feature representations of the specific tasks extracted by the teacher model and the student model for the boundary sample x i .
[0044] Furthermore, the preprocessing includes data cleaning, standardization, and preliminary feature extraction. The improved Z-score is used for preliminary cleaning, and the specific process is:
[0045] Z(x) = (x - μ rolling ) / σ rolling
[0046] Among them, x is the original data, μ rolling and σ rolling are the mean and standard deviation within the sliding window, Z(x) is the data value after cleaning, and the window size is dynamically adjusted according to the temporal characteristics of the data:
[0047] w = min(w max , max(w min , T cycle ))
[0048] Among them, w is the size of the final window, w max is the maximum allowable window size, w min is the minimum allowable window size, T cycle is the period estimation value of the data, obtained through autocorrelation analysis;
[0049] After cleaning, standardize different data types:
[0050] For continuous data, use improved MinMax standardization:
[0051] x' = (x - x min ) / (x max - x min + ε)
[0052] Among them, ε is a smoothing factor used to handle the influence of extreme values, x is the original data, x min , x max are the minimum and maximum values in the dataset, and x' is the standardized data value;
[0053] For periodic data, use phase standardization; for discrete data, use one-hot encoding conversion; the standardization parameters are updated online;
[0054] After standardization, perform preliminary feature extraction: Statistical features include the mean of data values, variance of data values, skewness, and kurtosis; Temporal features include the difference of data, mean and standard deviation of the sliding window; Frequency domain features take the top k frequency components;
[0055] Through feature importance scoring, filter respectively to obtain the preprocessed statistical features, temporal features, and frequency domain features.
[0056] Furthermore, the spatial feature extraction module adopts a Transformer architecture, a CNN architecture, or a hybrid architecture.
[0057] Furthermore, the temporal feature extraction module adopts a Bi-LSTM architecture, an LSTM architecture, a GRU architecture, or an RNN architecture.
[0058] Further, after the real-time industrial control network traffic data is preprocessed and then processed by the lightweight student model, spatial features, temporal features, and specific task features are obtained. In the anomaly detection module, a multi-dimensional scoring mechanism is adopted, comprehensively considering the anomaly degrees of spatial features, temporal features, and specific task features. The calculation formula for the total anomaly score Score(x) is:
[0059] Score(x) = β1S s (x) + β2S t (x) + β3S a (x)
[0060] where S s , S t and S a represent the anomaly scores of spatial feature representation, temporal features, and specific task features respectively; β1, β2, and β3 are adaptive weight coefficients, dynamically updated through m, i ∈ {s, t, a}, v s represents the importance score of spatial features, v t represents the importance score of temporal features, and v a represents the importance score of specific task features, dynamically adjusted according to historical detection effects;
[0061] The anomaly scores of spatial feature representation, temporal features, and specific task features are all calculated using a Mahalanobis distance-based scoring method,
[0062]
[0063] where m ∈ {s, t, a}, μ m and Σ m are the mean vector and covariance matrix of the corresponding features respectively, and x represents the input;
[0064] The covariance matrix Σ m is maintained by an incremental update method:
[0065]
[0066] where η is the learning rate, is the updated covariance matrix;
[0067] The detection decision rule is:
[0068] If
[0069] Score(x) > θ dynamic
[0070] Return anomaly
[0071] Otherwise:
[0072] Return to normal
[0073] Dynamic threshold θ dynamic Update in the following way:
[0074] θ dynamic = μ score + r·σ score
[0075] Where μ score and σ score are the mean and standard deviation of historical scores respectively, and r is an adjustable sensitivity parameter.
[0076] Furthermore, a sliding window mechanism is adopted to smooth the decision result, and an abnormal alarm is triggered only when Q consecutive samples exceed the dynamic threshold. At the same time, an exponential decay factor is introduced to update the historical statistics:
[0077]
[0078] Where α is the smoothing coefficient used to control the influence degree of historical information, and represent the updated mean and standard deviation, and Score current is the total abnormal score of the current sample;
[0079] Update the dynamic threshold with the updated mean and standard deviation.
[0080] Compared with the prior art, the beneficial effects of the present invention are:
[0081] The present invention innovatively proposes a hierarchical knowledge distillation framework (spatial knowledge distillation, temporal knowledge distillation, and task-specific knowledge distillation) and adopts a dynamic weight adjustment mechanism (using the softmax function to dynamically adjust the weights of knowledge distillation at each level) to ensure the effective transfer of knowledge in the three dimensions of space, time, and task, extract a comprehensive distillation loss function, balance the importance of knowledge at the three levels, adopt an adaptive weight mechanism, dynamically adjust the proportion of distillation at different levels, and realize the effective application of knowledge distillation in the field of industrial control network anomaly detection.
[0082] The present invention uses a spatial feature extraction module to capture the interaction patterns between devices, adopts a time feature extraction module to analyze the temporal dependence relationship, realizes the adaptive fusion of spatio-temporal features through a feature fusion module, then uses knowledge distillation technology to transfer the knowledge of a complex model to a lightweight model, and finally, realizes anomaly detection based on a multi-dimensional scoring mechanism and ensures the accuracy of detection through dynamic threshold update, fully considering the characteristics of the industrial control network and realizing efficient and reliable anomaly detection. Description of the Drawings
[0083] Figure 1It is a flow chart of the steps of a lightweight anomaly detection method for spatio-temporal features of industrial control networks based on knowledge distillation according to the present invention.
[0084] Figure 2 It is a schematic structural diagram of a teacher model and a student model in an embodiment. Specific embodiments
[0085] The technical solutions of the present invention will be further described in detail below in conjunction with the accompanying drawings and specific embodiments, but are not used to limit the protection scope of the present application.
[0086] The lightweight anomaly detection method for spatio-temporal features of industrial control networks based on knowledge distillation according to the present invention, as Figure 1 shown, first preprocesses the industrial control network traffic data, and then uses a spatial feature extraction module and a temporal feature extraction module to extract spatial features and temporal features from the preprocessed network traffic features (including the preprocessed frequency domain feature F freq , the preprocessed statistical feature F stat and the preprocessed time series feature F temp ), then fuses the two features using a fusion module, and then uses knowledge distillation technology to perform model compression for space, time, and specific tasks respectively. After training, a lightweight student model is obtained. The lightweight student model is used for anomaly detection. Real-time industrial control network traffic data is input into the lightweight student model. After the lightweight student model outputs spatial features, temporal features, and specific task features, multi-dimensional scoring anomaly detection is performed.
[0087] In the industrial control network environment, "real-time performance" is a key technical requirement, which means that the system needs to respond to events within a specified time limit. For anomaly detection, it is necessary to be able to quickly detect and report anomalies in the network so as to take measures in time. Due to the high computational complexity and large resource consumption of mainstream deep learning models, the detection time cannot be guaranteed to be completed within the millisecond level, and it cannot meet the application scenario requirements of the industrial control system control cycle, which usually requires 1-100 ms. The present invention uses knowledge distillation technology to meet the real-time performance requirements of this application scenario.
[0088] Embodiment 1
[0089] The lightweight anomaly detection method for spatio-temporal features of industrial control networks based on knowledge distillation in this embodiment includes the following steps:
[0090] Step 1: Processing of industrial control network traffic data
[0091] The data preprocessing module is mainly responsible for data cleaning, standardization, and preliminary feature extraction of the original industrial control network data. The industrial control network data has characteristics such as multi-source heterogeneity, noise interference, and time series dependence, and needs to undergo systematic preprocessing to provide a reliable data basis for subsequent analysis.
[0092] Data cleaning includes three links: outlier detection, missing value processing, and noise filtering.
[0093] For the collected industrial control network traffic data X = {x1, x2,..., x n}, the following data preprocessing is carried out, where n is the number of samples.
[0094] An improved Z-score method is used for preliminary cleaning:
[0095] Z(x) = (x - μ rolling ) / σ rolling
[0096] where x is the original data, μ rolling and σ rolling are the mean and standard deviation within the sliding window, Z(x) is the data value after cleaning, and the window size is dynamically adjusted according to the time series characteristics of the data:
[0097] w = min(w max , max(w min , T cycle ))
[0098] Here, w is the size of the final window, w max is the maximum allowed window size, w min is the minimum allowed window size. T cycle is the period estimation value of the data, obtained through autocorrelation analysis.
[0099] Through the preliminary cleaning process, for the detected outliers, the system will adopt different processing strategies according to the importance of the data: key data points are confirmed through manual review, and non-key data points are corrected using statistical interpolation methods. The processes of missing value processing and noise filtering can be implemented according to existing technologies and will not be elaborated here.
[0100] Considering the dimensional differences of different types of data in the industrial control network, a multi-mode standardization strategy is adopted for the cleaned data:
[0101] For continuous data, an improved MinMax standardization is used:
[0102] x' = (x - x min ) / (x max - x min + ε)
[0103] where ε is the smoothing factor used to handle the influence of extreme values, x is the original data, x min , x max are the minimum and maximum values in the dataset, and x' is the normalized data value.
[0104] For periodic data, phase normalization is adopted:
[0105] x′ = 2π * (x - x start ) / T period
[0106] where x is the original data, x start is the starting value, T period is the period length of the data, and x' is the normalized data value (converted to the [0, 2π] interval).
[0107] For discrete data, one-hot encoding transformation is used:
[0108] x′ = OneHot(x)
[0109] where x is the original data, x' is the normalized data value, and the transformed data will become a vector containing only 0s and 1s, with only one position in the vector being 1.
[0110] The normalization parameters are updated online:
[0111]
[0112] where are the updated minimum and maximum values, x current is the current data value, x min , x max are the original minimum and maximum values
[0113] After normalization, preliminary feature extraction is performed, mainly including:
[0114] Statistical features:
[0115] F stat = [mean(x), std(x), skew(x), kurt(x)]
[0116] where mean(x) represents the mean of the data values, std(x) represents the variance of the data values, skew(x) represents the skewness, describing the symmetry of the data distribution, kurt(x) represents the kurtosis, describing the peakedness of the data distribution, and F stat is the extracted statistical feature.
[0117] Temporal features:
[0118] F temp= [diff(x), rolling mean(x) , rolling std(x)
[0119] where diff(x) represents the difference of the data, reflecting the rate of change; rolling mean(x) and rolling std(x) represent the mean and standard deviation of the sliding window, and F temp is the extracted time-domain feature.
[0120] The frequency-domain features are the top k frequency components:
[0121] F freq = FFT(x)[1:k]
[0122] where FFT(x) is the fast Fourier transform, k represents the selected first k frequency components, and F freq is the extracted frequency-domain feature. These three types of features are respectively screened by the feature importance score Sc:
[0123] Sc(f) = MI(f, y) * (1 - max(|corr(f, f i |)))
[0124] where f represents the feature to be evaluated, y is the target variable, MI(f, y) is the mutual information between the feature f and the target variable y, corr(f, f i ) is the correlation coefficient matrix between features, and Sc(f) represents the final score of the feature. In this way, redundant features are removed.
[0125] Through the processing of the above three links, the data preprocessing module can effectively handle various quality problems of industrial control network data, provide a high-quality data basis for subsequent feature extraction and anomaly detection, can dynamically adjust the processing strategy according to the data characteristics, and improve the overall efficiency of the system through preliminary feature extraction.
[0126] In the data processing of the present invention, the window size is self-adaptive, the parameters can be self-adjusted after standardization, the appropriate standardization method can be selected according to different data types, and the Sc(f) function dynamically evaluates the feature importance to determine the final three types of features, namely the statistical features, time series features, and frequency domain features after preprocessing.
[0127] Step 2: Construct a spatio-temporal feature lightweight anomaly detection model,
[0128] The spatio-temporal feature lightweight anomaly detection model includes a spatial feature extraction module, a time series feature extraction module, a fusion module, a knowledge distillation module, and an anomaly detection module;
[0129] The spatial feature extraction module can adopt the Transformer architecture and use the frequency-domain features F freq and statistical features F stat generated by the data preprocessing module as inputs to extract spatial features and obtain the spatial feature representation F s .
[0130] The temporal feature extraction module uses the Bi-LSTM architecture to process the preprocessed temporal features F temp and statistical features F stat to obtain the temporal feature representation F t ;
[0131] The fusion module aims to effectively integrate the outputs of the spatial feature extraction module and the temporal feature extraction module to achieve the adaptive fusion of spatio-temporal features, and is used to fuse the temporal feature representation F t and the spatial feature representation F s to obtain the normalized fused feature F norm . In industrial control network anomaly detection, the importance of temporal features and spatial features will change dynamically with the changes in system states and anomaly types. For example, in detecting unauthorized communications, spatial features may be more important; while in detecting communication temporal anomalies, temporal features play a key role.
[0132] The knowledge distillation module aims to effectively transfer the knowledge of the complex teacher model to the lightweight student model to achieve the efficient deployment of the model. Using the normalized fused feature F norm as the input in the knowledge distillation process can enable the teacher model to obtain the complete context at the same time, thus making more accurate judgments and predictions. In the industrial control network anomaly detection scenario, knowledge distillation not only needs to consider general feature representations and prediction outputs, but also needs to pay special attention to the professional knowledge and constraints of industrial control systems. Therefore, the present invention adopts a hierarchical knowledge distillation framework, including three levels: spatial knowledge distillation, temporal knowledge distillation, and task-specific knowledge distillation. The teacher model includes a temporal module, a spatial module, and a specific task module that can process the input respectively to obtain spatial features, temporal features, and specific task features. The structure of the student model is similar to that of the teacher model. In this embodiment, the temporal module can also adopt the Bi-LSTM architecture, the spatial module can also adopt the Transformer architecture, and the specific task module can adopt the CNN architecture.
[0133] Spatial knowledge distillation mainly focuses on the transfer of interaction relationships and topological structure information between devices, including the distillation loss at the feature level, the relationship consistency loss, and the topology preservation loss. The distillation loss at the feature level is as follows:
[0134]
[0135] Among them, T s and S s respectively represent the spatial feature representations of the teacher model and the student model;
[0136] To maintain the structural relationship of the features, a relational consistency loss is introduced, and the relational consistency loss is:
[0137] L relation =∥G t (T s ) - G s (S s )∥1
[0138] Among them, G t (·) and G s (·) represent the construction functions of the feature relationship graphs of the teacher model and the student model, which are obtained by calculating the similarity matrix between feature vectors.
[0139] Meanwhile, considering the topological constraints of the industrial control network, a topological preservation loss is introduced, and the topological preservation loss is:
[0140] L topo = KL(P t ∥P s ) + λ t ·R topo (S s )
[0141] Among them, P t and P s are the topological prediction distributions of the teacher model and the student model respectively, R topo is the topological regularization term, KL is the Kullback-Leibler divergence, and λ t is the weight coefficient.
[0142] Temporal knowledge distillation focuses on transferring time series patterns and dynamic behavior features, including hidden layer state distillation loss, temporal relationship loss, and periodicity preservation loss, and is a multi-scale temporal knowledge transfer mechanism.
[0143] Hidden layer state distillation loss:
[0144] L hidden = MSE(h t , h s ) + λ h ·KL(P(h t ) ∥ P(h s ))
[0145] Among them, h t and h sare the hidden layer states of the teacher model and the student model respectively, P(·) represents the state distribution, and MSE is the mean squared error, λ h is the weight coefficient;
[0146] To maintain the temporal dependence relationship, a temporal relationship loss is introduced, and the temporal relationship loss is:
[0147] L temporal = ∑|r t (i,j) - r s (i,j)|
[0148] where r t (i,j) and r s (i,j) represent the relationship strength between time steps i and j of the teacher model and the student model.
[0149] The periodicity preservation loss is:
[0150]
[0151] where T seq and S seq represent the temporal features of the teacher and student models, and FFT(·) represents the fast Fourier transform, which is used to capture the periodic features of the signal.
[0152] For the special requirements of industrial control network anomaly detection, a task-specific knowledge transfer mechanism is adopted. Task-specific knowledge distillation includes anomaly discrimination knowledge distillation loss and boundary awareness loss.
[0153] The anomaly discrimination knowledge distillation loss is:
[0154] L anomaly = CE(y t , y s ) + λ a ·KL(q t ∥q s )
[0155] where y t 、y s are the anomaly discrimination results, q t 、q s are the anomaly probability distributions, CE is the cross entropy, and λ a is the weight coefficient.
[0156] To improve the recognition ability of boundary samples, a boundary awareness loss is introduced:
[0157]
[0158] where w i is the sample weight, and higher weights are assigned to boundary samples, ft (.) and f s (.) respectively represent the specific task feature extraction functions of the teacher model and the student model, and f t (x i ) and f s (x i ) are the feature representations of the specific tasks extracted by the teacher model and the student model for the boundary sample x i .
[0159] Then the comprehensive loss function is:
[0160] L total = α s (L feature + L relation + L topo ) + α t (L hidden + L temporal + L periodic ) + α a (L anomaly + L boundary )
[0161] Where α s , α t , α a are the weight coefficients of each level of loss, which are dynamically adjusted through the validation set:
[0162] α m = softmax(v m · performance m )
[0163] performance m represents the effect metric of knowledge transfer at each level, and v m is the weight vector parameter, which is a learnable parameter used to adjust the importance of knowledge transfer at different levels.
[0164] Adopt the soft label mechanism with temperature adjustment to constrain the probability of abnormal distribution and ensure the stability of knowledge transfer:
[0165] q i = softmax(z i / T)
[0166] Where z i is the logits output, and q i is the probability distribution after temperature adjustment. Here, i is the i-th element of the input vector, and T is the temperature parameter used to control the smoothness of the soft label. The temperature parameter is dynamically adjusted according to the training progress:
[0167]
[0168] where T max is the initial temperature, γ is the decay rate, epoch is the current round, and total epochs is the total number of rounds.
[0169] The above three layers respectively use the soft label mechanism. P in the topology-preserving loss in spatial knowledge distillation t and P s are in the form of soft labels. P(h t ) and P(h s ) in the hidden state distillation in temporal knowledge distillation are the soft labels of the teacher and student models respectively. In task-specific knowledge distillation, q t and q s use the soft label mechanism, representing the abnormal probability distributions of the teacher and student models.
[0170] Through the above multi-level knowledge distillation, the present invention realizes the efficient transfer of the knowledge of the teacher model, enabling the lightweight student model to significantly reduce the computational complexity and storage requirements while maintaining a high detection performance. Experimental results show that the student model after knowledge distillation has a parameter reduction of more than 60%, a 3-fold increase in inference speed, and maintains a detection accuracy of more than 90%.
[0171] A lightweight student model is obtained after training.
[0172] Step 6: Anomaly detection
[0173] Input the real-time industrial control network traffic data into the lightweight student model for real-time anomaly detection of industrial control network traffic.
[0174] The anomaly detection process of the present invention adopts a multi-dimensional scoring mechanism, comprehensively considering the anomaly degrees of spatial features, temporal features, and task-specific features. The calculation formula for the total anomaly score Score(x) is:
[0175] Score(x) = β1S s (x) + β2S t (x) + β3S a (x)
[0176] where S s , S t and S a respectively represent the anomaly scores of spatial features, temporal features, and task-specific features, and β1, β2, β3 are adaptive weight coefficients. The adaptive weight coefficient β m is dynamically updated in the following way:
[0177]
[0178] m, i ∈ {s, t, a}, vs The importance score representing spatial features, v t The importance score representing temporal features, v a The importance score representing task-specific features, dynamically adjusted according to historical detection results.
[0179] The calculation of the anomaly scores for spatial feature representations, temporal features, and task-specific features all adopts a scoring method based on Mahalanobis distance:
[0180]
[0181] where m ∈ {s, t, a}, μ m and Σ m are the mean vector and covariance matrix of the corresponding features respectively.
[0182] The covariance matrix is maintained by incremental update to improve calculation efficiency:
[0183]
[0184] where η is the learning rate, is the updated covariance matrix;
[0185] The detection decision rule is:
[0186] If
[0187] Score(x) > θ dynamic
[0188] Return anomaly
[0189] Otherwise:
[0190] Return normal
[0191] The dynamic threshold θ dynmaic is updated in the following way:
[0192] θ dynamic = μ score + r · σ score
[0193] where μ score and σ score are the mean and standard deviation of historical scores respectively, and r is an adjustable sensitivity parameter that can control the looseness or strictness of the threshold and determines the sensitivity of anomaly detection. The larger the r value, the more conservative the detection, and the smaller the r value, the more sensitive the detection. It needs to be adjusted according to the specific application scenario.
[0194] The present invention adopts a sliding window mechanism to smooth the decision results and improve the stability of detection. An abnormal alarm is triggered only when the prediction results of consecutive Q samples all exceed the dynamic threshold. Q = 10 is applicable to relatively strict judgment criteria, Q = 15 is applicable to scenarios with high reliability requirements, and Q = 20 is applicable to scenarios with extremely high reliability requirements. For the selection of the Q value, the tolerance of the system to false alarms needs to be considered. Generally, the larger Q is, the lower the false alarm rate is, and the smaller Q is, the more timely the response is. The sampling frequency of the data also needs to be considered. For high-frequency sampling, a larger Q can be selected.
[0195] Meanwhile, an exponential decay factor is introduced to update the historical statistics:
[0196]
[0197] where α is the smoothing coefficient, used to control the influence degree of historical information, and represent the updated mean and standard deviation, and Score current is the total abnormal score of the current sample.
[0198] In addition, the present invention adopts an adaptive alarm suppression mechanism. By recording the time and type of the recently triggered alarm, setting the minimum warning interval according to the recorded alarm time, and merging the same alarm types according to the alarm type, repeated alarms are avoided, and the practicality of detection is improved.
[0199] Embodiment 2
[0200] This embodiment is an industrial control network anomaly detection method based on spatio-temporal feature fusion, which realizes model lightweight by combining knowledge distillation technology.
[0201] Step 1: Deploy distributed data collectors to collect data of various sensors (such as time, temperature, vibration, etc.) in the industrial control network. The collected industrial control network traffic data items include measurement values, timestamps, sensor IDs, etc. For critical equipment, a higher sampling frequency and stricter data quality control are adopted to ensure the real-time performance and reliability of the data.
[0202] Step 2: Preprocess the collected original data.
[0203] Z(x) = (x - μ rolling ) / σ rolling
[0204] Data deviating from the mean by more than 3 standard deviations is marked as potential anomalies. For missing data, interpolation methods are used to complete it according to the time series characteristics. At the same time, the numerical data is standardized to unify data with different dimensions into the same scale range.
[0205] Construct a topological relationship graph between devices, and depict the spatial association through position encoding
[0206] TPE(i,j) = f(D ij )·cos(ω k ·d ij )
[0207] where D ij is the topological distance and d ij is the physical distance. This encoding method ensures the preservation of the spatial structure information between devices during the data processing, providing important prior knowledge for subsequent feature extraction.
[0208] Step 3 combines spatial knowledge distillation, temporal knowledge distillation, and task-specific knowledge distillation to construct a comprehensive loss function,
[0209] L total = α s (L feature + L relation + L topo ) + α t (L hidden + L temporal + L periodic ) + α a (L anomaly + L boundary )
[0210] The weight coefficients are dynamically adjusted through the validation set:
[0211] α m = softmax(v m ·performance m )
[0212] Meanwhile, a temperature adjustment mechanism is used to constrain the soft labels
[0213] Soft label calculation:
[0214] q i = softmax(z i / T)
[0215] Temperature parameter dynamic adjustment:
[0216]
[0217] Step 4 deploys the optimized lightweight student model to implement the anomaly detection function. A multi-dimensional scoring mechanism is used to calculate the total anomaly score, which can comprehensively consider the anomaly degree of the data in different feature spaces:
[0218] Score(x) = β1S spatial (x) + β2S temporal (x) + β3S fusion (x)
[0219] Step 5 implements a dynamic threshold update mechanism.
[0220] An exponential decay factor is introduced to update the historical statistics:
[0221]
[0222] where α is the smoothing coefficient used to control the influence degree of historical information, and represent the updated mean and standard deviation, and Score current is the total anomaly score of the current sample;
[0223] The dynamic threshold is updated with the updated mean and standard deviation. This adaptive threshold can be dynamically adjusted according to the system state.
[0224] Step 6 establishes an alarm management mechanism. Multiple levels of alarm thresholds are set to classify and process the detected anomalies. At the same time, an alarm suppression strategy is implemented to avoid repeated alarms, and detailed alarm context information is recorded to support subsequent cause analysis and processing. The system will also automatically count and analyze alarm patterns to optimize the alarm strategy.
[0225] Step 7 implements an incremental learning mechanism. The system continuously collects new labeled data during operation and periodically updates the model parameters. After each update, the optimization effect is evaluated on the validation set to ensure that the model performance will not degrade due to the update. This continuous learning mechanism enables the system to adapt to the dynamic changes of the industrial control network.
[0226] Step 8 establishes a performance monitoring system. Continuously monitor the key performance indicators of the system, including detection accuracy, false alarm rate, computing resource occupancy, response time, etc. Set performance warning thresholds and notify the operation and maintenance personnel in a timely manner when the indicators are abnormal. At the same time, record detailed performance logs to support subsequent optimization analysis.
[0227] The parts not described in this invention are applicable to the prior art.
Claims
1. A lightweight anomaly detection method for industrial control network spatiotemporal features based on knowledge distillation, characterized in that: The method includes the following: Obtain industrial control network traffic data and perform preprocessing to obtain preprocessed network traffic features, including preprocessed frequency domain features F freq , statistical features after preprocessing F stat And the preprocessed time series features F temp ; Construct a lightweight anomaly detection model based on spatiotemporal features, including spatial feature extraction module, temporal feature extraction module, fusion module, knowledge distillation module, and anomaly detection module; The time series feature extraction module is used to extract the preprocessed time series features F temp And the statistical characteristics F stat Processing is performed to obtain the time series feature representation F t ; The spatial feature extraction module is used to extract the preprocessed frequency domain features F freq And the statistical characteristics F stat Processing is performed to obtain the spatial feature representation F s ; The fusion module is used to represent the time series feature F t and spatial feature representation F s Fusion is performed to obtain the normalized fusion feature F norm ; The knowledge distillation module includes a teacher model and a student model. The teacher model is based on the normalized fusion feature F norm As input, the student model is trained. During training, three levels of loss constraints are set: spatial knowledge distillation, temporal knowledge distillation, and task-specific knowledge distillation. The comprehensive loss function L total for; L total =a s (L feature +L relation +L topo )+a t (L hidden +K temporal +L periodic )+a a (L anomaly +L boundary ) Among them, α s , α t , α a is the weight coefficient of each level loss, and the three are expressed by α m Unified representation, m∈{s,t,a}, through the validation set according to α m =softmax(v m ·performance m ) Dynamic adjustment, performance m Represents the effect measurement of knowledge transfer at each level, v m is a weight vector parameter used to adjust the importance of knowledge transfer at different levels; L feature is the distillation loss at the characteristic level; L relation is the relation consistency loss; L topo is the topology preservation loss; L hidden is the hidden layer state distillation loss; L temporal is the temporal relationship loss; L periodic To maintain the loss periodically; L anomaly is the knowledge distillation loss for abnormality discrimination; L boundary For boundary perception loss; The trained student model is used as a lightweight student model of the anomaly detection module, and the lightweight student model is used to perform anomaly detection on real-time industrial control network traffic data.
2. The detection method according to claim 1, characterized in that: Spatial knowledge distillation mainly focuses on the migration of interaction relationships and topological structure information between devices, including feature-level distillation loss, relationship consistency loss, and topology preservation loss. The feature-level distillation loss is: Among them, T s and S s Represent the spatial feature representations of the teacher model and the student model respectively; The relation consistency loss is: L relation =∥G t (T s )-G s (S s )∥1 Among them, G t (·) and G s (·) represents the construction function of the feature relationship graph between the teacher model and the student model, which is obtained by calculating the similarity matrix between the feature vectors; The topology preservation loss is: L topo =KL(P t ∥P s )+λ t ·R topo (S s ) Among them, P t and P s are the topological prediction distributions of the teacher model and the student model, R topo is the topological regularization term, KL is the Kullback-Leibler divergence, λ t is the weight coefficient; Time series knowledge distillation focuses on transferring time series patterns and dynamic behavior characteristics, including hidden state distillation loss, time series relationship loss, and periodicity retention loss. The hidden state distillation loss is: L hidden =MSE(h t ,h s )+λ h ·KL(P(h t )∥P(h s )) Among them, h t and h s are the hidden states of the teacher model and the student model, P(·) represents the state distribution, MSE is the mean square error, and λ h is the weight coefficient; The temporal relationship loss is: L temporal =∑|r t (i,j)-r s (i,j)| Among them, r t (i,j) and r s (i, j) represents the strength of the relationship between the teacher model and the student model at time steps i and j; The periodic holding loss is: Among them, S seq and S seq represents the time series characteristics of the teacher and student models, FFT(·) represents the fast Fourier transform, which is used to capture the periodic characteristics of the signal; Task-specific knowledge distillation includes anomaly discrimination knowledge distillation loss and boundary-aware loss. The abnormal discrimination knowledge distillation loss is: L anomaly =CE(y t ,y s )+λ a ·KL(q t ∥q s ) Among them, y t ,y s is the abnormality discrimination result, q t ,q s is the abnormal probability distribution, CE is the cross entropy, λ a is the weight coefficient; The boundary-aware loss is: Among them, w i is the sample weight, giving higher weight to boundary samples, f t (.) and f s (.) represent the task-specific feature extraction functions of the teacher model and the student model, respectively, and f t (x i ) and f s (x i ) is the teacher model and the student model for the boundary sample x i Extracted task-specific feature representations.
3. The detection method according to claim 1, characterized in that: The preprocessing includes data cleaning, standardization and preliminary feature extraction. The improved Z-score is used for preliminary cleaning. The specific process is: Z(x)=(x-μ rolling ) / s rolling Among them, x is the original data, μ rolling and σ rolling is the mean and standard deviation in the sliding window, Z(x) is the data value after cleaning, and the window size is dynamically adjusted according to the time series characteristics of the data: w=min(w max ,max(w min ,T cycle )) Among them, w is the size of the final window, w max is the maximum allowed window size, w min is the minimum allowed window size, T cycle is the estimated value of the period of the data, obtained through autocorrelation analysis; After cleaning, normalize the different data types: For continuous data, use the modified MinMax normalization: x’=(x-x min ) / (x max -x min +ε) Among them, ε is the smoothing factor, which is used to deal with the influence of extreme values, x is the original data, and x min 、x max are the minimum and maximum values in the data set, and x' is the standardized data value; For periodic data, phase normalization is used; for discrete data, one-hot encoding conversion is used; the normalization parameters are updated online; After standardization, preliminary feature extraction is performed: statistical features include data value mean, data value variance, skewness, and closedness; time series features include data difference, sliding window mean and standard deviation; frequency domain features take the top-k frequency components; The preprocessed statistical features, time series features, and frequency domain features are obtained by screening based on feature importance scores.
4. The detection method according to claim 1, characterized in that: The spatial feature extraction module adopts Transformer architecture, CNN architecture or hybrid architecture.
5. The detection method according to claim 1, characterized in that: The temporal feature extraction module adopts Bi-LSTM architecture, LSTM architecture, GRU architecture, and RNN architecture.
6. The detection method according to claim 1, characterized in that: After real-time industrial control network traffic data is preprocessed, it is processed by a lightweight student model to obtain spatial features, temporal features, and specific task features. The anomaly detection module adopts a multi-dimensional scoring mechanism to comprehensively consider the degree of anomaly of spatial features, temporal features, and specific task features. The calculation formula of the total anomaly score Score(x) is: Score(x)=β1S s (x)+β2S t (x)+β3S a (x) Among them, S s , S t and S a Represent the abnormal scores of spatial feature representation, temporal feature and specific task feature respectively; β1, β2, β3 are adaptive weight coefficients. Dynamic update, m, i∈{s,t,a},v s represents the importance score of the spatial feature, v t Represents the importance score of the time series feature, v a The importance score representing the feature of a specific task is dynamically adjusted based on the historical detection results; The calculation of the anomaly scores of spatial feature representation, temporal features, and specific task features all adopts the scoring method based on Mahalanobis distance. Among them, m∈{s,t,a}, μ m and Σ m are the mean vector and covariance matrix of the corresponding features respectively, and x represents the input; Covariance matrix Σ m Maintenance through incremental updates: Where η is the learning rate, is the updated covariance matrix; The detection decision rule is: if Score(x)>θ dynamic Return Exception otherwise: Return to Normal Dynamic threshold θ dynamic Updated via: θd ynamic =μ score +r·s score Among them, μ score and σ score are the mean and standard deviation of the historical scores respectively, and r is an adjustable sensitivity parameter.
7. The detection method according to claim 5, characterized in that: The sliding window mechanism is used to smooth the decision results. The abnormal alarm is triggered only when Q consecutive samples exceed the dynamic threshold. At the same time, an exponential decay factor is introduced to update the historical statistics: Among them, α is the smoothing coefficient, which is used to control the influence of historical information. and Represents the updated mean and standard deviation, Score current is the total anomaly score of the current sample; Update the dynamic threshold with the updated mean and standard deviation.
8. The detection method according to claim 5, characterized in that: Avoid duplicate alarms by recording the time and type of the most recently triggered alarm, setting the minimum warning interval based on the recorded alarm time, and merging the same alarm type based on the alarm type.
Citation Information
Cited By
Smelting control method, device and equipment based on TinyML and medium
CN120469320A
TinyML-based smelting control method, device, equipment and medium
CN120469320B
Network traffic anomaly detection method based on aggregation type mimicry distillation
CN121239497A
A network traffic anomaly detection method based on polymeric quasispecies distillation
CN121239497B
Operation and maintenance anomaly detection method and system based on core set selection
CN121436976A