Anomaly detection method and device based on adaptive multi-scale feature modeling
By adopting anomaly detection method of adaptive multi-scale feature modeling in industrial control systems, combined with Fourier transform, gated network, multi-scale expert network and graph convolution technology, the existing methods have insufficient ability to capture dynamic characteristics and time scale correlations, and achieve more efficient and accurate anomaly detection.
Patent Information
- Application Number
- CN202510412968.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-03
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2045-04-03
AI Technical Summary
The anomaly detection method of existing industrial control systems is insufficiently considered in the dynamic characteristics of the model, the time scale correlation and the context-dependent capture capabilities, making it difficult to cope with the dynamic changing industrial environment.
Anomaly detection method based on adaptive multi-scale feature modeling is adopted, and time component decomposition is performed through Fourier transform and sliding window. Combined with gated network and multi-scale expert network, features of different scales are adaptively weighted, spatial correlation features are captured using graph convolution, and potential spatial representation is learned through variational autoencoder, data reconstruction and model training are carried out.
It improves the accuracy of abnormal detection in industrial control systems, can dynamically process data characteristics on different time scales, adapt to complex industrial environments, and provides more flexible and accurate abnormal detection tools.
Smart Images

Figure CN119939476B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anomaly detection of time series data in industrial control systems, and in particular to an anomaly detection method and device based on adaptive multi-scale feature modeling. Background Art
[0002] Industrial control systems (ICS) are widely used in energy, manufacturing, transportation, water treatment and other fields to monitor and control various aspects of industrial processes. These systems usually rely on a large number of sensors, actuators and control devices. They ensure the normal operation of equipment by collecting relevant data and performing anomaly detection analysis, which is an important basis for complex system analysis and monitoring. Anomaly detection aims to identify behaviors that deviate from normal patterns. Such deviations usually indicate that the system may have potential failures, risks or abnormal events.
[0003] Traditional anomaly detection methods mostly rely on manually set thresholds or rules. Due to the diversity and complexity of industrial processes, traditional methods are limited by expert experience and model expression capabilities, and lack sufficient adaptability. Deep learning technology is widely used in the field of anomaly detection and is the latest direction of anomaly detection. However, in multi-dimensional time series anomaly detection, sequences on different time scales show different changes and fluctuations. Existing deep learning anomaly detection technology still has limitations. It does not consider the dynamic characteristics when modeling dependencies in different time ranges, ignores the flexibility of multi-scale feature extraction and fusion, fails to fully utilize the multi-level spatiotemporal information in the data, does not consider the dynamic characteristics when modeling dependencies in different time ranges, and lacks the ability to capture potential correlations and contextual dependencies on time scales, making it difficult to cope with dynamically changing industrial environments.
[0004] In view of this, the applicant filed this application after studying the existing technology. Summary of the invention
[0005] The present invention aims to provide an anomaly detection method and device based on adaptive multi-scale feature modeling to address the shortcomings of existing methods, such as insufficient consideration of model dynamic characteristics, insufficient ability to capture potential correlations and context dependencies on a time scale, and difficulty in coping with dynamically changing industrial environments.
[0006] In order to solve the above technical problems, the present invention is implemented through the following technical solutions:
[0007] An anomaly detection method based on adaptive multi-scale feature modeling, comprising:
[0008] S1, obtain the historical monitoring time series data of the industrial control system and preprocess the data to obtain the training data for the input anomaly detection model;
[0009] S2, combining Fourier transform with sliding windows of different sizes, decomposing the input training data by time component and fusing it with the training data to obtain a decomposition result;
[0010] S3, according to the decomposition result, a gating weight matrix is calculated in combination with a gating network, and the output of processing the training data by a plurality of expert networks of different scales is adaptively weighted and combined to obtain a multi-scale fusion feature;
[0011] S4, calculating the adjacency matrix of the multi-scale fusion feature at each time step, and aggregating the node features of the adjacency matrix using a graph convolution operation to capture the spatial correlation features between nodes to obtain spatial features;
[0012] S5, performing weighted summation of the spatial feature and the multi-scale fusion feature to obtain a final feature representation combining spatiotemporal information;
[0013] S6, sending the final feature representation to a variational autoencoder to learn a latent space representation, and performing data decoding and reconstruction to obtain reconstructed data;
[0014] S7, combining the reconstructed data error loss and KL divergence as the loss function to train the model and obtain a trained anomaly detection model;
[0015] S8, input the preprocessed data to be detected into the anomaly detection model, and combine it with the set threshold to obtain the anomaly detection result.
[0016] Preferably, the preprocessing includes processing missing values and duplicate values of the data.
[0017] Preferably, the input training data is decomposed into time components, specifically:
[0018] The frequency characteristics of the input data are analyzed by Fourier transform technology to capture periodic changes and obtain seasonal components; wherein the seasonal components are regular fluctuation data in the data that changes with a fixed period;
[0019] A moving average of the input data is calculated using multiple sliding windows of different sizes, and the overall upward or downward trend of the input data is captured by smoothing short-term fluctuations to obtain a trend component; wherein the trend component is the trend data of long-term changes in the data, including continuously growing data, continuously declining data, or unchanged data;
[0020] The seasonal component and the trend component are subtracted from the input data to obtain a residual component; the residual component includes random fluctuation data or noise data.
[0021] Preferably, S3 is specifically:
[0022] According to the decomposition results , extract the feature vector of each time step, and obtain the data decomposition result subset; generate the initial gating weight matrix according to the data decomposition result subset, the formula is:
[0023] ;
[0024] in, The decomposition result is represented by The gating weight matrix of the subsets; is the activation function; and are the learned weights and bias terms respectively; Represents the decomposition result No. The input features of the subset;
[0025] According to the training data, the feature vector of each time step is extracted to obtain a data subset; according to the gated weight matrix, the data subset is allocated to a plurality of expert networks of different scales to select the most suitable expert network to process the current data and obtain the output feature , the expression is:
[0026] ;
[0027] ;
[0028] in, Indicates the current data entered The corresponding gating weight matrix; Indicates from Select the output with the largest weight experts; represents the expert weight representation after noise injection; Indicates the number of selected experts; represents the gating weight; represents noise sampled from a standard normal distribution with mean 0 and variance 1; Represents a smooth ReLU function to enhance the nonlinear expression ability of the model; represents the noise weight;
[0029] Based on the gated weight matrix, the features of the outputs of the expert networks at different scales are weighted and summed to obtain the multi-scale fusion features. , the expression is:
[0030] ;
[0031] in, It is Output features of the expert network; Represents the total number of expert networks.
[0032] Preferably, expert networks of different scales define segments of different sizes, thereby extracting different scale features of the input data. When each expert network processes the data assigned to it, the processing steps are as follows:
[0033] Divide the allocated data into multiple time segments ;
[0034] By setting different time slice sizes , define the processing resolution of time series data for each expert network, each time segment represents a local time window, and use the local window to learn the time dependency of the data, then:
[0035] ;
[0036] in, represents the i-th time segment; Represents the time series data in the i-th time segment;
[0037] Each expert network extracts local features of each time segment through a deep grouped convolutional layer to obtain the short-term dependency features of continuous time steps in each time segment in the time series. , the expression is:
[0038] ;
[0039] in, Represents the convolution feature extraction operation;
[0040] Using multi-head self-attention to learn different time segments long-term dependencies between them, building interactions between different time segments and generating global features ; When calculating multi-head self-attention, the query matrix, key matrix and value matrix transformation are calculated, and the features of each time segment are aggregated through the attention weights;
[0041] Concatenate the features learned by each expert network, combining the assigned data and short-term dependency features and global features across time segments , get the fusion features output by each expert network .
[0042] Preferably, the S4 is specifically:
[0043] Calculate the cosine similarity between all node features of the multi-scale fusion feature in each time step, and for each time step Each pair of nodes and , calculate the cosine similarity, the formula is:
[0044] ;
[0045] in, Represents the feature vector and The cosine similarity between ; , Respectively represent nodes and nodes The eigenvector of represents the L2 norm;
[0046] If the similarity between two nodes is greater than the preset threshold , then establish edge connections and obtain the adjacency matrix, the expression is:
[0047] ;
[0048] in, Indicates the time step The corresponding adjacency matrix;
[0049] Based on the adjacency matrix, the GraphSAGE convolutional layer is used to aggregate the features of the nodes to capture the spatial correlation features between the nodes, namely the spatial features , the formula of the GraphSAGE convolutional layer is:
[0050] ;
[0051] in, Representation Node In the Feature representation of the layer; ) represents a node The neighbor node set of ; Aggregate is the aggregation operation; For the The weight matrix of the layer; is the activation function; Representation Node Neighbor nodes The feature representation of .
[0052] Preferably, the S6 is specifically:
[0053] The final feature representation is input into the variational autoencoder, and the latent space mean of the final feature representation is obtained through forward propagation learning of the fully connected network of the variational autoencoder. μ With log variance ; and use the reparameterization technique to sample the latent variables from the latent space , the expression is:
[0054] ;
[0055] in, It follows the standard normal distribution The noise sampled in represents the standard deviation of the latent space;
[0056] Through a fully connected layer, the latent variables Map back to the input space, reconstruct the input data to get the reconstructed data , to learn a compressed representation of the input data.
[0057] Preferably, the reconstructed data error loss is: the reconstructed data is measured by the mean square error The difference between the training data X input to the model is expressed as:
[0058] ;
[0059] in, represents the reconstructed data error loss;
[0060] The KL divergence It measures the difference between the distribution of the latent space and the standard normal distribution, and the expression is:
[0061] ;
[0062] in, represents the KL divergence; a total number of data subsets representing the training data; Indicates The standard deviation of a subset of data; Indicates The latent space mean of a subset of data;
[0063] The loss function is the weighted sum of the reconstructed data error loss and the KL divergence, and the formula is:
[0064] ;
[0065] in, represents the weight of the KL divergence.
[0066] Preferably, the method further includes testing and evaluating the anomaly detection model. When the test evaluation result meets the set requirements, the anomaly detection model is evaluated as qualified and used for subsequent anomaly detection; otherwise, the model is retrained; wherein the test evaluation result of the anomaly detection model is calculated by the following formula:
[0067] Accuracy The calculation formula is: ;
[0068] Accuracy The calculation formula is: ;
[0069] Recall The calculation formula is: ;
[0070] The calculation formula is: ;
[0071] Among them, TP represents the number of true positives, that is, the number of samples where both the actual situation and the test results are normal; FP represents the number of false positives, that is, the number of samples where the actual situation is abnormal and the test results are normal; TN represents the number of true negatives, that is, the number of samples where both the actual situation and the test situation are abnormal; FN represents the number of false negatives, that is, the number of samples where the actual situation is normal and the test results are abnormal.
[0072] The present invention also provides an anomaly detection device based on adaptive multi-scale feature modeling, comprising:
[0073] A data preprocessing unit is used to obtain the historical monitoring time series data of the industrial control system and preprocess the data to obtain the training data for the input anomaly detection model;
[0074] A component decomposition unit, used to combine Fourier transform with sliding windows of different sizes, decompose the input training data into time components, and then fuse them with the training data to obtain a decomposition result;
[0075] A multi-scale fusion unit, used to calculate a gating weight matrix based on the decomposition result in combination with a gating network, and adaptively weightedly combine the outputs of a plurality of expert networks of different scales processing the training data to obtain a multi-scale fusion feature;
[0076] A graph convolution unit is used to calculate the adjacency matrix of the multi-scale fusion feature at each time step, and aggregate the node features of the adjacency matrix using a graph convolution operation to capture the spatial correlation features between nodes and obtain spatial features;
[0077] A spatiotemporal combining unit, used for performing weighted summation of the spatial feature and the multi-scale fusion feature to obtain a final feature representation combining spatiotemporal information;
[0078] A data reconstruction unit, used for sending the final feature representation into a variational autoencoder to learn a latent space representation, and performing data decoding and reconstruction to obtain reconstructed data;
[0079] The model training unit is used to train the model by combining the reconstructed data error loss and KL divergence as the loss function to obtain a trained anomaly detection model;
[0080] The anomaly detection unit is used to input the preprocessed data to be detected into the anomaly detection model and obtain the anomaly detection result in combination with the set threshold.
[0081] The present invention also provides an anomaly detection device based on adaptive multi-scale feature modeling, comprising a processor and a memory, wherein a computer program is stored in the memory, and the computer program can be executed by the processor to implement an anomaly detection method based on adaptive multi-scale feature modeling as described above.
[0082] The present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, an anomaly detection method based on adaptive multi-scale feature modeling as described above is implemented.
[0083] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0084] The present invention combines a variety of advanced time series data analysis technologies such as time decomposition, hybrid multi-scale expert networks, and spatiotemporal graph convolution. While paying attention to the correlation between variables, it adaptively combines expert networks of different scales to dynamically process data features at different time scales, thereby making up for the shortcomings of existing methods in processing the dynamic characteristics of time scale changes.
[0085] Compared with traditional anomaly detection methods, this invention not only focuses on the spatial correlation in the multi-dimensional time series data of the industrial control system, but also takes into account the dynamic characteristics of the features over time. This invention improves the accuracy of anomaly detection in industrial control systems, provides a more flexible and accurate anomaly detection tool for practical applications, and adapts to the complex dynamic changes of time series data in industrial control systems without label data. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for use in the embodiments will be briefly introduced below. It should be understood that the following drawings only show certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.
[0087] Figure 1 A schematic diagram of an anomaly detection method based on adaptive multi-scale feature modeling provided in Example 1.
[0088] Figure 2 A flowchart of an anomaly detection method based on adaptive multi-scale feature modeling is provided in Example 1.
[0089] Figure 3 A schematic diagram of an anomaly detection device based on adaptive multi-scale feature modeling provided in Example 2.
[0090] The present invention is further described in detail below in conjunction with the accompanying drawings and specific embodiments. DETAILED DESCRIPTION
[0091] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the invention claimed for protection, but merely represents the selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0092] Embodiment 1
[0093] Embodiment 1 of the present invention provides an anomaly detection method based on adaptive multi-scale feature modeling, which can be implemented by an anomaly detection device based on adaptive multi-scale feature modeling (hereinafter referred to as an anomaly detection device), and in particular, executed by one or more processors in the anomaly detection device.
[0094] In this embodiment, the anomaly detection device may be an electronic device equipped with a processor, the processor having a computer program of the anomaly detection method based on adaptive multi-scale feature modeling and the computer program can be executed, such as a computer, a smart phone, a smart tablet, a workstation, etc., which is not limited here.
[0095] like Figure 1-Figure 2 As shown, an anomaly detection method based on adaptive multi-scale feature modeling includes steps S1 to S8.
[0096] S1, obtains the historical monitoring time series data of the industrial control system and preprocesses the data to obtain the training data for the input anomaly detection model.
[0097] Specifically, historical feature data of industrial control equipment is obtained to check whether the data is complete. Feature data is a collection of observations or data points associated with timestamps. Preprocessing operations include processing missing values and duplicate values of the data.
[0098] In this embodiment, the preprocessed data is input into an anomaly detection model based on adaptive multi-scale feature extraction and reconstruction to obtain a label indicating whether the industrial control equipment is in an abnormal state at the current moment, thereby obtaining an anomaly detection result.
[0099] Specifically, the preprocessed feature data Divided into training set and test set, the training set is used for model training, and the test set is used for model testing and evaluation.
[0100] definition: ;
[0101] in, Represents the maximum value of the timestamp; represents the characteristic data at time t, Represents the number of features, and the output is the anomaly detection result label , , Represents the label value of the anomaly detection result at time t.
[0102] S2, combining Fourier transform with sliding windows of different sizes, decomposing the input training data into time components and then fusing them with the training data to obtain a decomposition result.
[0103] The specific steps include S21 to S24.
[0104] S21. First, the frequency characteristics of the input data are analyzed using Fourier transform technology to capture periodic changes and obtain seasonal components; wherein the seasonal components are regular fluctuation data in the data that changes with a fixed period.
[0105] For example, the seasonal component is reflected in the fact that summer and winter are data peak periods each year, and spring and autumn are relatively low peak periods.
[0106] S22, using multiple sliding windows of different sizes to calculate the moving average of the input data, capturing the overall upward or downward trend of the input data by smoothing short-term fluctuations, and obtaining a trend component; wherein the trend component is the trend data of long-term changes in the data, including continuously growing data, continuously declining data, or unchanged data.
[0107] For example, the data shows an upward trend year by year.
[0108] S23, subtracting seasonal components and trend components from the input data to obtain residual components; the residual components are the parts of the data that cannot be explained by seasonality and trend, including random fluctuation data or noise data.
[0109] For example, data fluctuations caused by unexpected events (such as extreme weather).
[0110] S24, thereby decomposing the input data into seasonal components, trend components and residual components. Then, these components are combined with the input data to obtain the decomposition results The time decomposition method can more clearly understand the structure and changing patterns of the data, which can be used for subsequent expert selection and feature extraction, providing a scientific basis for decision-making.
[0111] S3, according to the decomposition result, a gating weight matrix is calculated in combination with a gating network, and the outputs of the training data processed by a plurality of expert networks of different scales are adaptively weighted and combined to obtain multi-scale fusion features.
[0112] The specific steps include S31 to S34.
[0113] S31, first, according to the decomposition result , extract the feature vector of each time step, and obtain the data decomposition result subset; generate the initial gating weight matrix according to the data decomposition result subset, the formula is:
[0114] ;
[0115] in, The decomposition result is The gating weight matrix of the subsets; is the activation function; and are the learned weights and bias terms respectively; Represents the decomposition result No. The input features of the subset.
[0116] S32, extracting the feature vector of each time step according to the training data to obtain a data subset; and assigning the data subset to a plurality of expert networks of different scales according to the gating weight matrix to select the most suitable expert network to process the current data.
[0117] In this embodiment, "Expert Network" refers to a specific neural network architecture or module, which is designed to be responsible for feature extraction of a subtask or a specific field in a complex task. This network can be regarded as an "expert" because it focuses on processing a specific type of data or features, thereby playing a greater role in the overall task. In the scenario of time series analysis, the task of the expert network is usually to process specific time segments or feature patterns.
[0118] Specifically, when processing a specific data subset, each expert network is assigned different weights according to its contribution to the task. The gated weight matrix determines which experts will be activated and their activation levels based on the weights, thereby determining how the data subsets are allocated to each expert for processing. The data selection process for experts is adaptive, and is dynamically adjusted based on the data features of each time step to ensure that in different time periods, the model can select the most appropriate expert network based on the characteristics of the data. The expression is:
[0119] ;
[0120] ;
[0121] in, Indicates the current data entered The corresponding gating weight matrix; Indicates from Select the output with the largest weight experts; represents the expert weight representation after noise injection; Indicates the number of selected experts; represents the gating weight; represents noise sampled from a standard normal distribution with mean 0 and variance 1; Represents a smooth ReLU function to enhance the nonlinear expression ability of the model; represents the noise weight;
[0122] S33, expert networks of different scales define time segments of different sizes, and then define the time resolution of the input data, so as to extract features of different scales of the data. Each expert network extracts local features through deep convolution within the divided time segments, uses the attention mechanism across time segments, models the relationship between different time segments, and improves the long-range dependency modeling capability of time series. The processing steps are as follows:
[0123] Divide the allocated data into multiple time segments ;
[0124] By setting different time slice sizes , define the processing resolution of time series data for each expert network, each time segment represents a local time window, and use the local window to learn the time dependency of the data, then:
[0125] ;
[0126] in, represents the i-th time segment; Represents the time series data in the i-th time segment; i represents the variable;
[0127] Each expert network extracts local features of each time segment through a deep grouped convolutional layer to obtain the short-term dependency features of continuous time steps in each time segment in the time series. , the expression is:
[0128] ;
[0129] in, Represents a convolutional feature extraction operation.
[0130] In this embodiment, deep grouped convolution is a special convolution operation that can divide the input channels into several groups and then perform convolution operations on each group separately. The advantage of this is that the amount of calculation can be greatly reduced while maintaining the effective extraction of features for each group. In this embodiment, the data of each time segment can be regarded as a multi-dimensional time series, and grouped convolution allows the network to process this data in parallel, thereby extracting features within each time segment.
[0131] Then, multi-head self-attention is used to learn different time segments long-term dependencies between them, building interactions between different time segments and generating global features ; When calculating multi-head self-attention, the matrix transformation of the query matrix, key matrix and value is calculated, and the features of each time segment are aggregated through the attention weight. Multi-head self-attention The expression is:
[0132] ;
[0133] in, is the query matrix, is the key matrix, is the value matrix, is the dimension of the key matrix, is the normalization function.
[0134] Concatenate the features learned by each expert network, combining the assigned data and short-term dependency features and global features across time segments , get the fusion features output by each expert network .
[0135] S34, based on the gated weight matrix, processing the features output by the expert networks of multiple different scales Perform weighted summation to obtain multi-scale fusion features , the expression is:
[0136]
[0137] in, It is Output features of the expert network; Represents the total number of expert networks.
[0138] S4, calculating the adjacency matrix of the multi-scale fusion feature at each time step, and aggregating the node features of the adjacency matrix using a graph convolution operation to capture the spatial correlation features between nodes and obtain spatial features.
[0139] The specific steps include S41 to S43.
[0140] S41, calculating the cosine similarity between all node features of the multi-scale fusion feature in each time step, for each time step Each pair of nodes and , calculate the cosine similarity, the formula is:
[0141] ;
[0142] in, Represents the feature vector and The cosine similarity between ; , Respectively represent nodes and nodes In time step The eigenvector of represents the L2 norm;
[0143] S42, if the similarity between the two nodes is greater than the preset threshold , then establish edge connections and obtain the adjacency matrix, the expression is:
[0144] ;
[0145] in, Indicates the time step The corresponding adjacency matrix.
[0146] S43, based on the adjacency matrix, uses the GraphSAGE (Graph Sample and Aggregate) convolutional layer to aggregate the node features to capture the spatial correlation features between nodes, namely the spatial features , the formula of the GraphSAGE convolutional layer is:
[0147] ;
[0148] in, Representation Node In the Feature representation of the layer; ) represents a node The neighbor node set of ; Aggregate is the aggregation operation; For the The weight matrix of the layer; is the activation function; Representation Node Neighbor nodes The feature representation of .
[0149] S5, performing weighted summation of the spatial feature and the multi-scale fusion feature to obtain a final feature representation combining spatiotemporal information.
[0150] In this step, the spatial features extracted by graph convolution are Multi-scale fusion features extracted by expert network Perform weighted summation to obtain the final feature representation that combines time and space information .
[0151] S6, sending the final feature representation to the variational autoencoder to learn the latent space representation, and performing data decoding and reconstruction to obtain reconstructed data.
[0152] The specific steps include S61 to S62.
[0153] S61, in the variational autoencoder part, the final feature representation is input into the variational autoencoder, and the latent space mean of the final feature representation is obtained through forward propagation learning of the fully connected network of the variational autoencoder. μ With log variance ; and use the reparameterization technique to sample the latent variables from the latent space , the expression is:
[0154] ;
[0155] in, is a latent variable, It follows the standard normal distribution The noise sampled in , represents the standard deviation of the latent space.
[0156] In this embodiment, the encoder of the variational autoencoder is composed of a fully connected network (FC), which maps the input features to the distribution parameters (mean and variance) of the latent space. The fully connected network of the encoder is usually composed of a multi-layer neural network, and each layer performs feature conversion through linear transformation and nonlinear activation function (such as ReLU).
[0157] Sampling directly from the latent space will result in the inability to backpropagate the gradient, so the reparameterization technique is needed. The core idea of the reparameterization technique is to separate the randomness from the parameters of the network so that the gradient can be backpropagated through the sampling process.
[0158] S62, in the decoder part, the latent variables are transformed through a fully connected layer Map back to the input space to reconstruct the input data and obtain the reconstructed data , to learn the compressed representation of the input data, the expression is:
[0159] ;
[0160] in, Represents a reconstruction compression operation.
[0161] In this embodiment, the decoder is also a fully connected network, which is used to map the latent variables back to the original feature space to generate reconstructed features.
[0162] S7, combining the reconstructed data error loss and KL divergence as the loss function to train the model and obtain the trained anomaly detection model.
[0163] In this embodiment, the loss function Combining the reconstruction data error loss and KL divergence, the formula is:
[0164] ;
[0165] in, represents the weight of the KL divergence, is the reconstruction data error loss, is the KL divergence.
[0166] The reconstructed data error loss is: the mean square error is used to measure the reconstructed data The difference between the training data X input to the model is expressed as:
[0167] ;
[0168] in, represents the reconstructed data error loss;
[0169] The KL divergence It measures the difference between the distribution of the latent space and the standard normal distribution, and the expression is:
[0170] ;
[0171] in, a total number of data subsets representing the training data; Indicates The standard deviation of a subset of data; Indicates The latent space mean of a subset of data.
[0172] S8, input the preprocessed data to be detected into the anomaly detection model, and combine it with the set threshold to obtain the anomaly detection result.
[0173] Specifically, after the model is trained, the trained anomaly detection model is tested and evaluated. When the test evaluation result meets the set requirements, the anomaly detection model is evaluated as qualified and used for subsequent anomaly detection; otherwise, the model is retrained; wherein the test evaluation result of the anomaly detection model is calculated by the following formula:
[0174] Accuracy The calculation formula is: ;
[0175] Accuracy The calculation formula is: ;
[0176] Recall The calculation formula is: ;
[0177] The calculation formula is: ;
[0178] Among them, TP represents the number of true positives, that is, the number of samples where both the actual situation and the test results are normal; FP represents the number of false positives, that is, the number of samples where the actual situation is abnormal and the test results are normal; TN represents the number of true negatives, that is, the number of samples where both the actual situation and the test situation are abnormal; FN represents the number of false negatives, that is, the number of samples where the actual situation is normal and the test results are abnormal.
[0179] Specifically, the training data is input into the anomaly detection model to obtain the training sample anomaly score. In this embodiment, the mean of the training sample anomaly score plus three times the standard deviation is used as the threshold. The test data set is input into the anomaly detection model, and the result obtained by weighting the reconstruction error loss and KL divergence in the total loss function is used as the anomaly detection score corresponding to each test sample in the test set. The anomaly score obtained in the test sample is compared with the threshold to obtain the anomaly category label corresponding to each test sample in the test set.
[0180] When the test evaluation result reaches the set threshold, the anomaly detection model is evaluated as qualified and used for subsequent unsupervised anomaly detection. The monitoring time series data of industrial control equipment is obtained in real time, and input into the anomaly detection model after preprocessing to obtain the anomaly detection result at the current moment.
[0181] In another preferred embodiment, in order to verify the effectiveness of the anomaly detection model and model solution proposed in the present invention, the SWaT data set used for industrial control system security research is selected as the research object. The data set has 51 dimensions and contains 11 days of operation data, of which the first 7 days are normal operation data, and the last 4 days receive typical network attack data of 36 attack scenarios. The data set contains data from 51 sensors and actuators, covering all aspects of the water treatment process, such as water level, flow, pressure, pH value, etc. The data is recorded in the form of a time series, recorded every 1 second, including a timestamp and the measurement value or status of each sensor / actuator.
[0182] The data of the first 7 days in the dataset are used as the training set, all of which are normal data, and the data of the last 4 days are used as the test set, which includes normal data and abnormal data. Then the model is trained, and the relevant parameter settings are shown in Table 1.
[0183] Table 1. Experimental parameter settings
[0184]
[0185] The experimental results are shown in Table 2.
[0186] Table 2. Experimental results
[0187]
[0188] It can be seen from Table 2 that the results of the detection using the anomaly detection model of the present invention are relatively high, and the characteristics are that it has obvious advantages in accuracy and precision, and can effectively improve the accuracy of anomaly detection in industrial control systems.
[0189] In summary, compared with the prior art, the present invention has the following beneficial effects:
[0190] The present invention combines multiple advanced time series data analysis technologies such as time decomposition, hybrid multi-scale expert networks, and spatiotemporal graph convolution. While focusing on the correlation between variables, it adaptively combines expert networks of different scales to dynamically process data features at different time scales, making up for the shortcomings of existing methods in processing dynamic characteristics of time scale changes. The method of the present invention takes into account the dynamic change characteristics of data at the time scale and can accurately model the anomaly detection task in industrial control systems.
[0191] Specifically, the present invention combines time decomposition with a gating network to assign appropriate expert network models to data at different time scales, thereby comprehensively modeling the temporal characteristics of data at multiple scales. In the expert layer, a deep convolutional network and an attention mechanism across time segments are used to mine the relationships within and between different time periods in the time series. Then, combined with the graph structure, the spatial interaction features between variables are modeled. Finally, the potential representation of the data is learned in the reconstruction module to generate reconstructed data, and the reconstruction error and the KL divergence of the latent space are calculated for model training, thereby performing efficient anomaly detection under unsupervised conditions.
[0192] Compared with traditional anomaly detection methods, this invention not only focuses on the spatial correlation in the multi-dimensional time series data of the industrial control system, but also takes into account the dynamic characteristics of the features over time, and can adaptively adjust the training strategy according to the characteristics of the data to ensure accurate anomaly detection and identification in different environments. This invention improves the accuracy of anomaly detection in industrial control systems, provides a more flexible and accurate anomaly detection tool for practical applications, and adapts to the complex dynamic changes of time series data in industrial control systems without labeled data.
[0193] Embodiment 2
[0194] like Figure 3 As shown, the second embodiment of the present invention further provides an abnormality detection device based on adaptive multi-scale feature modeling, comprising:
[0195] A data preprocessing unit is used to obtain the historical monitoring time series data of the industrial control system and preprocess the data to obtain the training data for the input anomaly detection model;
[0196] A component decomposition unit, used to combine Fourier transform with sliding windows of different sizes, decompose the input training data into time components, and then fuse them with the training data to obtain a decomposition result;
[0197] A multi-scale fusion unit, used to calculate a gating weight matrix based on the decomposition result in combination with a gating network, and adaptively weightedly combine the outputs of a plurality of expert networks of different scales processing the training data to obtain a multi-scale fusion feature;
[0198] A graph convolution unit is used to calculate the adjacency matrix of the multi-scale fusion feature at each time step, and aggregate the node features of the adjacency matrix using a graph convolution operation to capture the spatial correlation features between nodes and obtain spatial features;
[0199] A spatiotemporal combining unit, used for performing weighted summation of the spatial feature and the multi-scale fusion feature to obtain a final feature representation combining spatiotemporal information;
[0200] A data reconstruction unit, used for sending the final feature representation into a variational autoencoder to learn a latent space representation, and performing data decoding and reconstruction to obtain reconstructed data;
[0201] The model training unit is used to train the model by combining the reconstructed data error loss and KL divergence as the loss function to obtain a trained anomaly detection model;
[0202] The anomaly detection unit is used to input the preprocessed data to be detected into the anomaly detection model and obtain the anomaly detection result in combination with the set threshold.
[0203] Embodiment 3
[0204] The third embodiment of the present invention also provides an anomaly detection device based on adaptive multi-scale feature modeling, which includes a memory and a processor, wherein a computer program is stored in the memory, and the computer program can be executed by the processor to implement the anomaly detection method based on adaptive multi-scale feature modeling as described above.
[0205] Embodiment 4
[0206] The fourth embodiment of the present invention also provides a computer-readable storage medium, on which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor of a device where the computer-readable storage medium is located, the anomaly detection method based on adaptive multi-scale feature modeling as described above is implemented.
[0207] In several embodiments provided in the embodiments of the present invention, it should be understood that the disclosed apparatus and method can also be implemented in other ways. The apparatus and method embodiments described above are merely schematic. For example, the flowcharts in the accompanying drawings show the possible architecture, functions and operations of the apparatus, method and computer program product according to multiple embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of a code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.
[0208] In addition, the functional modules in the various embodiments of the present invention may be integrated together to form an independent part, or each module may exist independently, or two or more modules may be integrated to form an independent part.
[0209] If the function is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, an electronic device, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program code. It should be noted that in this article, the term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, article or device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, article or device. Without more constraints, an element defined by the phrase "comprising a..." does not exclude the existence of other identical elements in the process, method, article or apparatus comprising the element.
[0210] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0211] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0212] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0213] The "first\second" mentioned in the embodiments is only to distinguish similar objects, and does not represent a specific order for the objects. It is understandable that the "first\second" can be interchanged with the specific order or sequence where permitted. It should be understood that the objects distinguished by "first\second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than those illustrated or described herein.
[0214] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An anomaly detection method based on adaptive multi-scale feature modeling, characterized in that: include: S1, obtain the historical monitoring time series data of the industrial control system and preprocess the data to obtain the training data for the input anomaly detection model; S2, combining Fourier transform with sliding windows of different sizes, decomposing the input training data by time component and fusing it with the training data to obtain a decomposition result; S3, according to the decomposition result, a gating weight matrix is calculated in combination with a gating network, and the output of processing the training data by a plurality of expert networks of different scales is adaptively weighted and combined to obtain a multi-scale fusion feature; Specifically: According to the decomposition results , extract the feature vector of each time step, and obtain the data decomposition result subset; generate the initial gating weight matrix according to the data decomposition result subset, the formula is: ; in, The decomposition result is represented by The gating weight matrix of the subsets; is the activation function; and are the learned weights and bias terms respectively; Represents the decomposition result No. The input features of the subset; According to the training data, the feature vector of each time step is extracted to obtain a data subset; according to the gated weight matrix, the data subset is allocated to a plurality of expert networks of different scales to select the most suitable expert network to process the current data and obtain the output feature , the expression is: ; ; in, Indicates the current data entered The corresponding gating weight matrix; Indicates from Select the output with the largest weight experts; represents the expert weight representation after noise injection; Indicates the number of selected experts; represents the gating weight; represents noise sampled from a standard normal distribution with mean 0 and variance 1; Represents a smooth ReLU function to enhance the nonlinear expression ability of the model; represents the noise weight; Based on the gated weight matrix, the features of the outputs of the expert networks at different scales are weighted and summed to obtain the multi-scale fusion features. , the expression is: ; in, It is Output features of the expert network; represents the total number of expert networks; S4, calculating the adjacency matrix of the multi-scale fusion feature at each time step, and aggregating the node features of the adjacency matrix using a graph convolution operation to capture the spatial correlation features between nodes to obtain spatial features; S5, performing weighted summation of the spatial feature and the multi-scale fusion feature to obtain a final feature representation combining spatiotemporal information; S6, sending the final feature representation to a variational autoencoder to learn a latent space representation, and performing data decoding and reconstruction to obtain reconstructed data; S7, combining the reconstructed data error loss and KL divergence as the loss function to train the model and obtain a trained anomaly detection model; S8, input the preprocessed data to be detected into the anomaly detection model, and combine it with the set threshold to obtain the anomaly detection result.
2. The method for anomaly detection based on adaptive multi-scale feature modeling according to claim 1, characterized in that ,The preprocessing includes processing missing values and repeated values of the data.
3. The method for anomaly detection based on adaptive multi-scale feature modeling according to claim 1, characterized in that , the input training data is decomposed into time components, specifically: The frequency characteristics of the input data are analyzed by Fourier transform technology to capture periodic changes and obtain seasonal components; wherein the seasonal components are regular fluctuation data in the data that changes with a fixed period; A moving average of the input data is calculated using multiple sliding windows of different sizes, and the overall upward or downward trend of the input data is captured by smoothing short-term fluctuations to obtain a trend component; wherein the trend component is the trend data of long-term changes in the data, including continuously growing data, continuously declining data, or unchanged data; The seasonal component and the trend component are subtracted from the input data to obtain a residual component; the residual component includes random fluctuation data or noise data.
4. The method for anomaly detection based on adaptive multi-scale feature modeling according to claim 1, characterized in that ,Expert networks of different scales define fragments of different sizes, thereby extracting different scale features of the input data.,When each expert network processes its assigned data, the processing steps are as follows: Divide the allocated data into multiple time segments ; By setting different time slice sizes , define the processing resolution of time series data for each expert network, each time segment represents a local time window, and use the local window to learn the time dependency of the data, then: ; in, represents the i-th time segment; Represents the time series data in the i-th time segment; Each expert network extracts local features of each time segment through a deep grouped convolutional layer to obtain the short-term dependency features of continuous time steps in each time segment in the time series. , the expression is: ; in, Represents the convolution feature extraction operation; Using multi-head self-attention to learn different time segments long-term dependencies between them, building interactions between different time segments and generating global features ; When calculating multi-head self-attention, the query matrix, key matrix and value matrix transformation are calculated, and the features of each time segment are aggregated through the attention weights; Concatenate the features learned by each expert network, combining the assigned data and short-term dependency features and global features across time segments , get the fusion features output by each expert network .
5. The method for anomaly detection based on adaptive multi-scale feature modeling according to claim 1, characterized in that , the S4 is specifically: Calculate the cosine similarity between all node features of the multi-scale fusion feature in each time step, and for each time step Each pair of nodes and , calculate the cosine similarity, the formula is: ; in, Represents the feature vector and The cosine similarity between ; , Respectively represent nodes and nodes , calculate their dot product and normalize; represents the L2 norm; If the similarity between two nodes is greater than the preset threshold , then establish edge connections and obtain the adjacency matrix, the expression is: ; in, Indicates the time step The corresponding adjacency matrix; Based on the adjacency matrix, the GraphSAGE convolutional layer is used to aggregate the features of the nodes to capture the spatial correlation features between the nodes, namely the spatial features , the formula of the GraphSAGE convolutional layer is: ; in, Representation Node In the Feature representation of the layer; ) represents a node The neighbor node set of ; Aggregate is the aggregation operation; For the The weight matrix of the layer; is the activation function; Representation Node Neighbor nodes The feature representation of .
6. The method for anomaly detection based on adaptive multi-scale feature modeling according to claim 1, characterized in that , the S6 is specifically: The final feature representation is input into the variational autoencoder, and the latent space mean of the final feature representation is obtained through forward propagation learning of the fully connected network of the variational autoencoder. μ With log variance ; and use the reparameterization technique to sample the latent variables from the latent space , the expression is: ; in, It follows the standard normal distribution The noise sampled in represents the standard deviation of the latent space; Through a fully connected layer, the latent variables Map back to the input space, reconstruct the input data to get the reconstructed data , to learn a compressed representation of the input data.
7. The method for anomaly detection based on adaptive multi-scale feature modeling according to claim 6, characterized in that , the reconstructed data error loss is: the mean square error is used to measure the reconstructed data The difference between the training data X input to the model is expressed as: ; in, represents the reconstructed data error loss; The KL divergence It measures the difference between the distribution of the latent space and the standard normal distribution, and the expression is: ; in, represents the KL divergence; a total number of data subsets representing the training data; Indicates The standard deviation of a subset of data; Indicates The latent space mean of a subset of data; The loss function is the weighted sum of the reconstructed data error loss and the KL divergence, and the formula is: ; in, represents the weight of the KL divergence.
8. The method for anomaly detection based on adaptive multi-scale feature modeling according to any one of claims 1 to 7, characterized in that , and also includes testing and evaluating the anomaly detection model. When the test evaluation result meets the set requirements, the anomaly detection model is evaluated as qualified and used for subsequent anomaly detection; otherwise, the model is retrained; wherein the test evaluation result of the anomaly detection model is calculated by the following formula: Accuracy The calculation formula is: ; Accuracy The calculation formula is: ; Recall The calculation formula is: ; The calculation formula is: ; Among them, TP represents the number of true positives, that is, the number of samples where both the actual situation and the test results are normal; FP represents the number of false positives, that is, the number of samples where the actual situation is abnormal and the test results are normal; TN represents the number of true negatives, that is, the number of samples where both the actual situation and the test situation are abnormal; FN represents the number of false negatives, that is, the number of samples where the actual situation is normal and the test results are abnormal.
9. An anomaly detection device based on adaptive multi-scale feature modeling, used to implement an anomaly detection method based on adaptive multi-scale feature modeling according to any one of claims 1 to 7, characterized in that: include: A data preprocessing unit is used to obtain the historical monitoring time series data of the industrial control system and preprocess the data to obtain the training data for the input anomaly detection model; A component decomposition unit, used to combine Fourier transform with sliding windows of different sizes, decompose the input training data into time components, and then fuse them with the training data to obtain a decomposition result; A multi-scale fusion unit, used to calculate a gating weight matrix based on the decomposition result in combination with a gating network, and adaptively weightedly combine the outputs of a plurality of expert networks of different scales processing the training data to obtain a multi-scale fusion feature; Specifically: According to the decomposition results , extract the feature vector of each time step, and obtain the data decomposition result subset; generate the initial gating weight matrix according to the data decomposition result subset, the formula is: ; in, The decomposition result is represented by The gating weight matrix of the subsets; is the activation function; and are the learned weights and bias terms respectively; Represents the decomposition result No. The input features of the subset; According to the training data, the feature vector of each time step is extracted to obtain a data subset; according to the gated weight matrix, the data subset is allocated to a plurality of expert networks of different scales to select the most suitable expert network to process the current data and obtain the output feature , the expression is: ; ; in, Indicates the current data entered The corresponding gating weight matrix; Indicates from Select the output with the largest weight experts; represents the expert weight representation after noise injection; Indicates the number of selected experts; represents the gating weight; represents noise sampled from a standard normal distribution with mean 0 and variance 1; Represents a smooth ReLU function to enhance the nonlinear expression ability of the model; represents the noise weight; Based on the gated weight matrix, the features of the outputs of the expert networks at different scales are weighted and summed to obtain the multi-scale fusion features. , the expression is: ; in, It is Output features of the expert network; represents the total number of expert networks; A graph convolution unit is used to calculate the adjacency matrix of the multi-scale fusion feature at each time step, and aggregate the node features of the adjacency matrix using a graph convolution operation to capture the spatial correlation features between nodes and obtain spatial features; A spatiotemporal combining unit, used for performing weighted summation of the spatial feature and the multi-scale fusion feature to obtain a final feature representation combining spatiotemporal information; A data reconstruction unit, used for sending the final feature representation into a variational autoencoder to learn a latent space representation, and performing data decoding and reconstruction to obtain reconstructed data; The model training unit is used to train the model by combining the reconstructed data error loss and KL divergence as the loss function to obtain a trained anomaly detection model; The anomaly detection unit is used to input the preprocessed data to be detected into the anomaly detection model and obtain the anomaly detection result in combination with the set threshold.
Citation Information
Patent Citations
Adaptive correlation perception unsupervised deep learning anomaly detection method
CN115879505A
Multi-element sensor time series data anomaly detection method based on time series converter
CN118535981A