An abnormal flow detection method and system based on spatiotemporal modeling of industrial processes
Patent Information
- Application Number
- CN202510887692.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-06-30
Smart Images

Figure CN120434042B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of abnormal flow detection, and in particular relates to an abnormal flow detection method and system based on spatiotemporal modeling of industrial processes. Background Art
[0002] Industrial process anomaly detection is an important part of the development of industrial automation and intelligence. By analyzing and modeling the collected data and identifying abnormal data points, it can help companies promptly discover potential problems, formulate safety protection strategies, improve production efficiency, reduce maintenance costs and improve product quality. It is of great significance for companies to achieve safe, efficient and sustainable operations.
[0003] Industrial process data is essentially a multivariate time series. Anomaly detection in industrial process data is essentially time series detection. Traditional anomaly detection methods can be categorized as clustering-based, distance-based, density-based, and isolation-based. In recent years, with the development of deep learning technology, deep learning methods have been widely applied to time series anomaly detection. Existing deep learning-based anomaly detection methods can be divided into two types: reconstruction-based methods and prediction-based methods.
[0004] The paper "Miao J, Tao H, Xie H, et al. Reconstruction-based anomaly detection for multivariate time series using contrastive generative adversarial networks[J]. Information Processing & Management, 2024, 61(1):103569." proposes a new unsupervised anomaly detection framework that seamlessly integrates contrastive learning and generative adversarial networks and detects anomalies based on reconstruction errors.
[0005] The paper "Deng A, Hooi B. Graph neural network-based anomaly detection inmultivariate time series[C]. Proceedings of the AAAI conference on artificial intelligence. 2021, 35(5): 4027-4035." proposed a graph deviation network (GDN), which uses embedding to construct a graph structure and uses graph attention to learn the characteristics of time series. It obtains anomaly scores based on observed values and predicted values, thereby judging time series anomalies.
[0006] However, the method proposed in the paper "Miao J, Tao H, Xie H, et al. Reconstruction-based anomaly detection for multivariate time series using contrastive generative adversarial networks[J]. Information Processing & Management, 2024, 61(1):103569." is mainly designed for regularly sampled multivariate time series data. When faced with irregularly sampled data or data with a large number of missing values, it may be difficult to capture and model the temporal dependencies in the time series.
[0007] The method proposed in the paper "Deng A, Hooi B. Graph neural network-based anomaly detection inmultivariate time series[C]. Proceedings of the AAAI conference on artificialintelligence. 2021, 35(5): 4027-4035." takes into account the spatial relationship between variables, but ignores the characteristics of the industrial process itself and only starts from the data, resulting in inaccurate constructed spatial relationships, which limits the effectiveness of anomaly detection.
[0008] In summary, although various deep learning-based anomaly detection algorithms have emerged in recent years, most of the current methods are only data-driven and cannot accurately mine the spatial relationships of various automation components, and thus cannot accurately extract spatial features and temporal features. As a result, the constructed models do not conform to the actual scenarios, the anomaly detection effect is not ideal, and there is a lack of analysis and explanation of anomalies. Summary of the Invention
[0009] In response to the shortcomings of the existing technology, the present invention proposes an abnormal flow detection method and system based on industrial process spatiotemporal modeling to solve the problem that the existing methods are unable to accurately extract spatial features and temporal features, resulting in unsatisfactory anomaly detection results.
[0010] The technical solution of the present invention is:
[0011] A first aspect of the present invention provides a method for detecting abnormal flow based on spatiotemporal modeling of industrial processes, comprising the following specific steps:
[0012] Acquire industrial process data and preprocess the industrial process data, including data normalization and data windowing, to obtain windowed serialized data; the industrial process data is a multivariate time series including all sensor data and actuator data collected in the industrial process;
[0013] Constructing an anomaly detection model; the anomaly detection model is used to calculate an anomaly score based on industrial process data within an input time window;
[0014] Use the window serialization data to train the anomaly detection model to obtain a trained anomaly detection model;
[0015] Industrial process data is acquired and preprocessed, and the preprocessed industrial process data within a time window is input into the trained anomaly detection model to obtain an anomaly score. Based on the anomaly score, whether an anomaly occurs is detected.
[0016] Furthermore, the industrial process data is obtained and preprocessed, including data normalization and data windowing, to obtain window serialized data, specifically:
[0017] A1: Acquire industrial process data; the industrial process data is , Indicates The collection of all sensor data and actuator data collected at all times, Indicates the moment, Represents the length of industrial process data, represents the total number of automation components, including sensors and actuators;
[0018] A2: Perform normalization processing on the industrial process data to obtain normalized industrial process data;
[0019] A3: Convert the normalized industrial process data into window data, and then obtain window serialized data including multiple window data;
[0020] For the current moment , define the length as The window data is , including from Time has come Normalized industrial process data at time After normalization The collection of all sensor data and actuator data collected at all times; the window serialized data is represented as , including several lengths of Window data.
[0021] Furthermore, the anomaly detection model includes a spatial relationship construction module, a spatial feature extraction module, a time series feature extraction module, a time series prediction module, a reconstruction module and a scoring module;
[0022] The spatial relationship building module is used to construct a graph structure based on the acquired industrial process knowledge and modify the graph structure based on the input window data to obtain a directed graph G; the industrial process knowledge includes the spatial relationship between various automation components in the industrial process; the window data is the industrial process data within a time window;
[0023] The spatial feature extraction module is used to obtain the spatial features of each node in the directed graph G; the spatial feature extraction module is a graph attention network;
[0024] The temporal feature extraction module is used to extract short-term temporal features using the GRU model and long-term temporal features using the Informer model based on the spatial features of all nodes;
[0025] The time series prediction module is used to perform predictions based on short-term time series features and long-term time series features, respectively, to obtain short-term time series prediction results and long-term time series prediction results;
[0026] The reconstruction module is used to reconstruct the short-term time series features and obtain the reconstruction probability of the short-term time series features; a variational autoencoder (VAE) is used as the reconstruction module;
[0027] The scoring module is used to calculate anomaly scores based on the short-term time series prediction results, the long-term time series prediction results and the reconstruction probability of the short-term time series features obtained by the reconstruction module.
[0028] Furthermore, the graph structure construction method is specifically as follows: each automation component is regarded as a node and represented by a binary pair A=(c,r), where c represents the type of the automation component, 0 represents a sensor, 1 represents an actuator, and r represents the number of the sub-process to which the automation component belongs in the entire industrial process; then, the automation components that influence each other are determined through the spatial relationship of the automation components included in the industrial process knowledge, and the influence relationship between the automation components is represented by directed edges, a graph structure is constructed and saved using an adjacency matrix, and an edge set E is obtained at the same time; then, the edges in the constructed graph structure are verified, and the edges in the edge set E are corrected using the Pearson correlation coefficient and transfer entropy respectively, to obtain a directed graph G;
[0029] The process of modifying the edges in the edge set E using the Pearson correlation coefficient and the transfer entropy is specifically as follows:
[0030] Extract the time series data of different automation components from the input window data, calculate the Pearson correlation coefficient and transfer entropy between the time series data of different automation components, and set a Pearson correlation coefficient threshold θ p and a transfer entropy threshold θ t , if the absolute value of the Pearson correlation coefficient between the time series data of two automation components is greater than or equal to the Pearson correlation coefficient threshold θ p , then retain the edge between the two automation components in the graph structure if the absolute value of the Pearson correlation coefficient between the time series data of the two automation components is less than the Pearson correlation coefficient threshold θ p , then further determine whether the transfer entropy between the time series data of the two automation components is greater than or equal to the transfer entropy threshold θ t If so, the edge between the two automation components in the graph structure is retained; otherwise, the edge between the two automation components in the graph structure is deleted.
[0031] Furthermore, the method for obtaining the spatial features of each node in the directed graph G is as follows: first, the features of each node in the directed graph G are initialized to obtain the initial spatial features of each node: the time series data of different automation components extracted from the input window data are extracted through a linear layer. Convert to dimensional vector, and obtain the initial spatial characteristics of the node corresponding to the automation component ,in is the dimension of the initial spatial feature, Represents the number of the automation component; then, the multi-head attention mechanism of the graph attention network is used to update the initial spatial features of each node to obtain the final spatial features of each node;
[0032] The time series prediction module uses the short-term time series features output by the GRU model as Elements and the corresponding initial spatial features Multiply them one by one, and then input the results into the fully connected layer to perform short-term sequence prediction to obtain the short-term time series prediction results;
[0033] Long-term time series characteristics Input the fully connected layer and get arrive A vector of predicted values of data of all automation components at the moment , that is, the long-term time series prediction result, The time step length for the predicted future industrial process data;
[0034] Furthermore, the abnormality score is:
[0035] (19);
[0036] in, For abnormality score, is a hyperparameter, for The short-term time series prediction results and the long-term time series prediction results at the moment The average value of the predicted values of automation component data, for Moment The true value of the data of the automation component, Short-term time series characteristics No. The reconstruction probability of an element.
[0037] Furthermore, when the anomaly detection model is trained using window serialization data, the construction process of the joint loss function is:
[0038] Use mean squared error as the loss function for short-term time series forecasting:
[0039] (twenty one);
[0040] in, is the loss function for short-term time series prediction, for The short-term time series prediction results of The predicted value of data of an automation component;
[0041] Use mean squared error as the loss function for long-term time series forecasting:
[0042] (twenty two);
[0043] in, is the loss function for long-term time series prediction, is a positive integer, exist Moment The true value of the data of each automation component; Long-term time series prediction results middle Moment The predicted value of data of an automation component;
[0044] The reconstruction loss can be defined as:
[0045] (twenty three);
[0046] in, is the reconstruction loss, is the reconstruction error term, is to calculate the KL divergence, is a potential feature, For a given latent feature Generate short-term time series features The probability of For a scalable distribution, for The prior probability of
[0047] The joint loss function is defined as:
[0048] (twenty four);
[0049] in, is the joint loss function, These are hyperparameters that need to be set manually.
[0050] A second aspect of the present invention provides an abnormal flow detection system based on spatiotemporal modeling of industrial processes, which is used to implement an abnormal flow detection method based on spatiotemporal modeling of industrial processes, comprising:
[0051] Industrial process data acquisition module, used to obtain industrial process data;
[0052] Data preprocessing module, used to preprocess industrial process data to obtain window data;
[0053] Anomaly detection model, which is used to calculate anomaly scores based on the input window data;
[0054] The anomaly recognition module is used to determine whether an anomaly has occurred based on the anomaly score.
[0055] The third aspect of the present invention provides an electronic device, comprising: a processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate through the bus, and when the machine-readable instructions are executed by the processor, the steps of the abnormal flow detection method based on spatiotemporal modeling of industrial processes are performed.
[0056] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, executes the steps of the abnormal flow detection method based on spatiotemporal modeling of industrial processes.
[0057] Compared with the prior art, the present invention has the following beneficial effects:
[0058] (1) This paper proposes a method for establishing an industrial process model that combines industrial process knowledge and data-driven, thereby providing more accurate spatial and temporal features, building a model that is more in line with actual scenarios, and improving the anomaly detection effect.
[0059] (2) The present invention proposes a GRU and Informer joint time series feature extraction method for extracting short-term and long-term time series characteristics of industrial process data, which conforms to the characteristics of industrial process data and improves the accuracy of abnormal flow detection for industrial processes.
[0060] (3) This paper proposes a joint optimization strategy that combines the prediction model and the reconstruction model, which not only enables the model to identify the global distribution characteristics of the time series, but also pays attention to the individual characteristics of the variables, thereby enhancing the accuracy and robustness of the model in the anomaly detection task. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 Flowchart of anomaly detection method based on spatiotemporal modeling of industrial processes in an embodiment of the present invention;
[0062] Figure 2 This is a diagram of the overall architecture of the anomaly detection model based on spatiotemporal modeling of industrial processes in an embodiment of the present invention. DETAILED DESCRIPTION
[0063] The present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0064] This embodiment proposes an abnormal flow detection method based on industrial process spatiotemporal modeling, which includes six parts: data preprocessing, spatial relationship construction, spatial feature extraction, temporal feature extraction, joint optimization and anomaly scoring. Figure 1 As shown, the specific steps are as follows:
[0065] Step 1: Acquire and preprocess industrial process data, including data normalization and windowing, to obtain windowed serialized data for subsequent model training and testing. The industrial process data is a multivariate time series, which refers to a collection of various data and information related to the industrial process, including all sensor data and actuator data collected by the Industrial Control System (ICS) through the Supervisory Control and Data Acquisition (SCADA) system at different times during the industrial process.
[0066] Step 1.1: Obtain industrial process data; the industrial process data is , Indicates The collection of all sensor data and actuator data collected at all times, Indicates the moment, Represents the length of industrial process data, represents the total number of automation components, including sensors and actuators;
[0067] This example uses industrial process data from the SWaT dataset, which includes the states of sensors and actuators collected every second over 11 days;
[0068] Step 1.2: Perform normalization processing on the industrial process data to obtain normalized industrial process data;
[0069] The numerical range of industrial process data is adjusted to [0, 1] as follows:
[0070] (1);
[0071] in, represent The first time collected Data of automation components, Indicates the number of the automation component, Representative The time series data of each automation component, including the first Data of automation components, Represents the normalized The first time collected Data of automation components, and Respectively represent the minimum and maximum values in all time steps;
[0072] Step 1.3: Convert the normalized industrial process data into window data, and then obtain window serialization data including a plurality of window data;
[0073] For the current moment , define the length as The window data is , including from Time has come Normalized industrial process data at time After normalization The collection of all sensor data and actuator data collected at all times; the window serialized data is represented as , including several lengths of Window data; window serialization data This will be used to train the model. In this embodiment, data windowing is performed on both the training set and the test set, dividing them into multiple continuous windows of data. Only normal data is selected in the training set to allow the model to grasp normal data patterns. In addition to normal data, the test set also contains abnormal data to test and verify the accuracy and effectiveness of the model in identifying anomalies.
[0074] Step 2: Build an anomaly detection model; the anomaly detection model is used to calculate an anomaly score based on the industrial process data within a time window of the input;
[0075] like Figure 2 As shown, the anomaly detection model includes a spatial relationship construction module, a spatial feature extraction module, a time series feature extraction module, a time series prediction module, a reconstruction module and a scoring module;
[0076] The spatial relationship building module is used to construct a graph structure based on the acquired industrial process knowledge and modify the graph structure based on the input window data to obtain a directed graph G; the industrial process knowledge includes the spatial relationship between various automation components in the industrial process; the window data is the industrial process data within a time window;
[0077] Industrial process knowledge includes an in-depth understanding of the chemical reactions and physical changes involved in industrial processes, as well as the working principles and operating methods of the equipment and machinery used in production. However, this expert knowledge is extremely difficult to calculate and obtain, and even non-experts find it difficult to understand and analyze. Therefore, the industrial process knowledge described in this invention refers to simple industrial process knowledge obtained through information such as flow charts and piping and instrumentation diagrams. For example, a flow chart can be used to understand how many sub-processes an entire industrial process is divided into, the sequential relationships between sub-processes, and the number of sensors and actuators contained in each sub-process.
[0078] Based on this industrial process knowledge, the spatial relationship of automation components in the industrial system can be constructed. The spatial relationship refers to whether there is any influence between components. For the relationship between two automation components between sub-processes, only adjacent sub-processes are considered here. This is because the automation components of adjacent sub-processes are more likely to influence each other, and the relationship between the automation components of non-adjacent sub-processes can be transmitted through the sub-processes between them.
[0079] For example, for the automation component e in sub-process h and the automation component f in sub-process h+1, if both e and f are sensors, since sub-process h+1 is performed after sub-process h, automation component e will affect automation component f, while automation component f will not affect automation component e. If automation component e is a sensor and automation component f is an actuator, then the influence of automation component e on automation component f can be transmitted through the sensor of sub-process h+1, and no additional connection is required. If automation component e is an actuator and automation component f is a sensor, then the influence of automation component e on automation component f can be transmitted through the sensor of sub-process h, and no additional connection is required. If automation component e and automation component f are both actuators, then the influence of automation component e on automation component f can be transmitted through the associated sensors in sub-process h and sub-process h+1, and no additional connection is required.
[0080] The graph structure construction method is specifically as follows: each automation component is regarded as a node and represented by a binary tuple A = (c, r), where c represents the type of the automation component, 0 represents a sensor, 1 represents an actuator, and r represents the number of the sub-process to which the automation component belongs in the entire industrial process; then, the spatial relationship of the automation components included in the industrial process knowledge is used to determine the automation components that have an impact on each other, and the impact relationship between the automation components is represented by directed edges. The graph structure is constructed and saved using an adjacency matrix, and the edge set E is obtained at the same time;
[0081] Considering that the aforementioned graph structure construction method relies on the industrial process knowledge described by the industrial process diagram and the piping and instrumentation diagram, some irrelevant relationships may be introduced simply based on intuitive industrial process knowledge. Therefore, it is necessary to verify the edges in the constructed graph structure and delete the edges that fail the verification. Considering that linear and nonlinear relationships may exist at the same time, the Pearson correlation coefficient and transfer entropy are used to modify the edges in the edge set E, respectively, to achieve graph structure modification and obtain the directed graph G.
[0082] The process of modifying the edges in the edge set E using the Pearson correlation coefficient and the transfer entropy is specifically as follows:
[0083] Extract the time series data of different automation components from the input window data, calculate the Pearson correlation coefficient and transfer entropy between the time series data of different automation components, and set a Pearson correlation coefficient threshold θ p and a transfer entropy threshold θ t , if the absolute value of the Pearson correlation coefficient between the time series data of two automation components is greater than or equal to the Pearson correlation coefficient threshold θ p, then retain the edge between the two automation components in the graph structure if the absolute value of the Pearson correlation coefficient between the time series data of the two automation components is less than the Pearson correlation coefficient threshold θ p , which means that the influence of noise occupies a dominant position, and it cannot be used as a basis for verifying the spatial relationship. Further judgment is made on whether the transfer entropy between the time series data of the two automation components is greater than or equal to the transfer entropy threshold θ t If yes, the edge between the two automation components in the graph structure is retained; otherwise, the edge between the two automation components in the graph structure is deleted;
[0084] The spatial feature extraction module is used to obtain the spatial features of each node in the directed graph G; the spatial feature extraction module is a graph attention network;
[0085] The method for obtaining the spatial features of each node in the directed graph G is as follows: first, the features of each node in the directed graph G are initialized to obtain the initial spatial features of each node: the time series data of different automation components extracted from the input window data are converted into the time series data of the different automation components extracted from the input window data through a linear layer. Convert to dimensional vector, and obtain the initial spatial characteristics of the node corresponding to the automation component ,in is the dimension of the initial spatial feature;
[0086] Then, the multi-head attention mechanism of the graph attention network is used to update the initial spatial features of each node to obtain the final spatial features of each node;
[0087] (2);
[0088] in, Indicates the The final spatial features of the nodes after aggregation by the multi-head attention mechanism, represents the sigmoid activation function, Indicates the number of the attention head, represents the number of attention heads, Indicates the node number, Indicates the The set of neighbor nodes of a node, Representatives in the Attention head nodes and The attention coefficient between nodes, Indicates the The weight matrix of the attention head, Indicates the The initial spatial features of the nodes, Represents a connection;
[0089] The attention coefficient is obtained by first activating and then performing softmax normalization:
[0090] (3);
[0091] (4);
[0092] in, represents the attention score, Indicates the The initial spatial features of the nodes, represents the nonlinear activation function, Represents connecting two vectors. Indicates the The learning coefficient vector of the attention heads, is an exponential function, represents transpose;
[0093] The industrial process can be regarded as a complex and continuous system, in which the equipment, production lines and operating status of the entire system involved will change over time, and some of these changes may be short-term and instantaneous. In addition, due to the relatively fixed position behavior of industrial control equipment, the cyclic characteristics of the production process and the continuity requirements of product production, industrial process data often have periodic characteristics. In summary, industrial process data should have both short-term time series characteristics and long-term time series characteristics. The time series feature extraction module of the present invention is used to extract short-term time series features using the GRU model and long-term time series features using the Informer model based on the spatial characteristics of all nodes, so as to obtain richer industrial process features;
[0094] The method for obtaining the short-term time series features is specifically as follows:
[0095] At the current moment , GRU model extracts the spatial feature from the input Spatial characteristics of all nodes at the moment And the hidden state passed down by the GRU model at the previous moment , update the hidden state of the current moment as a short-term time series feature, and the update of the hidden state is controlled by the update gate inside the GRU model and reset gate Control, update gate and reset gate The formulas are as follows:
[0096] (5);
[0097] (6);
[0098] in, is the input weight matrix of the update gate, It acts on the hidden state at the previous moment The weight matrix, is the input weight matrix of the reset gate, It acts on the hidden state at the previous moment The weight matrix of
[0099] Then, the reset gate is used to calculate the candidate hidden state , used to compare with the hidden state passed down from the previous moment And the update gate calculates the subsequent output, the candidate hidden state The calculation method is:
[0100] (7);
[0101] Among them, ⊙ represents the multiplication of vector elements, and tanh is the activation function that can scale the data to the range of -1 to 1; is the weight of the input to the candidate hidden state, is the weight of the hidden state passed down from the previous moment to the candidate hidden state;
[0102] Finally, the update gate is used to obtain the output result, which is the short-term time series feature:
[0103] (8);
[0104] in, It is a short-term time series feature;
[0105] The method for obtaining the long-term time series features is specifically as follows:
[0106] First, the short-term time series features output by the GRU model As the original input of the encoder in the Informer model, the short-term time series features Perform one-dimensional convolution operation to obtain time embedding , which is the short-term time series characteristic Add positional encoding , embed the time of the one-dimensional convolution output With positional encoding Add together to get the final input of the encoder :
[0107] (9);
[0108] in, is the final input of the encoder in the Informer model, For time embedding, Encode for position;
[0109] Then, multi-head ProbSpare (sparse) self-attention calculation is performed, the formula is as follows:
[0110] (10);
[0111] in, For attention output, and is the final input The query matrix, key matrix, and value matrix obtained by multiplying different weight matrices, are the dimensions of the query matrix and key matrix, The highest score is selected query vectors, the scoring is calculated as follows:
[0112] (11);
[0113] in, is a score used to measure the Is the query vector important? For the query vectors, is the query vector number, It is key vectors, is the key vector number, is the number of key vectors;
[0114] Next, the self-attention distillation operation is performed, which is used to extract the most important attention. The formula is as follows:
[0115] (12);
[0116] Where, express The time encoder The hidden states of the encoder layers, express The time encoder The hidden states of the encoder layers, Represents the attention block, including necessary operations such as multi-head ProbSpare self-attention, residual connection and layer normalization, Indicates that a one-dimensional convolution filtering operation is performed using 1D convolution in the time dimension. Is an activation function used to introduce nonlinearity, It is a pooling layer with a stride of 2. After calculating one layer, the hidden state Halving is used to extract the main features, and after calculations by multiple encoder layers in the encoder, the output features of the encoder are obtained;
[0117] The decoder in the Informer model is a network consisting of a masked multi-head ProbSparse self-attention layer and a traditional self-attention layer. The decoder uses masked sparse self-attention.
[0118] The input features of the masked multi-head ProbSparse self-attention layer in the decoder are:
[0119] (13);
[0120] in, is the input feature of the masked multi-head ProbSparse self-attention layer in the decoder, For splicing operations, It is historical industrial process data, and its length is less than the length of the industrial process data within a time window of the input. , It is industrial process data whose value is all 0, and its length is the time step length of the future industrial process data to be predicted, and Processed with one-dimensional convolution and positional encoding embedding;
[0121] Finally, the output features of the encoder and the output of the masked multi-head ProbSparse self-attention layer are used as the input of the traditional self-attention layer in the decoder. After calculation, the final long-term temporal features are obtained. ;
[0122] The time series prediction module is used to perform predictions based on short-term time series features and long-term time series features, respectively, to obtain short-term time series prediction results and long-term time series prediction results;
[0123] Specifically: the short-term time series features output by the GRU model Elements and the corresponding initial spatial features Multiply them one by one and input the results into the fully connected layer for short-term sequence prediction. The formula is as follows:
[0124] (14);
[0125] in, represent The vector of predicted values of the data of all automation components at the moment, i.e. the short-term time series prediction result, Indicates prediction through the fully connected layer;
[0126] Long-term time series characteristics Input the fully connected layer and get arrive A vector of predicted values of data of all automation components at the moment , which is the long-term time series prediction result, the formula is as follows:
[0127] (15);
[0128] in, Represents a vector middle The predicted value of the data of all automation components at the moment, is a positive integer;
[0129] The reconstruction module is used to reconstruct short-term time series features. The present invention uses a variational autoencoder (VAE) as the reconstruction module. Considering that VAE performs poorly in periodic scenarios, which means that VAE is not suitable for processing long-term time series features, the present invention only reconstructs short-term time series features. Input reconstruction module;
[0130] Specifically, through the posterior distribution To reconstruct short-term time series features;
[0131] (16);
[0132] in, is a potential feature, For a given latent feature Generate H t The probability of VAE training is to make the latent features The posterior distribution of maximize, is a short-term time series feature The reconstruction probability of for The prior probability of
[0133] (17);
[0134] in, Short-term time series characteristics No. The reconstruction probability of an element;
[0135] Short-term time series characteristics The reconstruction probability The calculation formula is:
[0136] (18);
[0137] The time series prediction module excels at capturing short-term dynamic changes in time series and is sensitive to local trends and short-term dependencies between variables. The reconstruction module focuses on learning the global data structure and patterns of the entire time window and is capable of modeling long-term dependencies and stable statistical characteristics. By combining the two, it can not only identify the global distribution characteristics of time series but also focus on the individual characteristics of variables, thereby enhancing the accuracy and robustness of the model in anomaly detection tasks.
[0138] The scoring module is used to predict the results based on the short-term time series , long-term time series forecast results and the short-term timing characteristics obtained by the reconstruction module The reconstruction probability of ,calculate the anomaly score;
[0139] (19);
[0140] in, For abnormality score, is a hyperparameter to balance prediction and loss, for The short-term time series prediction results and the long-term time series prediction results at the moment The average value of the predicted values of automation component data, for Moment The true value of the data of each automation component;
[0141] (20);
[0142] in, for The short-term time series prediction results of The predicted value of the data of the automation components, for The long-term time series prediction results of the moment The predicted value of data of an automation component;
[0143] Step 3: Use the window serialization data to train the anomaly detection model to obtain the trained anomaly detection model;
[0144] Loss function during training:
[0145] The present invention uses mean square error as the loss function for short-term time series prediction:
[0146] (twenty one);
[0147] in, is the loss function for short-term time series prediction;
[0148] Use mean squared error as the loss function for long-term time series forecasting:
[0149] (twenty two);
[0150] in, is the loss function for long-term time series prediction, exist Moment The true value of the data of each automation component; is a vector middle Moment The predicted value of data of an automation component;
[0151] Since the latent features are directly calculated The posterior distribution of Very difficult, use another scalable distribution To approximate , optimized through deep networks Make it with Very similarly, the reconstruction loss can be defined as:
[0152] (twenty three);
[0153] in, is the reconstruction loss, is the reconstruction error term, Is to calculate KL divergence;
[0154] Assign different weights to the three loss functions, and the joint loss function is defined as the formula:
[0155] (twenty four);
[0156] in, These are hyperparameters that need to be set manually;
[0157] Step 4: Obtain and preprocess the industrial process data. Input the preprocessed industrial process data within a time window into the trained anomaly detection model to obtain an anomaly score. Based on the anomaly score, detect whether an anomaly has occurred.
[0158] The method for detecting whether an abnormality occurs is specifically:
[0159] A threshold is set. When the anomaly score exceeds the threshold, an anomaly exists. For the selection of the threshold, the present invention adopts an algorithm called Peaks Over Threshold (POT), which is specifically used to select a suitable anomaly threshold from the validation set.
[0160] This paper uses two public datasets and a physical platform built by iTrust Labs to simulate real-world process industry scenarios. The public datasets are the Secure Water Treatment (SWaT) dataset and the Water Distribution (WADI) dataset. The SWAT dataset contains 51 sensors and actuators, recording data for 11 days. The system operated normally for the first seven days, totaling 496,800 samples. The system was attacked for the last four days, and the data includes both normal and abnormal data, totaling 449,919 samples. The WADI dataset contains data from 123 sensors and actuators in the water distribution system. The data collection process lasted for 16 days, of which the first 14 days were collected under normal operation, totaling 1,048,571 samples. The last two days were collected under an attack scenario, totaling 172,081 samples. Dataset information is shown in Table 1.
[0161] Table 1 Dataset information;
[0162] Dataset Number of training sets Number of test sets Number of features Abnormal ratio (%) SWaT 496800 449919 51 11.98 WADI 1048571 172081 123 5.99
[0163] To validate the effectiveness of our proposed method, we conducted a series of experiments comparing our anomaly detection model with several popular baseline models, including DAGMM, LSTM-VAE, USAD, OmniAnomaly, BeatGAN, GDN, and TranAD. Our proposed method was compared with these seven baseline models on the SWaT and WADI datasets.
[0164] Table 2 shows the anomaly detection results of the proposed method and seven baseline models on the SWaT dataset. Table 3 shows the anomaly detection results of the proposed method and seven baseline models on the WADI dataset.
[0165] Table 2 Anomaly detection results of the proposed method and seven baseline models on the SWaT dataset;
[0166] method Precision Recall F1 DAGMM 0.2746 0.6952 0.3936 LSTM-VAE 0.9624 0.5991 0.7384 USAD 0.9851 0.6618 0.7917 OmniAnomaly 0.9825 0.6497 0.7822 BeatGAN 0.9897 0.6374 0.7754 GDN 0.9935 0.6812 0.8082 TranAD 0.9760 0.6997 0.8119 Ours 0.9863 0.7256 0.8361
[0167] Table 3 Anomaly detection results of the proposed method and seven baseline models on the WADI dataset;
[0168] method Precision Recall F1 DAGMM 0.5444 0.2699 0.3609 LSTM-VAE 0.8779 0.1445 0.2482 USAD 0.6451 0.3220 0.4296 OmniAnomaly 0.9947 0.1298 0.2296 BeatGAN 0.4144 0.3392 0.3730 GDN 0.9750 0.4019 0.5692 TranAD 0.3529 0.8296 0.4952 Ours 0.9837 0.5315 0.6901
[0169] On the SWaT dataset, in terms of precision, the method of the present invention outperforms all other baseline models except BeatGAN and GDN. In terms of recall, the method of the present invention outperforms all other baseline models. In terms of F1 score, the method of the present invention also outperforms all other baseline models. On the WADI dataset, in terms of precision, the method of the present invention outperforms all other baseline models except OmniAnomaly. In terms of recall, the method of the present invention outperforms all other baseline models except TranAD. In terms of F1 score, the method of the present invention outperforms all other baseline models. In summary, the modeling method proposed in the present invention achieved the best results in F1 scores on both datasets.
[0170] The traditional DAGMM method performed poorly on both datasets. This is because DAGMM fails to consider spatiotemporal correlations, meaning it fails to extract both spatial and temporal features. While LSTM-VAE uses LSTM to extract long-term temporal features, it ignores spatial correlations, so its performance lags behind other methods. The WADI dataset involves more automated components and has more complex spatial correlations, so LSTM-VAE performs worse on the WADI dataset than on the SWaT dataset. USAD and OmniAnomaly capture temporal features within time series and perform well on the SWaT dataset. However, because they fail to consider spatial correlations, their performance is poor in more complex scenarios like WADI. GDN uses a graph attention network to learn spatial correlations between variables, but it is not adept at extracting temporal features from time series. TranAD uses a Transformer to extract temporal features, but the Transformer is not adept at extracting long-term temporal features.
[0171] The method proposed in the present invention has the best effect. This is because the present invention first combines industrial process knowledge to construct a graph structure and uses correlation calculation to correct the graph structure. The constructed graph structure is more in line with the actual scenario, and the spatial features extracted subsequently are more accurate. Secondly, this method uses a multi-head attention mechanism to capture richer spatial features from different feature spaces. GRU and Informer extract short-term time series features and long-term time series features, respectively, obtaining more complete time series features from two aspects. In addition, this method combines the advantages of reconstruction and prediction to effectively capture the distribution characteristics and time series trends of normal data.
[0172] To demonstrate the importance of each component used in this paper, we designed an ablation experiment to observe how the model's performance degrades by excluding or replacing components. The details are as follows:
[0173] (1) Replace Spatial (RS) module, which constructs the graph structure by embedding a random vector for each node, calculating the similarity, and then selecting the TopK method;
[0174] (2) Replace Multi-Head Attention (RMHA) and use the ordinary self-attention mechanism;
[0175] (3) Remove GRU (RG) and only use Informer to extract long-term temporal features;
[0176] (4) Exclude Informer (Remove Informer, RI1) and only use GRU to extract short-term time series features;
[0177] (5) Replace Informer (RI2) and use Transformer;
[0178] (6) Remove VAE (RV) from the reconstruction module and only use the prediction module;
[0179] The experimental results are shown in the following table:
[0180] Table 4. SWaT anomaly detection ablation experiment results;
[0181] method Precision Recall F1 Ours 0.9863 0.7256 0.8361 RS 0.9136 0.5943 0.7201 RMHA 0.9523 0.6247 0.7545 RG 0.9385 0.6281 0.7525 RI1 0.9124 0.6037 0.7266 RI2 0.8521 0.5462 0.6657 RV 0.9617 0.6824 0.7983
[0182] Table 5. WADI anomaly detection ablation experiment results;
[0183] method Precision Recall F1 Ours 0.9837 0.5315 0.6901 RS 0.9011 0.4503 0.6005 RMHA 0.9691 0.5166 0.6740 RG 0.9428 0.5184 0.6670 RI1 0.9237 0.5034 0.6517 RI2 0.8933 0.4856 0.6292 RV 0.9712 0.5278 0.6840
[0184] Experimental results show that when the spatial relationship building module is replaced with a standard TopK strategy, precision, recall, and F1 scores on both datasets decrease significantly. This demonstrates that the proposed process knowledge space construction method can better simulate the spatial characteristics of industrial processes. When multi-head attention is replaced with standard self-attention, all metrics also decrease, demonstrating that multi-head attention can better capture spatial features. Removing the GRU also reduces all three metrics to a certain extent, demonstrating the effectiveness of using the GRU in extracting short-term temporal features. Removing the Informer also reduces all three metrics, demonstrating that the Informer can effectively extract long-term temporal features. Replacing the Informer with the Tranformer results in a significant decrease in all metrics, indicating that the Tranformer is not suitable for extracting long-term temporal features in this task. Finally, removing the VAE from the reconstruction module also results in a slight decrease in each metric, demonstrating that joint optimization using the reconstruction module can improve model performance.
[0185] This embodiment further provides an abnormal flow detection system based on industrial process spatiotemporal modeling, which is used to implement an abnormal flow detection method based on industrial process spatiotemporal modeling, including:
[0186] Industrial process data acquisition module, used to obtain industrial process data;
[0187] Data preprocessing module, used to preprocess industrial process data to obtain window data;
[0188] Anomaly detection model, which is used to calculate anomaly scores based on the input window data;
[0189] The anomaly recognition module is used to determine whether an anomaly has occurred based on the anomaly score.
[0190] This embodiment also provides an electronic device, comprising: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus, and when the machine-readable instructions are executed by the processor, the steps of the abnormal flow detection method based on spatiotemporal modeling of industrial processes are performed.
[0191] This embodiment further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the abnormal flow detection method based on spatiotemporal modeling of industrial processes are executed.
[0192] The above preferred embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable technicians in this field to understand the content of the present invention and implement it. It cannot be used to limit the scope of protection of the present invention. Any equivalent changes or modifications made according to the essence of the present invention fall within the scope of protection of the present invention.
Claims
1. A method for detecting abnormal flow based on spatiotemporal modeling of industrial processes, characterized in that: The specific steps include: Acquire industrial process data and preprocess the industrial process data, including data normalization and data windowing, to obtain windowed serialized data; the industrial process data is a multivariate time series including all sensor data and actuator data collected in the industrial process; Constructing an anomaly detection model; the anomaly detection model is used to calculate an anomaly score based on industrial process data within an input time window; Use the window serialization data to train the anomaly detection model to obtain a trained anomaly detection model; Acquire and preprocess industrial process data, input the preprocessed industrial process data within a time window into the trained anomaly detection model to obtain an anomaly score, and detect whether an anomaly has occurred based on the anomaly score; The anomaly detection model includes a spatial relationship construction module, a spatial feature extraction module, a time series feature extraction module, a time series prediction module, a reconstruction module and a scoring module; The spatial relationship construction module is used to construct a graph structure based on the acquired industrial process knowledge and modify the graph structure based on the input window data to obtain a directed graph G; the industrial process knowledge includes the spatial relationship between various automation components in the industrial process; the window data is the industrial process data within a time window; The spatial feature extraction module is used to obtain the spatial features of each node in the directed graph G; the spatial feature extraction module is a graph attention network; The temporal feature extraction module is used to extract short-term temporal features using the GRU model and long-term temporal features using the Informer model based on the spatial features of all nodes; The time series prediction module is used to perform predictions based on short-term time series features and long-term time series features, respectively, to obtain short-term time series prediction results and long-term time series prediction results; The reconstruction module is used to reconstruct the short-term time series features and obtain the reconstruction probability of the short-term time series features; a variational autoencoder (VAE) is used as the reconstruction module; The scoring module is used to calculate the anomaly score based on the short-term time series prediction results, the long-term time series prediction results and the reconstruction probability of the short-term time series features obtained by the reconstruction module; The method for constructing the graph structure is specifically as follows: each automation component is regarded as a node and represented by a binary tuple A=(c,r), where c represents the type of the automation component, 0 represents a sensor, 1 represents an actuator, and r represents the number of the sub-process to which the automation component belongs in the entire industrial process; then, the spatial relationship of the automation components included in the industrial process knowledge is used to determine the automation components that have an impact on each other, and the impact relationship between the automation components is represented by directed edges. A graph structure is constructed and saved using an adjacency matrix, and an edge set E is obtained at the same time; then, the edges in the constructed graph structure are verified, and the edges in the edge set E are corrected using the Pearson correlation coefficient and transfer entropy, respectively, to obtain a directed graph G; The process of modifying the edges in the edge set E using the Pearson correlation coefficient and the transfer entropy is specifically as follows: Extract the time series data of different automation components from the input window data, calculate the Pearson correlation coefficient and transfer entropy between the time series data of different automation components, and set a Pearson correlation coefficient threshold θ p and a transfer entropy threshold θ t , if the absolute value of the Pearson correlation coefficient between the time series data of two automation components is greater than or equal to the Pearson correlation coefficient threshold θ p , then retain the edge between the two automation components in the graph structure if the absolute value of the Pearson correlation coefficient between the time series data of the two automation components is less than the Pearson correlation coefficient threshold θ p , then further determine whether the transfer entropy between the time series data of the two automation components is greater than or equal to the transfer entropy threshold θ t If yes, the edge between the two automation components in the graph structure is retained; otherwise, the edge between the two automation components in the graph structure is deleted; The method for obtaining the spatial features of each node in the directed graph G is as follows: first, the features of each node in the directed graph G are initialized to obtain the initial spatial features of each node; a linear layer is used to extract the time series data of different automation components from the input window data. Convert to dimensional vector, and obtain the initial spatial characteristics of the node corresponding to the automation component ,in is the dimension of the initial spatial feature, Indicates the number of the automation component; Then, the multi-head attention mechanism of the graph attention network is used to update the initial spatial features of each node to obtain the final spatial features of each node; The time series prediction module uses the short-term time series features output by the GRU model as Elements and the corresponding initial spatial features Multiply them one by one, and then input the results into the fully connected layer to perform short-term sequence prediction to obtain the short-term time series prediction results; Long-term time series characteristics Input the fully connected layer and get arrive A vector of predicted values of data of all automation components at the moment , that is, the long-term time series prediction result, The time step length for the predicted future industrial process data; The abnormality scores are: (19); in, For abnormality score, is a hyperparameter, for The short-term time series prediction results and the long-term time series prediction results at the moment The average value of the predicted values of automation component data, for Moment The true value of the data of the automation component, Short-term time series characteristics No. The reconstruction probability of an element.
2. The abnormal flow detection method based on spatiotemporal modeling of industrial processes according to claim 1 is characterized in that: The industrial process data is obtained and preprocessed, including data normalization and data windowing, to obtain window serialized data, specifically: A1: Acquire industrial process data; the industrial process data is , Indicates The collection of all sensor data and actuator data collected at all times, Indicates the moment, Represents the length of industrial process data, represents the total number of automation components, including sensors and actuators; A2: Perform normalization processing on the industrial process data to obtain normalized industrial process data; A3: Convert the normalized industrial process data into window data, and then obtain window serialized data including multiple window data; For the current moment , define the length as The window data is , including from Time has come Normalized industrial process data at time After normalization The collection of all sensor data and actuator data collected at all times; the window serialized data is represented as , including several lengths of Window data.
3. The abnormal flow detection method based on spatiotemporal modeling of industrial processes according to claim 1 is characterized in that: When the anomaly detection model is trained using window serialization data, the construction process of the joint loss function is: Use mean squared error as the loss function for short-term time series forecasting: (21); in, is the loss function for short-term time series prediction, for The short-term time series prediction results of The predicted value of data of an automation component; Use mean squared error as the loss function for long-term time series forecasting: (22); in, is the loss function for long-term time series prediction, is a positive integer, exist Moment The true value of the data of each automation component; Long-term time series prediction results middle Moment The predicted value of data of an automation component; The reconstruction loss can be defined as: (23); in, is the reconstruction loss, is the reconstruction error term, is to calculate the KL divergence, is a potential feature, For a given latent feature Generate short-term time series features The probability of For a scalable distribution, for The prior probability of The joint loss function is defined as: (24); in, is the joint loss function, These are hyperparameters that need to be set manually.
4. An abnormal flow detection system based on spatiotemporal modeling of industrial processes, used to implement the abnormal flow detection method based on spatiotemporal modeling of industrial processes according to any one of claims 1 to 3, characterized in that: include: Industrial process data acquisition module, used to obtain industrial process data; Data preprocessing module, used to preprocess industrial process data to obtain window data; Anomaly detection model, which is used to calculate anomaly scores based on the input window data; The anomaly recognition module is used to determine whether an anomaly has occurred based on the anomaly score.
5. An electronic device, characterized in that: include: A processor, a memory and a bus, wherein the memory stores machine-readable instructions executable by the processor. When the electronic device is running, the processor and the memory communicate via the bus. When the machine-readable instructions are executed by the processor, the steps of the abnormal flow detection method based on spatiotemporal modeling of industrial processes as described in any one of claims 1 to 3 are performed.
6. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, executes the steps of the abnormal flow detection method based on spatiotemporal modeling of industrial processes as described in any one of claims 1 to 3.