A method for detecting data anomalies in a sugar factory based on a graph structure mask autoencoder
Patent Information
- Application Number
- CN202211038693.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-29
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2042-08-29
AI Technical Summary
Existing technologies make it difficult to effectively use deep learning models to detect anomalies in high-dimensional time series data in sugar factories. In particular, graph neural networks are insufficient in capturing complex relationships and temporal context, making anomaly diagnosis and location difficult.
A graph-structured masked autoencoder method is adopted. By constructing a causal directed graph and a graph attention network, combined with the Transformer algorithm, an autoencoder model is built. The KL loss is used to train the model, and the moment-level reconstruction loss is calculated for anomaly detection and localization.
It achieves accurate anomaly diagnosis and location of sugar factory data, improves the effect of anomaly detection, and can detect anomalies and distinguish normal data from abnormal data in a short time.
Smart Images

Figure CN115456055B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data anomaly detection in the production process of a sugar refinery, and in particular to a method for detecting data anomaly in a sugar refinery based on a graph structure masked autoencoder. Background Art
[0002] In modern process industries, sensor technology and information systems are widely used for data collection and storage. Time series anomaly detection can provide rich insights into catastrophic events such as sensor failures in process industries, which is crucial for maintaining normal system operation. Unsupervised time series anomaly detection is extremely challenging in practice. This method typically learns information representations from complex temporal dynamics through unsupervised tasks. Classic unsupervised methods include statistical, probability, similarity, and prediction-based methods. Classical methods remain a viable option most of the time. However, as the target system becomes larger and more complex, these methods are limited in their inability to operate on multidimensional data.
[0003] Thanks to the representational learning capabilities of neural networks, deep learning can be applied to anomaly detection in high-dimensional time series data. Generally speaking, DL-based time series anomaly detection can be roughly divided into prediction-based models and reconstruction-based models. Long Short-Term Memory (LSTM) networks can determine whether an anomaly has occurred based on the difference between actual measurements and predicted values. However, prediction-based models are not suitable for unpredictable time series and their application is limited by the high cost of supervised labeled data. Among reconstruction-based models, recurrent networks demonstrate excellent performance in capturing the temporal characteristics of data. Variational autoencoders based on recurrent networks detect anomalies based on reconstruction error. However, due to their structural characteristics, recurrent neural networks are plagued by vanishing / exploding gradient problems. Furthermore, the reconstruction error is calculated point by point, failing to provide a comprehensive description of the temporal context. Clearly, there is still a significant gap between deep learning models and practical applications.
[0004] Graph neural networks (GNNs), a branch of deep learning, are becoming increasingly popular and have demonstrated powerful learning capabilities across multiple domains. Graphs can be used to extract a rich variety of real-world relationships between entities. A graph is typically defined by two sets: vertices and edges. Vertices represent entities in the graph, while edges represent relationships between these entities. As graphs leverage both basic and relevant relationships between vertices, GNNs are becoming increasingly popular for capturing complex relationships. As a variant of traditional GNNs, graph convolutional networks (GCNs) exhibit permutation invariance, local connectivity, and compositionality. GCNs allow each graph node to identify its neighborhood context through structural propagation of information. However, while extensive research has been conducted on anomaly occurrence, research on locating anomalies by leveraging the powerful temporal properties and variable relationship modeling capabilities of graph networks and transformers remains scarce. Summary of the Invention
[0005] In view of the above problems existing in the prior art, the purpose of the present invention is to provide a sugar factory data anomaly detection method based on a graph structure masked autoencoder to achieve diagnosis and location of anomalies.
[0006] The technical solution adopted by the present invention to solve its technical problem is:
[0007] A method for detecting anomaly in sugar factory data based on a graph structure masked autoencoder, the method comprising the following steps:
[0008] 1) Data acquisition and integration:
[0009] By using the DMDS sugar factory dataset, each data file contains data obtained from the DMDS sugar factory in one day, and clearly indicates the abnormal time and type of data, the dataset is divided, collected, and integrated.
[0010] 2) Establish a causal directed graph:
[0011] Based on the sugar factory distiller flow chart, a causal directed graph is constructed between its variables.
[0012] 3) Modeling training:
[0013] A data anomaly detection model based on graph structure masked autoencoder is constructed and trained using KL loss.
[0014] 4) Anomaly Detection
[0015] The reconstruction loss of the data at each moment between the model reconstructed data and the input data is calculated to complete the detection and location of anomalies.
[0016] Further, the process of step 1) is:
[0017] Step 1.1: Select the data of a sugar factory over a period of time as the data set for testing.
[0018] Step 1.2: Take out the abnormal data from the dataset and use it as the test set, and use the remaining normal data as the training set.
[0019] Step 1.3: In order to detect anomalies in the shortest possible time, set a sliding window and select the length and step size of the sliding window.
[0020] Step 1.4: Since the data differences between different feature variables are large, it is necessary to standardize the data to facilitate model processing and calculation. The specific formula is as follows:
[0021]
[0022] Among them, X' is the data after standardization, X is the original data without standardization, μ is the mean of the data, and γ is the standard deviation of the data.
[0023] Furthermore, the process of step (2) is as follows:
[0024] Step 2.1: Find out the required units and create nodes for each unit. These nodes are called unit nodes.
[0025] Step 2.2: Find all the corresponding variables of the cells in the first step and create nodes for each variable; these nodes are called variable nodes.
[0026] Step 2.3: Create a directed edge from each cell node to the corresponding variable node.
[0027] Step 2.4: Divide the variable nodes into process nodes and manipulation nodes according to the corresponding control loop of each unit; add arrows between the process nodes and manipulation nodes according to the control loop.
[0028] Step 2.5: Delete the independent nodes in the graph to simplify the variable topology graph; Figure 1 The topological graph of the selected variables.
[0029] Further, the process of step 3) is:
[0030] Step 3.1: The masked autoencoder model based on the graph structure is mainly based on the autoencoder model built on the graph attention network and the Transformer algorithm. The input of the graph attention layer is a set of node features. For the graph G, X={x1,x2,…,x V}, Where V is the number of nodes and T is the number of features in each node. In order to obtain sufficient expressive power, at least one learnable linear transformation is required. To do this, a weight matrix Perform a linear transformation on each node, where F represents the dimension after the linear transformation. Then perform the self-attention mechanism a on the node: To calculate the attention coefficient e ij :
[0031] e ij =a(x i W,x j W) (1)
[0032] e ij Indicates the importance of the feature of node j to node i; this formula allows each node to pay attention to all other nodes, so the graph structure is incorporated into the mechanism through masking, and only the nodes e ij ,in is the neighborhood of node i in the graph. To make the coefficients between different nodes easier to compare, all choices of j are normalized using the softmax function, as follows:
[0033]
[0034] In order to stabilize the learning process of self-attention, it is beneficial to adopt multi-head attention. Specifically, the results of h independent attention mechanisms are averaged to produce the following output representation:
[0035]
[0036] Among them, σ(.) is the LeakyReLU activation function.
[0037] At its core, the Transformer is a multi-head attention mechanism, where multiple heads can learn relevant information in different representation subspaces. The attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, key, value, and output are all vectors. In multi-head attention, the query Q, key K, and value V are each mapped to different linear spaces through a fully connected FC layer. Each projection is calculated in parallel to obtain the correlation between any two moments. Then, all attention values are concatenated together:
[0038] h-Head(Q,K,V)=FC(Concat(head1,…,head h )) (4)
[0039]
[0040]
[0041] Among them, h represents the number of heads, α ij Represents the attention coefficient between two nodes, W h Q、W h K and W h V Project Q, K and V onto h respectively th The mapping weight of the linear space; B is the scale factor, which represents the data time step; Figure 2 This is the structure of the multi-head attention mechanism. After Q, K, and V are linearly transformed, their long-term dependencies are constructed using the multi-head attention mechanism.
[0042] Step 3.2: By stacking the attention encoding layer and the attention decoding layer, an asymmetric autoencoder structure is obtained; by restricting the multi-head attention mechanism through the multi-head variable attention map, the encoder can be forced to learn the intrinsic features under the relational constraints; Figure 2 It is a masked autoencoder framework based on graph structure.
[0043] Further, the process of step 4) is:
[0044] Step 4.1: Construct variable topology map A based on industrial processes and industrial data T and variable learning graph A L , and take the variable topology graph as the main factor, combine the variable learning graph to generate the variable relationship graph A G At the same time, the industrial data is divided into different tokens through the patch and mask mechanism and 50% of the tokens are randomly masked to obtain masked data. The masked data is embedded in the position and the graph attention mechanism is used to calculate the multi-head variable attention graph A. A ; In the multi-head variable attention map A A The multi-head attention mechanism is restricted based on ; by stacking the encoding attention layer, the mask data X is finally obtained M The intrinsic embedding of d is the embedding dimension and c is the number of tokens.
[0045] Step 4.2: Flatten the intrinsic embedding to obtain Randomly initialize a mask vector And copy V×c / 2 times in the first dimension to get the masked data The intrinsic embedding and the masked data are concatenated according to the first dimension, and then the multi-head attention mechanism with multiple layers of attention decoding layers is stacked to decode the intrinsic embedding; the model is trained with KL loss to make the distribution of the reconstructed data consistent with the input data; finally, the moment-level reconstruction loss between the reconstructed data and the input data is calculated to complete the detection and location of anomalies. Figure 2 Heatmap of moment-level reconstruction loss.
[0046] By adopting the above technology, compared with the prior art, the beneficial effects of the present invention are as follows:
[0047] The present invention proposes a sugar factory data anomaly detection method based on a graph structure masked autoencoder. When encountering an anomaly, the masked part can be reconstructed from the normal data, which increases the reconstruction loss at the moment level, indicating that an anomaly exists in the current moment and variable, thereby realizing the diagnosis and location of the anomaly. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a variable topology diagram selected in an embodiment of the present invention;
[0049] Figure 2 The present invention is based on the graph structure masked autoencoder framework;
[0050] Figure 3 is the moment-level reconstruction loss heat map of the present invention; DETAILED DESCRIPTION
[0051] A method for detecting anomaly in sugar factory data based on a graph structure masked autoencoder, the method comprising the following steps:
[0052] (1) Obtain the DMDS sugar factory dataset. The process is as follows:
[0053] Step 1.1: Test the GMAE-based time series anomaly detection method using data from a sugar factory on a single day. The data consists of 86,400 time points and 32 variables, i.e., 86,400 × 32 data.
[0054] Step 1.2: The dataset description clearly states that there were 10 time series anomalies that day: three positioner supply pressure drops, two valve unexpected pressure changes, one bypass valve fully or partially open, and four flow sensor failures. After removing the anomaly data from the dataset, we obtained 81,600 × 32 normal data as the training set and 4,800 × 32 abnormal data as the test set.
[0055] Step 1.3: Each anomaly affects only 2-5 variables and lasts between 14-100 seconds. To detect anomalies as quickly as possible, a sliding window of length 200 and step size 10 is set.
[0056] Step 1.4: Process each data according to the standardization formula.
[0057] (2) Draw a variable causal directed graph. The process is as follows:
[0058] Step 2.1: Draw a device-level causal directed graph.
[0059] Step 2.2: Based on the device-level causal directed graph, draw the causal directed graph of all variables.
[0060] (3) Training the anomaly detection model. The process is as follows:
[0061] Step 3.1: According to the causal directed graph, set the number of nodes of the simplified causal directed graph to 32 and simplify the causal directed graph. Table 1 shows the subordinate relationship between the nodes and the simplified nodes.
[0062] Table 1 Affiliation relationships between nodes
[0063]
[0064] Step 3.2: According to the causal directed graph and the simplified causal directed graph, the training set is input into the model to obtain the reconstruction result.
[0065] Step 3.3: Based on the reconstruction results and input data, adjust the model parameters so that the difference between the reconstruction results and the input data is reduced.
[0066] Step 3.4: Repeat steps 3.2 to 3.3 until the reconstruction error of the model is within the allowable error.
[0067] (4) Use the test set data to perform data anomaly detection. The process is as follows:
[0068] Step 4.1: Input the test machine data into the model and calculate the moment-level reconstruction loss between the reconstruction result and the input data.
[0069] Step 4.2: For the moment-level reconstruction loss, points within 3 standard deviations of the mean are considered normal, while points outside 3 standard deviations are considered abnormal.
[0070] The method of the present invention adopts a data anomaly detection method based on a graph structure masked autoencoder, which improves the effect of anomaly detection, can accurately find abnormal time and abnormal variables, and has universality and versatility.
[0071] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The scope of protection of the present invention should not be regarded as limited to the specific forms described in the embodiments. The scope of protection of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. A sugar factory data anomaly detection method based on graph structure masked autoencoder, characterized by The method comprises the following steps: 1) Data acquisition and integration: extract normal data as the training set, extract abnormal data as the test set, and set a sliding window; 2) Establish a causal directed graph: Based on the production flow chart, construct a causal directed graph between its variables; 3) Modeling and Training: Build a data anomaly detection model based on a graph-structured masked autoencoder. The specific process of step 3) is as follows: Step 3.1: Build an anomaly detection model based on graph attention and Transformer; for graph G, Where V is the number of nodes and T is the number of features in each node; using the weight matrix Perform linear transformation on each node, F represents the dimension after linear transformation; perform self-attention mechanism on the node To calculate the attention coefficient e ij : e ij =a(x i W,x j W) (1) e ij It shows the importance of the feature of node j to node i; this formula allows each node to pay attention to all other nodes, and integrates the graph structure into the mechanism through masking, only calculating the node e ij ,in is the neighborhood of node i in the graph; in order to make the coefficients between different nodes easier to compare, all choices of j are normalized using the softmax function, as follows: Among them, σ(.) is the LeakyReLU activation function; Among them, h represents the number of heads, α ij Represents the attention coefficient between two nodes, W h Q 、W h K and W h V Project Q, K and V onto h respectively th The mapping weight of the linear space; B is the scaling factor, which represents the data time step; after Q, K, and V are linearly transformed, their long-term dependencies are constructed using the multi-head attention mechanism; Step 3.2: Stack the layers in step 3.1 sequentially to obtain an asymmetric autoencoder structure; 4) Detect anomalies: Calculate the difference between the reconstruction result obtained by inputting the test set into the model and the input data, and determine whether the input data is abnormal. The specific process of step 4) is as follows: Step 4.1: Create variable topology map A T and variable learning graph A L , and take the variable topology graph as the main factor, combine the variable learning graph to generate the variable relationship graph A G ; Divide the data into different tokens and randomly mask 50% of the tokens to obtain masked data d is the embedding dimension, c is the number of tokens; Step 4.2: Flatten the intrinsic embedding to obtain Randomly initialize a mask vector And copy V×c / 2 times in the first dimension to get the masked data The intrinsic embedding is concatenated with the masked data along the first dimension, and then a multi-head attention mechanism with multiple layers of attention decoding layers is used to decode the intrinsic embedding. Step 4.3: Train the model using KL loss to make the distribution of the reconstructed data consistent with the input data; finally, calculate the moment-level reconstruction loss between the reconstructed data and the input data to complete the detection and location of anomalies.
2. The sugar factory data anomaly detection method based on graph structure masked autoencoder according to claim 1 is characterized in that: The specific process of step 1) is as follows: Step 1.1: Select the data of a sugar factory over a period of time as the data set for testing; Step 1.2: Take out the abnormal data in the dataset and use it as the test set, and use the remaining normal data as the training set; Step 1.3: Set the sliding window and select the length and step size of the sliding window; Step 1.4: Standardize the acquired data to facilitate model processing and calculation. The specific formula is as follows: Among them, X' is the data after standardization, X is the original data without standardization, μ is the mean of the data, and γ is the standard deviation of the data.