Industrial process fault detection method based on space-time causal graph auto-encoder
By constructing a spatiotemporal causal graph autoencoder, utilizing spatial self-attention mechanism and graph convolutional long short-term memory network, and combining the principle of causal invariance and process mechanism, the problem of difficulty in characterizing causal relationships in existing technologies is solved, and high reliability and high interpretability of industrial process fault detection are achieved.
Patent Information
- Application Number
- CN202511505237.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-21
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-10-21
AI Technical Summary
Existing industrial process fault detection methods based on graph neural networks cannot effectively characterize the causal relationships between variables, leading to a decline in fault detection performance. Furthermore, existing causal discovery algorithms perform poorly in cases of multidimensional variable coupling, and are difficult to interpret and lack reliability.
A spatiotemporal causal graph autoencoder-based approach is adopted, which constructs a causal graph through a spatial self-attention mechanism and a graph convolutional long short-term memory network. Combined with the principle of causal invariance and knowledge of process mechanisms, the adaptive learning and fault detection of the causal graph are realized.
It improves the reliability and interpretability of industrial process fault detection, effectively distinguishes between faults and normal operating conditions, reduces false alarm rates, and enhances the accuracy of fault detection and the interpretability of the model.
Smart Images

Figure CN120974245A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of process monitoring and fault diagnosis, specifically relating to an industrial process fault detection method based on a spatiotemporal causal graph autoencoder. Background Technology
[0002] With the development of deep learning and artificial intelligence, data-driven multivariate statistical process monitoring methods have become powerful tools for monitoring complex industrial processes, ensuring operational safety and production efficiency. However, these methods suffer from low reliability and poor interpretability, leading to a significant gap between laboratory outputs and industrial applications. For example, monitoring model performance degrades under changing operating conditions, resulting in high false alarm rates; the decision-making logic of data-driven black-box models hinders trust from on-site personnel and makes it difficult to pinpoint the root cause of faults. Graph Neural Networks (GNNs) can utilize neural networks to process graph-structured data. In a real industrial process, variables are often closely related. By representing the relationships between variables as a graph, fault detection models based on GNNs can enhance the interpretability of process monitoring. However, existing technologies have two limitations. First, most GNN-based fault detection methods often build correlation graphs between variables, only capturing correlations and failing to characterize causal relationships. Correlation becomes invalid with fluctuations in operating conditions, resulting in poor generalization ability and decreased fault detection performance. In contrast, fault detection models based on causal graphs have significant advantages in reliability and interpretability. Secondly, existing industrial process causal discovery algorithms rely on transfer entropy and Granger causal analysis methods, without considering confounding factors caused by multidimensional variable coupling, resulting in poor causal discovery performance and the established causal graphs containing a large number of spurious correlations. Summary of the Invention
[0003] Objective of the Invention: The technical problem to be solved by the present invention is to address the shortcomings of existing technologies by providing an industrial process fault detection method based on a spatiotemporal causal graph autoencoder, comprising the following steps: Step 1: For the target industrial process, preprocess the spatiotemporal process data of all process variables collected during the operation of the target industrial process to obtain a standardized time series data for each process variable; generate dynamic time series data through a fixed-length sliding window to form standardized dynamic input time series data. Step 2: Establish a causal graph spatiotemporal autoencoder (CGSTAE). The causal graph spatiotemporal autoencoder (CGSTAE) includes a correlation graph structure learning module based on the spatial self-attention mechanism (SSAM) and a spatiotemporal encoder-decoder module based on the graph convolutional long short-term memory network (GCLSTM). Step 3: Execute the three-step causal graph structure learning algorithm to train the causal graph spatiotemporal autoencoder CGSTAE, including three steps: pre-training, causal extraction, and fine-tuning. Step 4: Based on the hidden layer features of the causal graph spatiotemporal autoencoder CGSTAE and the residual data of the reconstructed output, calculate the feature space statistics and residual space statistics respectively. Determine the control limit at a given significance level by kernel density estimation. Samples with statistics exceeding the control limit are regarded as fault samples, and the fault detection results are obtained.
[0004] In step 1, the data preprocessing includes outlier removal, missing value imputation, and max-min standardization. The spatiotemporal process data of all process variables are processed into the following form: X represents the training data consisting of N normal samples, and T represents the matrix transpose. Let X represent the t-th normal sample; let the i-th column of X be denoted as . , This represents the time-series data vector of the i-th process variable. This represents the measurement data of the i-th process variable at time t; By reconstructing the data using a sliding window, the time series data matrix of the input model at time t is obtained. , where w represents the length of the sliding window.
[0005] In step 2, the causal graph is defined as a directed unweighted graph. ,in It is a set of nodes. It is a set of directed edges; nodes It is the i-th variable in a given process, and its node attribute is a standardized time-series data vector. Directed edge Representing variables and variables The causal dependency between them.
[0006] In step 2, the preprocessed time-series data is input into the Spatial Self-Attention Mechanism (SSAM). Attention weights between variables are obtained by calculating the query matrix and key-value matrix, forming the adjacency matrix of a dynamic correlation graph, thus achieving adaptive learning of time-varying correlations between process variables. Using the graph structure as input, a spatiotemporal encoder-decoder module based on a graph convolutional long short-term memory (GCLSTM) network is constructed. The reconstructed values of the process data are obtained using the hidden features and fully connected layer mapping of the spatiotemporal encoder-decoder module, achieving topology-guided spatiotemporal dynamic modeling of industrial processes. The Spatial Self-Attention Mechanism (SSAM) then processes the time-series data matrix... Convert to attention matrix : , The query matrix Key-value matrix , and These are two trainable weight matrices for the Spatial Self-Attention Mechanism (SSAM). Use the Sigmoid activation function; Since the spatial self-attention mechanism (SSAM) can model the correlation between changes in variables, the attention matrix can be used... It is considered as an adjacency matrix of a dynamic correlation graph.
[0007] In step 2, a spatiotemporal encoder-decoder module based on a graph convolutional long short-term memory network (GCLSTM) is constructed. The encoder in the spatiotemporal encoder-decoder module uses a temporal data matrix. Attention Matrix As input, along the forward timing Circular execution of a graph convolutional long short-term memory network (GCLSTM) to update the forget gate at time k. Input gate Output gate Candidate status Unit status and hidden layer features ,in: , , , , , , Where tanh represents the hyperbolic tangent function. Represents the Hadama product. , , , These represent the graph convolutional layers of the encoder: the forget gate, the input gate, the output gate, and the candidate states, respectively. , , , These represent the learnable bias parameters of the forget gate, input gate, output gate, and candidate state in the encoder, respectively. The encoder in the spatiotemporal encoder-decoder module extracts spatiotemporal information by updating the unit state and hidden features through a graph convolutional long short-term memory (GCLSTM) network, and then passes this information to the decoder. The decoder proceeds along the reverse temporal sequence. The GCLSTM graph convolutional long short-term memory network is executed in a loop to update the unit states and hidden features: , , , , , , The calculation method for GC in graph convolutional layers is as follows: , Where D is the adjacency matrix. The degree matrix, I is the degree matrix with Unit arrays of the same shape For trainable weight parameters, Z represents the node features of the input graph convolutional layer: in the encoder In the decoder ; Finally, by utilizing the hidden layer features, reconstructed values of the process data at time k are generated through a fully connected layer. : , in and These are the trainable weight parameters and bias parameters of the fully connected layer, respectively.
[0008] In step 3, firstly, dynamic correlation maps under different operating conditions are generated using the pre-trained spatial self-attention mechanism SSAM; then, based on the principle of causal invariance, stable structures are extracted from the dynamic correlation maps, and causal graphs are generated by integrating process mechanism knowledge; finally, the learned causal graphs are used to fine-tune the parameters of the spatiotemporal encoder decoder module, establishing a causal-based fault detection model; the trainable model parameters in the causal graph spatiotemporal autoencoder CGSTAE are used... It means that among them These represent the parameters of the correlation graph structure learning module. The parameters of the spatiotemporal encoder decoder module are represented by: and This represents the function mappings for the correlation graph structure learning module and the spatiotemporal encoder-decoder module.
[0009] In step 3, the pre-training includes: minimizing the mean square error (MSE) of the reconstructed values by jointly optimizing the parameters of the correlation graph structure learning module and the spatiotemporal encoder-decoder module. , in express and The mean square error; Reconstructed value The correlation graph structure learning module and the spatiotemporal encoder decoder module are used to calculate the dynamic correlation between industrial process variables, enabling the spatial self-attention mechanism (SSAM) to adaptively learn the dynamic correlation between industrial process variables. The causal extraction includes: freezing the parameters of the pre-trained spatiotemporal encoder-decoder module and the correlation graph structure learning module, introducing a trainable causal graph adjacency matrix A, and learning the causal graph adjacency matrix A by optimizing the following objectives: , in , , , Indicates the weighting coefficient; The term ensures that the reconstruction capability of the causal graph is consistent with that of the correlation graph by constraining the root mean square error (MSE). The calculation method is as follows: , The term extracts causal relationships based on invariance, and the calculation method is as follows: , in, This represents the element in the i-th row and j-th column of the adjacency matrix A of a trainable causal graph. Adjacency matrix representing a dynamic correlation graph The element in the i-th row and j-th column; The term guarantees that the learned cause-effect graph conforms to the prior knowledge of the process, and the calculation method is as follows: , Where M is the mask matrix representing whether knowledge exists in the representation process, if ,but ;like , "Unknown" indicates that the process knowledge is unknown. This represents the element in the i-th row and j-th column of the mask matrix M; Item and The term is used to improve the sparsity and discreteness of the causal graph, and is calculated as follows: , , The fine-tuning includes: fine-tuning the parameters of the spatiotemporal encoder-decoder module based on the mean square error loss and using the learned causal graph, with the loss function... for: .
[0010] Step 4 includes: performing fault detection, which consists of two phases: offline modeling and online monitoring. Offline modeling stage: The causal graph spatiotemporal autoencoder CGSTAE is trained based on normal training samples, and the feature space statistics T² and residual space statistics SPE are constructed. The feature space statistics at time t are... The calculation formula is: , in Σ and Σ are the mean and covariance matrices of the hidden layer features of normal samples, respectively; It is the final hidden layer feature output by the encoder in the spatiotemporal encoder decoder module; Residual spatial statistics at time t The calculation formula is: , Control limits of feature space statistics Control limits for residual space statistics Determined at a given significance level by kernel density estimation; Online monitoring phase: at every moment The input matrix is obtained by reorganizing the process data through a sliding window. Spatial statistics are calculated using a trained causal graph spatiotemporal autoencoder (CGSTAE). and residual space statistics ,if and If the process is normal, then the fault is determined; otherwise, the fault is determined.
[0011] The present invention also provides an electronic device, including a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method.
[0012] The present invention also provides a storage medium storing a computer program or instructions that, when the computer program or instructions are run on a computer, execute the steps of the method described.
[0013] The network architecture employed in this method consists of a correlation graph structure learning module and a spatiotemporal encoder-decoder. Based on this, a three-step causal graph structure learning algorithm is provided for training a spatiotemporal causal graph autoencoder. Through three steps—pre-training, causal extraction, and fine-tuning—the algorithm discovers invariant causal graph structures from changing correlations, thereby controlling confounding factors in industrial data. Furthermore, with the help of the spatiotemporal encoder-decoder, causal relationship-based industrial process fault detection is achieved, improving the reliability and interpretability of industrial process monitoring.
[0014] The present invention has the following advantages: First, the industrial process fault detection method based on spatiotemporal causal graph autoencoder of the present invention constructs a graph neural network model based on causal graph. The causal graph describes the causal relationship between process variables, which conforms to the physical mechanism of industrial process, so that the spatiotemporal causal graph autoencoder model has high reliability and interpretability.
[0015] Second, the industrial process fault detection method based on spatiotemporal causal graph autoencoder of the present invention has a network architecture including a correlation graph structure learning module based on spatial self-attention mechanism and a spatiotemporal encoder-decoder module based on graph convolutional long short-term memory network. The correlation graph structure learning module can describe the correlation of changes in process data by learning dynamic graphs, and the spatiotemporal encoder-decoder module can process dynamic graphs in the pre-training stage and causal graphs obtained by causal extraction. This network architecture is conducive to realizing causal discovery and fault detection.
[0016] Third, the industrial process fault detection method based on spatiotemporal causal graph autoencoder of the present invention proposes a three-step causal graph structure learning algorithm, which innovatively utilizes the principle of causal invariance to discover invariant causal graphs from changing correlations. At the same time, it integrates process mechanism knowledge to constrain the causal graph, which can control the confounding factors in industrial data and improve the effectiveness of causal discovery. Attached Figure Description
[0017] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.
[0018] Figure 1 This is a schematic diagram of the network architecture of the Spacetime Causal Graph Autoencoder (CGSTAE).
[0019] Figure 2 This is a schematic diagram of a graph convolutional long short-term memory (GCLSTM) unit of a spatiotemporal encoder.
[0020] Figure 3 This is a flowchart of the fault detection process based on the spatiotemporal causal graph autoencoder CGSTAE.
[0021] Figure 4 This is a process flow diagram of the Tennessee-Eastman process.
[0022] Figure 5 Different methods are used to address the T11 fault in the Tennessee-Eastman process. 2 Statistical graph.
[0023] Figure 6 These are SPE statistics plots for the Tennessee-Eastman process fault 11 using different methods.
[0024] Figure 7It is a graph structure of the Tennessee-Eastman process. Detailed Implementation
[0025] This invention provides a fault detection method based on a spatiotemporal causal graph autoencoder, comprising the following steps S1 to S4, the specific process of which is as follows: Figure 3 As shown below, the implementation method of each step will be described in detail.
[0026] S1. After determining the target industrial process, first understand the process flow to be modeled, clarify the process variables involved in the industrial process modeling, and then collect spatiotemporal process data over a period of time through a distributed control system. For the target industrial process, preprocess the spatiotemporal process data of all process variables collected during the operation of the target industrial process, obtaining a standardized time series data for each process variable. Dynamic time series data is generated through a fixed-length sliding window to form standardized dynamic input time series data. Step S1 is detailed as follows: S11: Perform data preprocessing, including outlier removal based on the 3-sigma principle, missing value imputation based on interpolation, and max-min standardization. The spatiotemporal process data of all process variables will be in the following form after preprocessing: X represents the training data consisting of N normal samples, and T represents the matrix transpose. Let X represent the t-th normal sample; let the i-th column of X be denoted as . , This represents the time-series data vector of the i-th process variable. This represents the measurement data of the i-th process variable at time t; S12: Reassemble the data using a sliding window to obtain the time series data matrix of the input model at time t. , where w represents the length of the sliding window.
[0027] S2. Input the preprocessed time-series data into a spatial self-attention mechanism (SSAM). Calculate the query matrix and key-value matrix to obtain the attention weights between variables, thereby constructing the adjacency matrix of a dynamic correlation graph, achieving adaptive learning of time-varying correlations between process variables. Using the graph structure as input, construct a spatiotemporal encoder-decoder based on a graph convolutional long short-term memory network (GCLSTM). Utilize the hidden layer features of the spatiotemporal encoder-decoder and the mapping of fully connected layers to obtain the reconstructed values of the process data, realizing topology-guided spatiotemporal dynamic modeling of industrial processes. Step S2 is detailed below: S21: Utilizing the Spatial Self-Attention Mechanism (SSAM) to integrate time-series data matrices Convert to attention matrix : , The query matrix Key-value matrix , and These are two trainable weight matrices for the Spatial Self-Attention Mechanism (SSAM). Use the Sigmoid activation function; S22: Construct a spatiotemporal encoder-decoder module based on a graph convolutional long short-term memory (GCLSTM) network. The encoder in the spatiotemporal encoder-decoder module uses a temporal data matrix. Attention Matrix As input, along the forward timing Circular execution of a graph convolutional long short-term memory network (GCLSTM) to update the forget gate at time k. Input gate Output gate Candidate status Unit status and hidden layer features ,in: , , , , , , Where tanh represents the hyperbolic tangent function. Represents the Hadama product. , , , These represent the graph convolutional layers of the encoder: the forget gate, the input gate, the output gate, and the candidate states, respectively. , , , These represent the learnable bias parameters of the forget gate, input gate, output gate, and candidate state in the encoder, respectively. The encoder in the spatiotemporal encoder-decoder module extracts spatiotemporal information by updating the unit state and hidden features through a graph convolutional long short-term memory (GCLSTM) network, and then passes this information to the decoder. The decoder proceeds along the reverse temporal sequence. The GCLSTM graph convolutional long short-term memory network is executed in a loop to update the unit states and hidden features: , , , , , , The calculation method for GC in graph convolutional layers is as follows: , Where D is the adjacency matrix. The degree matrix, I is the degree matrix with Unit arrays of the same shape For trainable weight parameters, Z represents the node features of the input graph convolutional layer: in the encoder In the decoder ; S23: Utilizing hidden layer features, reconstructed values of the process data at time k are generated through a fully connected layer. : , in and These are the trainable weight parameters and bias parameters of the fully connected layer, respectively.
[0028] S3. The three-step causal graph structure learning algorithm is used to train the CGSTAE model, including pre-training, causal extraction, and fine-tuning. First, a dynamic correlation graph under different operating conditions is generated through pre-training of the spatial self-attention mechanism (SSAM). Then, based on the principle of causal invariance, a stable structure is extracted from the dynamic correlation graph, and process mechanism knowledge is integrated to generate a causal graph. Finally, the learned causal graph is used to fine-tune the parameters of the spatiotemporal encoder-decoder module, establishing a causal-based fault detection model. The trainable model parameters in the causal graph spatiotemporal autoencoder CGSTAE are used... It means that among them These represent the parameters of the correlation graph structure learning module. The parameters of the spatiotemporal encoder decoder module are represented by: and This represents the function mapping between the correlation graph structure learning module and the spatiotemporal encoder-decoder module. Step S3 is as follows: S31: Pre-training: Minimize the mean square error (MSE) of the reconstructed values by jointly optimizing the parameters of the correlation graph structure learning module and the spatiotemporal encoder-decoder module. , in express and The mean square error; Reconstructed value The correlation graph structure learning module and the spatiotemporal encoder decoder module are used to calculate the dynamic correlation between industrial process variables, enabling the spatial self-attention mechanism (SSAM) to adaptively learn the dynamic correlation between industrial process variables. S32: Causal Extraction: Freeze the parameters of the pre-trained spatiotemporal encoder-decoder module and the correlation graph structure learning module, introduce a trainable causal graph adjacency matrix A, and learn the causal graph adjacency matrix A by optimizing the following objectives: , in , , , Indicates the weighting coefficient; The term ensures that the reconstruction capability of the causal graph is consistent with that of the correlation graph by constraining the root mean square error (MSE). The calculation method is as follows: , The term extracts causal relationships based on invariance, and the calculation method is as follows: , in, This represents the element in the i-th row and j-th column of the adjacency matrix A of a trainable causal graph. Adjacency matrix representing a dynamic correlation graph The element in the i-th row and j-th column; The term guarantees that the learned cause-effect graph conforms to the prior knowledge of the process, and the calculation method is as follows: , Where M is the mask matrix representing whether knowledge exists in the representation process, if ,but ;like , "Unknown" indicates that the process knowledge is unknown. This represents the element in the i-th row and j-th column of the mask matrix M; Item and The term is used to improve the sparsity and discreteness of the causal graph, and is calculated as follows: , , S33: Fine-tuning: Based on the mean squared error loss, the parameters of the spatiotemporal encoder-decoder module are fine-tuned using the learned causal graph. The loss function... for: .
[0029] S4. Based on the hidden layer features of the model and the residual data of the reconstructed output, calculate the feature space statistics and residual space statistics respectively. Determine the control limits at a given significance level through kernel density estimation. Samples with statistics exceeding the control limits are considered fault samples, thus obtaining the fault detection results. Step S4 is detailed below: S41: Offline modeling stage: Train the causal graph spatiotemporal autoencoder CGSTAE based on normal training samples, and construct the feature space statistic T² and the residual space statistic SPE. The feature space statistic at time t... The calculation formula is: , in Σ and Σ are the mean and covariance matrices of the hidden layer features of normal samples, respectively; It is the final hidden layer feature output by the encoder in the spatiotemporal encoder decoder module; Residual spatial statistics at time t The calculation formula is: , Control limits of feature space statistics Control limits for residual space statistics Determined at a given significance level by kernel density estimation; S42: Online monitoring phase: at every moment The input matrix is obtained by reorganizing the process data through a sliding window. Spatial statistics are calculated using a trained causal graph spatiotemporal autoencoder (CGSTAE). and residual space statistics ,if and If the process is normal, then the fault is determined; otherwise, the fault is determined.
[0030] Based on the industrial process fault detection method based on spatiotemporal causal graph autoencoders shown in S1-S4 above, a three-step causal graph structure learning algorithm is proposed for CGSTAE training, improving the reliability and interpretability of industrial process monitoring. It should be noted that the key feature of the industrial process fault detection method based on spatiotemporal causal graph autoencoders lies in the causal graph-based CGSTAE process monitoring model, which exhibits significant advantages in both reliability and interpretability.
[0031] The method described below will be applied to a specific example to demonstrate its implementation and technical effects.
[0032] In this embodiment, the Tennessee-Eastman process (TEP) is used as an example to illustrate the effectiveness of the present invention. The TEP involves 52 variables, including 22 continuous process measurements, 19 component measurements, and 11 manipulated variables. The TEP allows for the simulation of 21 types of faults, facilitating the evaluation of fault detection performance. The proposed method is evaluated using publicly available datasets from the TEP. The TEP process flow is as follows: Figure 4 As shown.
[0033] For TEP industrial processes, this embodiment provides a fault detection method based on a spatiotemporal causal graph autoencoder, and the implementation steps are as follows: Step 1: Data preprocessing; By collecting data from 52 variables in the TEP (Transmission of Processes), using 960 normal operating condition samples for training and 21 fault conditions as the test set, the spatiotemporal input sequence was reconstructed using a sliding window of length w=5. The spatiotemporal process data of the process variables, after data preprocessing, are in the following form: Let represent the training data consisting of N normal samples. The dynamic spatiotemporal input matrix is obtained by reorganizing the data through a sliding window. .
[0034] Step 2: Train CGSTAE; This embodiment utilizes a three-step causal graph structure learning algorithm to train CGSTAE, in order to discover invariant causal graphs from changing correlations.
[0035] 2.1, Pre-training; The parameters of the correlation graph structure learning module and the spatiotemporal encoder-decoder module are jointly optimized by the Adam optimizer with a learning rate of 0.05 and a batch size of 32 during pre-training.
[0036] 2.2 Causal Extraction; Causal graph learning freezes parameters by introducing a trainable causal graph adjacency matrix A, selecting an Adam optimizer with a learning rate of 0.1, and learning the causal graph by minimizing the loss function. The four balancing hyperparameters are as follows: , , , .
[0037] 2.3, Fine-tuning; Remove the SSAM module, fix the resulting causal graph A, and use the Adam optimizer with a learning rate of 0.05 to fine-tune the parameters of the spatiotemporal encoder-decoder to optimize its reconstruction capability.
[0038] Step 3: Fault detection; In TEP process monitoring applications, CGSTAE achieves accurate fault detection through T² and SPE statistics. Taking fault 11 as an example, the application of this method in TEP fault detection is divided into an offline control limit calculation stage and an online monitoring stage.
[0039] 3.1, Offline control limit calculation stage; Based on normal training data, kernel density estimation was used to determine the control limits at a significance level of 0.01. Figure 5 and Figure 6 The red dashed line represents the control limit.
[0040] 3.2 Online monitoring phase; Taking fault 11 as an example, Figure 5 and Figure 6 The paper presents T2 and SPE statistics for test data using different methods, with the fault introduced from the 161st sample. The CGSTAE statistics proposed in this invention can clearly distinguish between fault conditions and normal conditions.
[0041] Step 4: Model evaluation and validation; In this embodiment, process data from the publicly available TEP dataset is used to train and validate the process detection performance of CGSTAE and its comparative models. Comparison methods include autoencoders (AE), LSTM-AE (based on long short-term memory networks), GAE-I (Graph Autoencoder with Pearson correlation coefficient for graph construction), GAE-II (Graph Autoencoder with transfer entropy for graph construction), DGSTAE (Dynamic Graph Spatiotemporal Autoencoder with Spatial Self-Attention Mechanism for Dynamic Graph Construction), and CGSTAE, an industrial process fault detection method based on spatiotemporal causal graph autoencoders proposed in this invention. Detection rate (FDR), false alarm rate (FAR), and F1 score are selected as performance evaluation criteria for fault detection, where the F1 score is a comprehensive performance indicator that considers both FDR and FAR.
[0042] Table 1 presents the fault detection performance of all methods in TEP. CGSTAE achieves the best overall performance with an F1 score of 0.883, significantly reducing the false alarm rate while maintaining high detection accuracy. Thanks to its causal graph learning and spatiotemporal modeling capabilities, CGSTAE is the most reliable method for fault detection in TEP. DGSTAE also performs excellently, with an F1 score of 0.819 and the lowest FAR, demonstrating the importance of graph learning and spatiotemporal modeling. LSTM-AE and GAE-II achieve fault detection results with F1 scores of 0.813 and 0.774, respectively, reflecting their effectiveness in handling temporal and spatial dependencies. In contrast, GAE-I and AE show poor fault detection performance.
[0043] Table 1. Comparison of Fault Detection Performance of All Methods in TEP
[0045]
[0046] Figure 5 and Figure 6 The T² statistic and SPE statistic for all methods are compared for fault 11. Figure 5 and Figure 6 In the figure, (a) to (f) represent AE, LSTM-AE, GAE-I, GAE-II, DGSTAE, and CGSTAE, respectively. The dashed line represents the control limit, and the fault is introduced from the 161st sample. It can be seen that the T² statistic of the comparison methods such as AE and LSTM-AE frequently remains below the control limit after the fault occurs, making it difficult to distinguish between faulty and normal operating conditions. In contrast, the T² statistic of CGSTAE proposed in this invention can clearly distinguish between faulty and normal operating conditions. The SPE statistic of the comparison methods is below the control limit at many points after the fault occurs, resulting in a low detection rate. The CGSTAE proposed in this invention uses T²... 2 The SPE statistic successfully detected TEP fault 11 while maintaining a low false alarm rate.
[0047] Figure 7 A comparison of graph structure learning results on TEP. Figure 7 (a) shows a priori causal graph built from process knowledge. Figure 7 Figures (b) through (d) respectively show the graph structure learning results for GAE-I, GAE-II, and CGSTAE. It can be seen that neither the Pearson correlation coefficient nor the transitivity is sufficient to identify prior causality from process data. Figure 1 The causal relationships learned by CGSTAE are consistent with the underlying process mechanisms, thanks to the causal graph structure learning algorithm. This greatly improves the interpretability of the model.
[0048] As described above, the method of the present invention has satisfactory reliability and interpretability, and can accomplish practical industrial process monitoring tasks.
[0049] This invention provides a method for industrial process fault detection based on a spatiotemporal causal graph autoencoder. Many methods and approaches exist for implementing this technical solution; the above description is merely a preferred embodiment. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this invention, and these improvements and modifications should also be considered within the scope of protection of this invention. All components not explicitly stated in this embodiment can be implemented using existing technologies.
Claims
1. A method for industrial process fault detection based on spatiotemporal causal graph autoencoder, characterized in that, Includes the following steps: Step 1: For the target industrial process, preprocess the spatiotemporal process data of all process variables collected during the operation of the target industrial process to obtain a standardized time series data for each process variable; generate dynamic time series data through a fixed-length sliding window to form standardized dynamic input time series data. Step 2: Establish a causal graph spatiotemporal autoencoder (CGSTAE). The causal graph spatiotemporal autoencoder (CGSTAE) includes a correlation graph structure learning module based on the spatial self-attention mechanism (SSAM) and a spatiotemporal encoder-decoder module based on the graph convolutional long short-term memory network (GCLSTM). Step 3: Execute the three-step causal graph structure learning algorithm to train the causal graph spatiotemporal autoencoder CGSTAE, including three steps: pre-training, causal extraction, and fine-tuning. Step 4: Based on the hidden layer features of the causal graph spatiotemporal autoencoder CGSTAE and the residual data of the reconstructed output, calculate the feature space statistics and residual space statistics respectively. Determine the control limit at a given significance level by kernel density estimation. Samples with statistics exceeding the control limit are regarded as fault samples, and the fault detection results are obtained.
2. The method of claim 1, wherein, In step 1, The data preprocessing includes abnormal data rejection, missing value filling and maximum minimum standardization, and the spatio-temporal process data of all process variables after data preprocessing is in the form of , X represents training data composed of N normal samples, T represents matrix transposition, , the i-th column of X is denoted as , , and the time series data vector of the i-th process variable is denoted as , and the measurement data of the i-th process variable at time t is denoted as The time series data matrix of time t inputting the model is obtained by reorganizing data through a sliding window wherein w represents the length of the sliding window.
3. The method as described in claim 2, characterized in that, In step 2, the preprocessed time series data is input into the spatial self-attention mechanism SSAM. The attention weights between variables are obtained by calculating the query matrix and key matrix, which form the adjacency matrix of the dynamic correlation graph, thereby realizing adaptive learning of the time-varying correlation between process variables. Using a graph structure as input, a spatiotemporal encoder-decoder module based on a graph convolutional long short-term memory network (GCLSTM) is constructed. The reconstructed values of process data are obtained by mapping the hidden layer features and fully connected layers of the spatiotemporal encoder-decoder module, enabling topology-guided spatiotemporal dynamic modeling of industrial processes. The causal graph is defined as a directed unweighted graph. ,in It is a set of nodes. It is a set of directed edges; nodes It is the i-th variable in a given process, and its node attribute is a standardized time-series data vector. Directed edge Representing variables and variables The causal dependency between them.
4. The method as described in claim 3, characterized in that, In step 2, the spatial self-attention mechanism SSAM will process the temporal data matrix. Convert to attention matrix : , The query matrix Key-value matrix , and These are two trainable weight matrices for the Spatial Self-Attention Mechanism (SSAM). Use the Sigmoid activation function; Attention matrix It is considered as an adjacency matrix of a dynamic correlation graph.
5. The method as described in claim 4, characterized in that, In step 2, a spatiotemporal encoder-decoder module based on a graph convolutional long short-term memory network (GCLSTM) is constructed. The encoder in the spatiotemporal encoder-decoder module uses a temporal data matrix. Attention Matrix As input, along the forward timing Circular execution of a graph convolutional long short-term memory network (GCLSTM) to update the forget gate at time k. Input gate Output gate Candidate status Unit status and hidden layer features ,in: , , , , , , Where tanh represents the hyperbolic tangent function. Represents the Hadama product. , , , These represent the graph convolutional layers of the encoder: the forget gate, the input gate, the output gate, and the candidate states, respectively. , , , These represent the learnable bias parameters of the forget gate, input gate, output gate, and candidate state in the encoder, respectively. The encoder in the spatiotemporal encoder-decoder module extracts spatiotemporal information by updating the unit state and hidden features through a graph convolutional long short-term memory (GCLSTM) network, and then passes this information to the decoder. The decoder proceeds along the reverse temporal sequence. The GCLSTM graph convolutional long short-term memory network is executed in a loop to update the unit states and hidden features: , , , , , , The calculation method for GC in graph convolutional layers is as follows: , Where D is the adjacency matrix. The degree matrix, I is the degree matrix with Unit arrays of the same shape For trainable weight parameters, Z represents the node features of the input graph convolutional layer: in the encoder In the decoder ; Finally, by utilizing the hidden layer features, reconstructed values of the process data at time k are generated through a fully connected layer. : , in and These are the trainable weight parameters and bias parameters of the fully connected layer, respectively.
6. The method as described in claim 5, characterized in that, In step 3, firstly, dynamic correlation maps under different operating conditions are generated using the pre-trained spatial self-attention mechanism SSAM; then, based on the principle of causal invariance, stable structures are extracted from the dynamic correlation maps, and causal graphs are generated by integrating process mechanism knowledge; finally, the learned causal graphs are used to fine-tune the parameters of the spatiotemporal encoder decoder module, establishing a causal-based fault detection model; the trainable model parameters in the causal graph spatiotemporal autoencoder CGSTAE are used... It means that among them These represent the parameters of the correlation graph structure learning module. The parameters of the spatiotemporal encoder decoder module are represented by: and This represents the function mappings for the correlation graph structure learning module and the spatiotemporal encoder-decoder module.
7. The method as described in claim 6, characterized in that, In step 3, the pre-training includes: minimizing the mean square error (MSE) of the reconstructed values by jointly optimizing the parameters of the correlation graph structure learning module and the spatiotemporal encoder-decoder module. , in express and The mean square error; Reconstructed value The correlation graph structure learning module and the spatiotemporal encoder decoder module are used to calculate the dynamic correlation between industrial process variables, enabling the spatial self-attention mechanism (SSAM) to adaptively learn the dynamic correlation between industrial process variables. The causal extraction includes: freezing the parameters of the pre-trained spatiotemporal encoder-decoder module and the correlation graph structure learning module, introducing a trainable causal graph adjacency matrix A, and learning the causal graph adjacency matrix A by optimizing the following objectives: , in , , , Indicates the weighting coefficient; The term ensures that the reconstruction capability of the causal graph is consistent with that of the correlation graph by constraining the root mean square error (MSE). The calculation method is as follows: , The term extracts causal relationships based on invariance, and the calculation method is as follows: , in, This represents the element in the i-th row and j-th column of the adjacency matrix A of a trainable causal graph. Adjacency matrix representing a dynamic correlation graph The element in the i-th row and j-th column; The term guarantees that the learned cause-effect graph conforms to the prior knowledge of the process, and the calculation method is as follows: , Where M is the mask matrix representing whether knowledge exists in the representation process, if ,but ;like , "Unknown" indicates that the process knowledge is unknown. This represents the element in the i-th row and j-th column of the mask matrix M; Item and The term is used to improve the sparsity and discreteness of the causal graph, and is calculated as follows: , , The fine-tuning includes: fine-tuning the parameters of the spatiotemporal encoder-decoder module based on the mean square error loss and using the learned causal graph, with the loss function... for: 。 8. The method as described in claim 7, characterized in that, Step 4 includes: performing fault detection, which consists of two phases: offline modeling and online monitoring. Offline modeling stage: The causal graph spatiotemporal autoencoder CGSTAE is trained based on normal training samples, and the feature space statistics T² and residual space statistics SPE are constructed. The feature space statistics at time t are... The calculation formula is: , in Σ and Σ are the mean and covariance matrices of the hidden layer features of normal samples, respectively; It is the final hidden layer feature output by the encoder in the spatiotemporal encoder decoder module; Residual spatial statistics at time t The calculation formula is: , Control limits of feature space statistics Control limits for residual space statistics Determined at a given significance level by kernel density estimation; Online monitoring phase: at every moment The input matrix is obtained by reorganizing the process data through a sliding window. Spatial statistics are calculated using a trained causal graph spatiotemporal autoencoder (CGSTAE). and residual space statistics ,if and If the process is normal, then the fault is determined; otherwise, the fault is determined.
9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, It stores a computer program or instructions that, when run on a computer, perform the steps of the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Complex industrial process fault detection method based on space-time variation graph attention auto-encoder
CN116520799A
Industrial process fault monitoring method based on time information enhancement graph convolution auto-encoder
CN119758925A
Steel production system fault diagnosis method based on ConvLSTM-AE
CN120011985A
Performance causal discovery and fault diagnosis method and system in strip steel hot rolling process
CN120540259A
Virtual power plant fault early warning method based on hierarchical interactive causality graph Transform
CN120611270A
Cited By
Energy-saving and consumption-reducing method for air blower of sewage treatment plant based on machine learning
CN121541475A
Two-channel coal chemical engineering fault diagnosis method based on graph space-time fusion modeling
CN122045782A
A double-channel coal chemical fault diagnosis method based on graph space-time fusion modeling
CN122045782B