Prediction reconstruction framework causal perception space-time network for explaining anomaly monitoring in complex industrial process

Through the causal-aware spatiotemporal network, combined with the time-frequency fusion network of temporal convolution and Fourier analysis, and using residual graph attention and variational autoencoder, the problem of high false alarm rate of traditional monitoring methods in high-dimensional nonlinear industrial processes is solved, and explainable anomaly monitoring and cause analysis are achieved.

CN120744571APending Publication Date: 2025-10-03CENT SOUTH UNIV
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510778770.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional industrial process monitoring methods have high false alarm rates under high dimensions, nonlinearity and noise interference, have difficulty in accurately identifying the complex spatiotemporal relationships between variables, and lack interpretability.

Method used

A sequential causal graph reasoning framework is adopted to combine the time-frequency fusion network of temporal convolutional network and Fourier analysis network, spatiotemporal feature learning is performed through residual graph attention network, and the prediction error is reconstructed using variational autoencoder to construct a causal-aware spatiotemporal network for anomaly monitoring.

Benefits of technology

It effectively reduces the false alarm rate and provides explainable abnormal monitoring results. It can accurately identify the cause of abnormalities and reduce the false alarm rate, thereby improving monitoring efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744571A_ABST
    Figure CN120744571A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of fault detection, and particularly discloses a prediction reconstruction framework causal perception space-time network for explaining anomaly monitoring in a complex industrial process, comprising the following steps: S01, constructing graph data E (V) and a causal graph; automatically adjusting the fusion proportion of the time-frequency characteristics according to the data characteristics so as to ensure that the model can comprehensively capture the information of the data in the time domain and the frequency domain; secondly, introducing a residual image attention network (RGAT), and converting the image data E (V) into image structure data G (S (V), E (V)); and S03, reconstructing a prediction error by adopting a variational automatic encoder (VAE), learning an error mode of normal data, providing an anomaly judgment AD (V) for anomaly detection, analyzing a causal relationship between data in combination with a causal graph, and positioning an anomaly reason according to an anomaly score, so as to form a prediction result. The network solves the problem that a traditional monitoring network is high in false alarm rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of fault detection, and specifically discloses a causal-aware spatiotemporal network, a predictive reconstruction framework for interpretable anomaly monitoring of complex industrial processes. Background Art

[0002] With the continuous expansion of production scale and the deep integration of production processes, production processes in industrial scenarios are becoming increasingly complex, showing high-dimensional characteristics, accompanied by high-intensity noise interference, and exhibiting significant nonlinear characteristics. In this context, extremely complex spatiotemporal coupling relationships are generated between multivariate time series variables collected on-site, making the interactions and inherent connections between variables extremely hidden. Accurately understanding the complex relationships between these variables and achieving efficient and explainable anomaly monitoring has become a difficult problem that needs to be overcome in the current industrial field.

[0003] As modern complex industrial production continues to expand and processes become increasingly complex, the spatiotemporal coupling relationships between variables have become intricate. The data generated by these processes is characterized by high dimensionality, strong noise, and nonlinearity, which increases the false alarm rate of process monitoring methods and reduces detection efficiency. Accurately revealing the complex interdependencies between variables and capturing normal process patterns to reduce false alarm rates and improve process monitoring efficiency has become a crucial task. Furthermore, reasonable explanations should be provided for abnormal situations to support decision-making for safe production operations. Therefore, it is particularly important to introduce an efficient and interpretable process monitoring method that comprehensively considers the complex data characteristics and the intricate spatiotemporal relationships between variables.

[0004] In the temporal dimension, when process data exhibits nonlinear and nonstationary characteristics, traditional time-domain analysis is susceptible to noise, while frequency-domain analysis struggles to locate abnormal nodes. Time-frequency fusion technology offers a novel approach to improving process detection accuracy by integrating the time-frequency characteristics of gradual trends, cyclical fluctuations, and sudden anomalies. In the spatial dimension, the variable correlation mechanisms within a process stage are determined by both physical layout and control logic. Changes in key variables at a given stage follow physical / chemical laws and can trigger chain reactions as they propagate downstream along the process flow. Within a specific equipment unit, multiple variables form functional spatial coupling relationships around core operating parameters. These coupling relationships are closely mapped to the anomaly-causing processes. Deviations form anomaly propagation paths through spatial coupling relationships or variable transmission chains, leading to cascading deviations at the system level. Therefore, it is necessary to establish a variable correlation model to identify variable nodes with key influences.

[0005] Graph neural networks (GNNs) have shown unique advantages in anomaly detection tasks by explicitly modeling the complex correlations between variables, making them particularly suitable for processing network structure data [. Peng et al. proposed a dynamic spatiotemporal graph neural network that transforms physical connections into graph structures. By combining convolutional neural networks (CNNs) and graph neural networks, this approach achieves dynamic spatiotemporal modeling of complex systems. Lin et al. used context-enhanced graph autoencoders to process spatiotemporal graph structured data, and constructed data by fusing features with correlation coefficients. Compared with the above methods, the graph attention network (GAT) adaptively captures the spatiotemporal dependencies between nodes through a dynamic attention mechanism, showing higher flexibility and interpretability in spatiotemporal modeling. Deng et al. proposed a graph bias network that combines structural learning and graph attention mechanisms to capture complex relationships between sensors for anomaly detection. Although the above methods have shown excellent performance on multiple datasets, research on model interpretability has received relatively little attention. Although attention weights can explain the dependencies between nodes, deep models with stacked layers make it difficult to intuitively explain the causal logic between variables. Causal spatiotemporal graph modeling has become a promising direction to improve the interpretability of spatiotemporal graph models;

[0006] This study proposes a causal-aware spatiotemporal network with predictive reconstruction capabilities for interpretable anomaly monitoring in complex industrial processes. The network achieves efficient fault detection for industrial processes with complex variable coupling relationships through temporal dynamic feature extraction, spatial interaction information conversion, and integrated spatiotemporal feature learning. First, the sequential causal graph inference (SCGI) framework is used to extract the potential causal relationships between multivariate time series data. These relationships reflect the complex interactions between variables in system operation. Then, a graph topology structure is constructed based on the potential causal relationships, which characterizes the entity dependencies between variables. The constructed causal graph avoids false relationships and effectively represents

[0007] This paper demonstrates the causal transmission links between variables, providing a basis for explaining anomalies. A time-frequency fusion network (TFFN) based on a temporal convolutional network (TCN) and a Fourier analysis network (FAN) is proposed to extract high-quality fused time and frequency domain information. Next, a residual graph attention network (RGAT) is proposed to represent and learn spatiotemporal causal interactions, thereby effectively predicting normal operating conditions. Finally, a variational autoencoder (VAE) is used to reconstruct the prediction error as a criterion for anomaly judgment, and interpretability analysis is performed based on the causal graph. Summary of the Invention

[0008] The purpose of the present invention is to solve the problem of high false alarm rate in traditional monitoring networks.

[0009] In order to achieve the above object, the present invention provides the following basic scheme:

[0010] A predictive reconstruction framework for interpretable anomaly monitoring of complex industrial processes using causal-aware spatiotemporal networks, including the following steps:

[0011] S01: Adopting the sequential causal graph reasoning framework model, organically combining the gated causal mechanism with the prediction function to construct graph data E(V) and causal graph

[0012] S02: First, a time-frequency fusion network (TFFN) method based on a temporal convolutional network (TCN) and a Fourier analysis network (FAN) is used to extract and fuse the high-level temporal features S(·) of industrial process data. The fusion ratio of the time-frequency features is automatically adjusted according to the data characteristics to ensure that the model can fully capture the data information in the time and frequency domains.

[0013] Next: S(·) obtained through time-frequency feature learning is combined with E(V) representing the spatial structure to form spatiotemporal domain data;

[0014] Then: introduce the residual graph attention network (RGAT) to convert the graph data E(V) into graph structure data G(S(V),E(V));

[0015] Finally: formed by stacking time and space Construct a prediction head module that takes the features extracted in the spatiotemporal representation learning phase as input to accurately predict the variables of the industrial process;

[0016] S03: Use variational autoencoder (VAE) to reconstruct prediction errors, learn the error pattern of normal data, and provide anomaly judgment AD(V) for anomaly detection. Combined with causal graph to analyze the causal relationship between data And locate the cause of the anomaly based on the anomaly score.

[0017] Furthermore, the gated causal mechanism is organically combined with the prediction function to form an architecture that integrates

[0018] The causal gating mechanism and LSTM are as follows: For the data X collected by m sensors during t∈[t0,t1], X=[X1(t),X2(t),…,X m (t)], for each target variable X k , all variables X1,X2,…,X m are included in the prediction task, through the learnable gating vector g=[g1,g2,…,g m ], select the prediction X k Variables with reference value serve as potential causal variables.

[0019] Furthermore, the causal threshold g i As a multiplicative weight, the input is the window history observation value X t-w:t-1 =[X t-w ,X t-w+1 ,…,X t-1 ], where X t =[X 1t ,X 2t ,…,X mt ], the output Z generated by the causal gating mechanism t-w:t-1,i (i=1,2,…,m) is given by the following formula:

[0020] Z t-w:t-1,i =g i ·X i,t-w:t-1 #;

[0021] Among them, g i represents the learnable parameter in the causal gating mechanism, which quantifies the variable X i For the target variable X k The causal contribution of is learned by back propagation and gradient descent, and its update rule is:

[0022]

[0023] Furthermore, LSTM processes the m windowed historical observation sequences Z output by the causal gating mechanism. t-w:t-1,i (i=1,2,…,m) are processed to predict the target variable X k The current value of X kt , the input window size is w, the feature dimension is m, the batch size is N, and the structure of the input data is The input with time step t is About LSTM architecture:

[0024] h i =LSTM(x i )i=1,2,…,m#;

[0025] When calculating each gated unit and cell state, all input variables are simultaneously involved in the operation and their information is integrated with each other. Subsequently, all dimensional features are merged:

[0026] h=Concatenate(h1,h2…h m )#;

[0027] Finally, output the prediction results:

[0028] y=Dense(h)#;

[0029] Define the vertex set of the graph as Construct an n×n causal correlation matrix E, where element E ij Representative node X i With node X j A causal relationship between; Definition:

[0030]

[0031] Among them, E ij =1 means variable X i With variable X j There is a directed causal relationship between them;

[0032] About the two-stage training strategy:

[0033] Phase 1: Initial phase (training iterations < αP): In the early stages of training, the model has limited understanding of the data; let P denote the total number of training iterations, α∈(0,1), and when the learned gating vector g conflicts with the prior knowledge, the corresponding gating value g is reset. i , to limit the search space of the model to the area consistent with the prior knowledge, which can reduce invalid exploration and speed up convergence;

[0034] Phase 2: Advanced phase (number of training iterations ≥ αP): As training progresses, data-driven learning becomes dominant, previous knowledge-based adjustments are no longer performed, and the causal gating coefficients are continuously updated only through data.

[0035] Furthermore, in step S02, the spatiotemporal feature fusion network is constructed based on TFFN-RGAT, where TFFN is a hybrid architecture integrating TCN and FAN, and RGAT integrates the causal spatial association learned by the causal structure reasoning framework and the temporal features extracted by the time-frequency feature fusion network.

[0036] Furthermore, the extracted and fused industrial process variable node S i The features are applied to the Kth th layer, its feature fusion output can be expressed as:

[0037]

[0038] Among them, S represents the set of all industrial nodes; E i Represents the neighborhood spatial relationship set of node i;

[0039] Represents the initial features S of the industrial variable node i (t); W k Indicates the Kth th The weight matrix of the hidden layer from the input of the encoder layer feature map to the hidden layer; σ(·) represents the activation function; It is Kth The relative correlation attention coefficient between neighboring node j and node i is obtained by the following expression:

[0040]

[0041] in, It represents the correlation coefficient between node i and its neighboring node j, which can be calculated by the following formula:

[0042]

[0043] in, It is K th The trainable parameters of the encoder layer; LeakyReLU represents the activation function.

[0044] Furthermore, in step S02, a residual structure is introduced, as follows:

[0045]

[0046] in, is the updated node feature, is the original node feature.

[0047] Furthermore, given that the features obtained by time-frequency domain fusion have multi-channel characteristics, a method of interleaving spatiotemporal convolutional layers is adopted, and the definition of feature update is:

[0048]

[0049]

[0050]

[0051] in, and represents the features extracted and fused by the first layer of network, Represents the features extracted by the second layer, and the output state is processed by the multi-layer perceptron (MLP) Make a prediction:

[0052]

[0053] Where W mlp represents the learnable parameters of the MLP.

[0054] Furthermore, regarding the use of variational autoencoder (VAE) to reconstruct the prediction error: for each variable X i (t), calculate the absolute prediction error, the calculation formula is:

[0055]

[0056] set up For E i Normalized error matrix of the normalized error. In VAE, the error matrix output by the encoder is mapped to the mean μ and logarithmic variance logσ in the space 2 ;

[0057]

[0058] To implement backpropagation training, we use the reparameterization technique to sample from the latent distribution and obtain its representation z:

[0059]

[0060] Here, ∈ is a random vector sampled from a standard normal distribution and ⊙ represents element-wise multiplication.

[0061] Furthermore, the decoder of the encoder represents the reconstruction of the VAE:

[0062]

[0063] The training process of VAE is completed by an N(0,1) Gaussian encoder. The loss function L consists of two parts: one is the reconstruction loss, which is used to measure the difference between the input error matrix and the reconstruction error matrix; the other is the KL divergence, which is used to regularize the potential space distribution. The loss function is defined as follows:

[0064]

[0065] Among them L MSE represents the mean square error (MSE) loss, KL(·) represents the KL divergence, N(μ,σ 2 ) means the mean is μ and the variance is σ 2 Normal distribution, N(0,I) represents the standard normal distribution; using the reconstructed standardized error, calculate the L1-norm distance between the denoised error of the current timestamp and the original standardized error, and use it as the final anomaly score, assuming is the raw standardized error, is the normalized error after reconstruction, then the anomaly score S can be expressed as:

[0066]

[0067] L1 emphasizes the sum of the absolute values ​​of each element in the error vector, making the abnormal information more prominent and facilitating accurate identification of abnormal situations. Since the final abnormal score is a linear combination of the reconstruction error, it can be

[0068] By sorting the single variables that contribute the most to the anomaly score, we can further determine the contribution percentage of each single variable to the final anomaly score. Assume that the anomaly score S is the reconstruction error e1, e2, ..., e n The linear combination of , that is:

[0069]

[0070] where w i is the weight of each single variable, the percentage C of single variable i's contribution to the anomaly score i The calculation formula is:

[0071]

[0072] The principle and effect of this solution are:

[0073] 1. Compared with existing technologies, this network conducts in-depth causal relationship analysis on dynamic data. By designing a sequential causal graph reasoning framework, it explores the potential spatial causal relationship between multivariate time series variables, constructs causal graphs and graph data, and builds a time-frequency fusion network that integrates time series convolution and Fourier analysis to extract and fuse time-frequency features. The residual graph attention network is then used to characterize the spatiotemporal causal interaction and predict the normal operation status of the system. Finally, the variational autoencoder is used to reconstruct the prediction error. Combined with the causal graph, abnormal state judgment and cause analysis are completed, thereby achieving highly interpretable abnormal monitoring results. Through effective spatiotemporal representation learning and system state prediction, the overall monitoring efficiency is improved.

[0074] 2. Compared with the existing technology, the method proposed in this scheme can effectively reduce the false alarm rate and provide interpretable opinions for the diagnosis results. Moreover, the network not only retains the absolute information of the original anomaly score, but also reflects the influence intensity between variables through the weight of the edge. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0076] Figure 1 A flowchart of constructing a causal-aware spatiotemporal network for predictive reconstruction of explainable anomaly monitoring in complex industrial processes, as proposed in an embodiment of the present application, is shown;

[0077] Figure 2A schematic diagram of the purification process of a causal-aware spatiotemporal network for a predictive reconstruction framework for interpretable anomaly monitoring of complex industrial processes, as proposed in an embodiment of the present application, is shown;

[0078] Figure 3 A monitoring result diagram of a purification process of a causal-aware spatiotemporal network for interpretable anomaly monitoring of complex industrial processes, proposed in an embodiment of the present application, is shown;

[0079] Figure 4 A diagram showing the monitoring limitation results of a purification process of a causal-aware spatiotemporal network for interpretable anomaly monitoring of complex industrial processes, proposed in an embodiment of the present application;

[0080] Figure 5 An attention coefficient matrix diagram of a causal-aware spatiotemporal network for a predictive reconstruction framework for interpretable anomaly monitoring of complex industrial processes, as proposed in an embodiment of the present application, is shown;

[0081] Figure 6 A causal relationship diagram of a causal-aware spatiotemporal network of a predictive reconstruction framework for interpretable anomaly monitoring of complex industrial processes proposed in an embodiment of the present application is shown;

[0082] Figure 7 The following figure shows a variable anomaly contribution diagram of a causal-aware spatiotemporal network of a predictive reconstruction framework for interpretable anomaly monitoring in complex industrial processes proposed in an embodiment of the present application;

[0083] Figure 8 A visualization diagram of anomaly contributions of a causal-aware spatiotemporal network of a predictive reconstruction framework for interpretable anomaly monitoring of complex industrial processes proposed in an embodiment of the present application is shown. DETAILED DESCRIPTION

[0084] In order to further illustrate the technical means and effects adopted by the present invention to achieve the predetermined purpose of the invention, the specific implementation methods, structures, features and effects of the present invention are described in detail below in conjunction with the accompanying drawings and preferred embodiments.

[0085] A predictive reconstruction framework for interpretable anomaly monitoring of complex industrial processes Causally-aware spatiotemporal networks, implementing e.g. Figures 1-8 As shown:

[0086] First, we establish a temporal causal discovery framework: we propose a causal discovery method for multivariate time series variables. This method uses the sequential causal graph inference (SCGI) framework model to organically combine the gated causal mechanism with the prediction function to achieve in-depth mining of the spatial causal relationship between variables, so as to construct graph data E(V) and causal graphs. In addition, the two-stage training strategy can significantly accelerate the convergence of the model by incorporating prior knowledge, thereby greatly improving the efficiency and accuracy of causal discovery;

[0087] Then, a spatiotemporal feature fusion framework is established: a time-frequency fusion network (TFFN) method based on a temporal convolutional network (TCN) and a Fourier analysis network (FAN) is proposed to extract and fuse the high-level temporal features S(·) of industrial process data and automatically adjust the time-frequency features according to the data characteristics.

[0088] The fusion ratio is set to ensure that the model can fully capture the information of the data in the time domain and frequency domain. Subsequently, S(·) obtained by learning the time-frequency features is combined with E(V) representing the spatial structure to form spatiotemporal domain data. Then, the residual graph attention network (RGAT) is introduced to convert the graph data E(V) into graph structure data G(S(V), E(V)), thereby systematically integrating the spatial correlation features embedded in the data. This method implements spatiotemporal representation learning of data and effectively mines the potential information of data in the spatiotemporal dimension. Finally, a prediction head module is constructed, which takes the features extracted in the spatiotemporal representation learning phase as input to accurately predict the variables of the industrial process.

[0089] Finally, anomaly detection and root cause analysis: Variational Autoencoder (VAE) is used to reconstruct the prediction error, learn the error pattern of normal data, and provide anomaly judgment AD(V) for anomaly detection. Combined with causal graph to analyze the causal relationship between data The cause of the anomaly is located based on the anomaly score, enabling traceable diagnosis of abnormal data and clearly presenting the causal chain of the anomaly.

[0090] For two time series X t and Y t , if X has Granger causality on Y, then the past values ​​of X should contain predictive information that is useful for predicting Y. Mathematically speaking, given the historical observations of Y (Y t-1 ,Y t-2 ,…,Y t-p ), add the past value of X (X t-1 ,X t-2 ,…,X t-q ) can significantly improve Y t If the prediction accuracy is good, then X is said to Granger cause Y:

[0091]

[0092] Among them, α i and β j is the unknown coefficient to be estimated, ∈ t is a white noise process with zero mean.j Conducting a statistical significance test can confirm that there is a Granger causal relationship from X to Y;

[0093] Based on this, we propose a prediction-contribution transformation architecture that integrates a causal gating mechanism and LSTM. We further use this framework to develop the SCGI model for time series data analysis.

[0094] First, for the data X = [X1(t), X2(t), ..., X m (t)]. For each target variable X k , all variables X1,X2,…,X m In order to capture the causal relationship between variables, an adaptive causal

[0095] Gating mechanism, through the learnable gating vector g=[g1,g2,…,g m ], select the prediction X k Variables with reference value serve as potential causal variables.

[0096] Next: Causal threshold g i As a multiplicative weight, the dimension and structure of the input are preserved. The input consists of the window history observation value X t-w:t-1 =[X t-w ,X t-w+1 ,…,X t-1 ], where X t =[X 1t ,X 2t ,…,X mt The output Z generated by the causal gating mechanism t-w:t-1,i (i=1,2,…,m) is given by the following formula:

[0097] Z t-w:t-1,i =g i ·X i,t-w:t-1 #;

[0098] Among them, g i represents the learnable parameter in the causal gating mechanism, which quantifies the variable X i For the target variable X k The causal contribution of . This parameter is learned through back propagation and gradient descent, and its update rule is:

[0099]

[0100] LSTM outputs the m windowed historical observation sequence Z of the causal gating mechanism t-w:t-1,i (i=1,2,…,m) are processed to predict the target variable Xk The current value of X kt As a variant of recurrent neural network (RNN), LSTM captures long-term dependencies in sequence data by regulating the information flow of input, forget and output gates. The input window size is w, the feature dimension is m, the batch size is N, and the structure of the input data is The input with time step t is Regarding the detailed description of the LSTM architecture, for clarity, the multi-input processing pipeline can be simplified as follows:

[0101] h i =LSTM(x i )i=1,2,…,m#;

[0102] It is important to note that when calculating each gate unit and cell state, all input variables will participate in the operation at the same time and their information will be integrated with each other. Subsequently, all dimensional features will be merged:

[0103] h=Concatenate(h1,h2…h m )#;

[0104] Finally: output the prediction results:

[0105] y=Dense(h)#;

[0106] Define the vertex set of the graph as Construct an n×n causal correlation matrix E, where element E ij Representative node X i With node X j The causal relationship between. Definition:

[0107]

[0108] Among them, E ij =1 means variable X i With variable X j There is a directed causal relationship between them.

[0109] For processes with existing domain knowledge, a two-stage training strategy that integrates previous information is proposed:

[0110] Phase 1: Initial phase (training iterations < αP): In the early stages of training, the model has limited understanding of the data. Let P denote the total number of training iterations, and α∈(0,1). If the learned gating vector g conflicts with the prior knowledge, the corresponding gating value g is reset. i , to limit the model's search space to an area consistent with prior knowledge. This can reduce invalid exploration and speed up convergence.

[0111] Phase 2: Advanced (number of training iterations ≥ αP): As training progresses, data-driven learning becomes dominant. Previous knowledge-based adjustments are no longer performed, and the causal gating coefficients are continuously updated solely based on data. This increases the model's information entropy and improves its adaptability to complex data distributions.

[0112] About establishing a spatiotemporal feature fusion framework:

[0113] 1. Time-frequency feature fusion based on TFFN. The process variables in complex industrial systems show complex temporal and spectral dynamic changes. To this end, TFFN, a hybrid architecture integrating TCN and FAN, is proposed. TFFN jointly processes time domain and frequency domain data, extracts complementary features, and captures variables.

[0114] The complex time-frequency dependencies between them.

[0115] FAN is a novel structure based on Fourier series decomposition, which aims to enhance the modeling of periodic phenomena

[40] . By incorporating periodicity into the network structure and computation, FAN can accurately represent and predict periodic patterns. For X = [X1(t), X2(t), …, X m (t)] for each variable X i (t), collected by m sensors during t∈[t0,t1], performs the same processing:

[0116] F i (t) = FAN(X i (t))#;

[0117] T i (t) = TCN (X i (t))#;

[0118] S i (t) = S(T i (t),F i (t))#;

[0119] Among them, F i (t),T i (t),S i (t) represents X i (t) feature extraction and fusion process, for the convenience of description, X i (t) is denoted as x in the following. FAN layer The mathematical definition of is:

[0120]

[0121] in:

[0122] and is a learnable parameter. The hyperparameter d p and Represents W p and The operator [·||·] represents a concatenation along the first dimension. The complete FAN consists of Multiple layers stacked together:

[0123]

[0124] in:

[0125]

[0126] 2. The spatiotemporal convolutional network consists of multiple residual spatiotemporal convolutional layers, which extract and aggregate high-level spatiotemporal features in a non-recursive manner, thus avoiding some of the shortcomings of recurrent neural networks (RNNs), such as time-consuming iterations and gradient explosion. The temporal convolution layer is defined as:

[0127]

[0128] in, In order to solve the problem of gradually decreasing sequence length caused by low-level temporal information aggregation, the final Q l+1 Elements will be drawn from Z l The sequence length is truncated in the middle, and the remaining input is Z l will be truncated to TCN(Z l ). The temporal convolution function uses a gating mechanism to aggregate information flow:

[0129] TCN(Z l ,Φ l )=f C (Z l )⊙f G (Z l )#;

[0130] Among them, Z^0=x, Φ^l are learnable parameters, ⊙ represents matrix multiplication, f C (·) and f G (·) are filtered convolution and gated convolution respectively, which are defined as:

[0131]

[0132] where h(·,·) represents the convolution operation.

[0133] The extracted temporal and frequency features are aggregated through a dynamic gating mechanism. A gating unit is introduced:

[0134] g=σ(W g [FAN(x),TCN(x)]+bg )#;

[0135] Where σ is the activation function, W g is the gating weight matrix, b g is the bias term. The gate vector g has a value between 0 and 1, which can dynamically control the weight of time and frequency characteristics. Finally, the fusion feature of the variable is obtained:

[0136]

[0137] Where ⊙ represents element-wise multiplication. After the union operation, multi-channel data is obtained.

[0138] The dynamic gating mechanism adaptively adjusts the weights according to the pattern of data changes over time, allowing the model to effectively capture the most informative features in the data.

[0139] Causal Correlation Embedding-Spatiotemporal Feature Fusion Based on RGAT:

[0140] The details are as follows:

[0141] Each unit in the system (such as a device or monitoring point) is abstracted as a node in the graph, and the relationships between them are defined as edges. Edge relationships are inferred by a causal structural inference framework that captures the underlying causal dependencies between variables. At the same time, the time-frequency domain features fused by TFN-FAN are used as node features. From a temporal perspective, these features comprehensively and meticulously reflect the characteristics of the data in both time and frequency dimensions, providing rich information for subsequent analysis. Spatially, GAT dynamically weighs the importance of connections between nodes through an attention mechanism, considers the structural relationships between adjacent nodes, and fully captures the spatial interactions between variables while fusing features. This approach enables the model to effectively represent the normal operating mode of the system.

[0142] The industrial process variable node S extracted and fused by TFN-FAN i The features are applied to the Kth th layer, its feature fusion output can be expressed as:

[0143]

[0144] Among them, S represents the set of all industrial nodes; E i Represents the neighborhood spatial relationship set of node i; Represents the initial features S of the industrial variable node i (t); W k Indicates the Kth th The weight matrix of the hidden layer from the input of the encoder layer feature map to the hidden layer; σ(·) represents the activation function; It is Kth The relative correlation attention coefficient between neighboring node j and node i is obtained by the following expression:

[0145]

[0146] in, It represents the correlation coefficient between node i and its neighboring node j, which can be calculated by the following formula:

[0147]

[0148] in, It is K th The trainable parameters of the encoder layer. LeakyReLU represents the activation function.

[0149] During the attention mechanism and aggregation operation, the complexity of the data and the multi-step nature of the calculation may

[0150] This will lead to the loss of important information. To this end, a residual structure is introduced. In the traditional GAT, the feature update of a node is achieved by aggregating the features of its neighboring nodes and applying an attention mechanism. The updated features of the node and the original features are combined through residual connections to retain key information. The definition of feature update is:

[0151]

[0152] in, is the updated node feature, is the original node feature. β is a learnable parameter that balances the contribution of the original and updated features. This residual mechanism ensures that key information in the original features is preserved during the update process, thereby improving the stability and accuracy of the model.

[0153] Given the multi-channel nature of the features obtained by time-frequency fusion, a multi-head attention mechanism is used to process each channel independently. This is because the features of different channels may contain information of different granularities, and separate processing can better explore the unique value of each channel. In order to accurately simulate the spatiotemporal dynamics of a series of historical observations in a multivariate time series, a method of interleaving spatiotemporal convolutional layers is used instead of simply concatenating or averaging the results of multi-head attention. The definition of feature update is:

[0154]

[0155] in, and represents the features extracted and fused by the first layer of network, Represents the features extracted by the second layer. Finally, the output state is processed by multi-layer perceptron (MLP) Make a prediction:

[0156]

[0157] Where W mlp represents the learnable parameters of the MLP. This step is the key to detecting anomalies in multivariate time series, providing an accurate data foundation and decision-making reference for subsequent anomaly detection.

[0158] Finally, VAE is used for anomaly monitoring and source tracing analysis:

[0159] The details are as follows:

[0160] After making a prediction, this embodiment considers that the reconstruction of outliers may cause deviations in the learning process within the neural network module, thereby affecting the prediction results to varying degrees. In other words, if an abnormal sample cannot be obtained, the node may randomly generate or produce erroneous data. Using VAE can learn the distribution of normal prediction errors, effectively separating abnormal errors from normal prediction errors, thereby improving the overall ability to detect anomalies. For each variable X i (t), calculate the absolute prediction error, the calculation formula is

[0161]

[0162] set up For E i Normalized error matrix. In VAE, the error matrix output by the encoder is mapped to the mean μ and logarithmic variance logσ in the space 2 .

[0163]

[0164] To implement backpropagation training, this example uses a reparameterization technique to sample from the potential distribution and obtain its representation z:

[0165]

[0166] Here, ∈ is a random vector sampled from a standard normal distribution and ⊙ represents element-wise multiplication.

[0167] The decoder of the encoder represents the reconstruction of the VAE:

[0168]

[0169] The training process of VAE is completed by an N(0,1) Gaussian encoder. The loss function L consists of two parts: one is the reconstruction loss, which is used to measure the difference between the input error matrix and the reconstruction error matrix; the other is the KL divergence, which is used to regularize the potential space distribution. The loss function is defined as follows:

[0170]

[0171] Among them L MSE represents the mean square error (MSE) loss, KL(·) represents the KL divergence, N(μ,σ 2 ) means the mean is μ and the variance is σ 2 Normal distribution, N(0,I) represents the standard normal distribution.

[0172] Using the reconstructed normalized error, this embodiment calculates the L1-norm distance between the denoised error of the current timestamp and the original normalized error, and uses it as the final anomaly score. Assume is the raw standardized error, is the normalized error after reconstruction. Then the anomaly score S can be expressed as:

[0173]

[0174] L1 emphasizes the sum of the absolute values ​​of each element in the error vector, making the abnormal information more prominent and facilitating accurate identification of abnormal situations.

[0175] Since the final anomaly score is a linear combination of the reconstruction errors, this embodiment can find the important causes of the abnormal event by sorting the single variables that contribute the most to the anomaly score. In practical applications, this embodiment needs to further determine the contribution percentage of each single variable to the final anomaly score. Assume that the anomaly score S is the reconstruction error e1, e2, ..., e n The linear combination of , that is:

[0176]

[0177] where w i is the weight of each single variable. Then, the contribution percentage C of single variable i to the abnormal score is i The calculation formula is:

[0178]

[0179] Then, combined with important anomaly contributing variables and cause-and-effect diagrams, operators can gain a detailed understanding of the root causes of abnormal events, effectively identifying the root causes of abnormal events and taking targeted mitigation measures.

[0180] Testing and experimental verification data:

[0181] The purification system consists of a complex multi-reaction unit designed to effectively remove copper and cobalt ions, with complex coupling relationships between its components. The dynamic characteristics of the system are mainly due to three key factors:

[0182] 1) Inherent fluctuations in the physical properties of raw materials;

[0183] 2) Environmental disturbances including temperature and pressure changes;

[0184] 3) Complex nonlinear interactions between multiple process variables. Figure 2 As shown in the figure, the system achieves the selective removal of impurity metal ions through a precisely controlled displacement-precipitation reaction chain. It is worth noting that changes in the raw material composition will cause the time evolution of key process parameters (including reaction kinetic coefficients, mass transfer rates and crystallization kinetics), resulting in cascade dynamic interference between the operating units. This multi-physics coupling effect significantly amplifies the prediction error of the traditional static model. The industrial-scale process purification system contains a complex multi-reaction unit designed to effectively remove copper and cobalt ions, and there is a complex coupling relationship between its various components. The dynamic characteristics of the system are mainly due to three key factors:

[0185] 1) Inherent fluctuations in the physical properties of raw materials;

[0186] 2) Environmental disturbances including temperature and pressure changes;

[0187] 3) Complex nonlinear interactions between multiple process variables. Figure 2 As shown in Figure 1, the system achieves the selective removal of impurity metal ions through a precisely controlled replacement-precipitation reaction chain. It is worth noting that changes in the raw material composition will cause the time evolution of key process parameters (including reaction kinetic coefficients, mass transfer rates, and crystallization kinetics), resulting in cascade dynamic interference between the various operating units. This multi-physics coupling effect significantly amplifies the prediction error of traditional static models, posing challenges to industrial-scale process control and optimization, as shown in Table 1. The physical meaning of purification process variables

[0188]

[0189] Table 1

[0190] As shown in Table 1, this example considers 16 key monitoring variables in the system configuration. Specifically, the copper removal unit has six online monitoring parameters and three offline analysis indicators, while the cobalt removal unit has six real-time monitoring points and one artificial chemical analysis parameter. The experimental dataset consists of 16,000 training samples and 2,000 test samples, of which the test set includes 300 anomalous samples to evaluate the model's robustness under non-ideal conditions.

[0191] about Figure 2The naming rules for other marks not shown are as follows: the specific device or equipment is given in the box in the figure, and the first digit of the subscript "1" and "2" correspond to the copper removal and cobalt removal processes respectively; the second digit of the subscript is used to distinguish multiple similar devices or condition variables under the same process, for example, "R24" represents the fourth material bin in the cobalt removal reaction process.

[0192] Two comprehensive experiments are conducted to rigorously validate the proposed method. First, a comparative experiment is conducted to evaluate the performance of the causal-aware spatiotemporal network anomaly monitoring framework with prediction-reconstruction capabilities and the baseline monitoring method. Then, the detection results of the proposed method are analyzed for traceability and interpretability.

[0193] 1. Comparative Experiments: Four unsupervised monitoring methods were compared: PCA, AE, STGCN and GDN. The false alarm rate (FAR) and fault detection rate (FDR) of these methods are shown in Figure 3 and Figure 4 and Table 2;

[0194] Abnormal monitoring performance of the purification process

[0195]

[0196] Table 2

[0197] The conclusion is as follows:

[0198] Monitoring results for the purification process show that PCA achieves a relatively high FAR value of 87.41%, significantly exceeding the normal characteristics of the purification process. PCA exhibits significant performance limitations in anomaly monitoring for complex industrial processes. This limitation may be primarily attributed to PCA's lack of adaptability to dynamic processes. AE also exhibits limitations, with a FAR of 47.05%, likely due to redundant information in multivariate time series variables. Compared to PCA and AE, STGCN achieves a significantly lower FAR of only 15.38%, demonstrating its superior ability to distinguish normal from anomalous samples. However, STGCN also has significant limitations. Its reliance on a fixed adjacency matrix fails to capture the dynamic correlations in industrial data that change over time and operating conditions. GDN offers advantages such as a low failure rate of 6.37% and high sensitivity to significant anomalies. Its graph neural network architecture exploits complex correlations between multidimensional data to accurately identify anomalous patterns that differ from known operating conditions. However, this model lacks adaptability to dynamic changes in normal operating conditions and tends to misclassify gradual parameter fluctuations as anomalies. The proposed method achieves excellent fault detection performance with no false positives and an FDR of 99%. This demonstrates that the method can effectively analyze and exploit the spatiotemporal correlations between variables while adapting to the dynamic changes in industrial data, thereby improving detection accuracy.

[0199] In order to better study the superiority of the proposed method, in the above comparison methods, this example selects TGCN and GDN, two methods based on the prediction framework, for further analysis of the supplementary experiments. Supplementary experiments: This example explores the impact of the prediction performance of these two methods on the anomaly detection performance and interpretability. Table 3 lists the MAE (mean absolute error) and R 2 Comparison of the coefficient of determination (R) where R is not calculated for variables with a shift in the distribution between the training and test sets 2 , and the average does not include the maximum and minimum values. This embodiment draws the following main conclusions: STGCN captures the long-term dependencies in the time series through its temporal convolutional network module. In addition, it also uses the graph structure to simulate the intricate relationships between nodes. Compared with traditional statistics-based methods, STGCN has significant advantages in simulating the dynamic changes of industrial processes. In terms of accurate prediction of the process, STGCN performs generally (median R 2 The GDN can reveal the structural relationship between modeling variables and make predictions based on the obtained structure. The effective learning of dynamic relationships enables GDN to approximately track process changes (median R 2 0.15, MAE 0.5455), which is consistent with the complex dynamics inherent in industrial processes. However, due to the lack of temporal feature extraction and modeling, GDN shows limitations in predicting complex industrial processes. Although it can still distinguish the prediction deviation between normal data and abnormal data to a large extent, the lack of accurate capture of model operation greatly limits the credibility of its interpretability. The proposed method can effectively learn the potential causal relationship between variables. By integrating time-frequency and spatial domain features, the method significantly enhances the modeling ability of complex dynamic systems (median R 2 =0.97, and MAE is 0.074). Therefore, this method can predict complex processes with high accuracy, thus achieving low false alarm rate and high misdiagnosis rate;

[0200]

[0201] Table 3

[0202] Explainability of anomaly root cause analysis:

[0203] Interpretability is revealed through the following three dimensions: multivariate time series causal graph structure inference technology that reveals the variable causal relationship network, residual graph attention network that extracts variable dynamic attention coefficient matrix, and variable anomaly score decomposition mechanism based on contribution analysis. Figure 5-Figure 8 This is summarized. Figure 5, the attention coefficient matrix not only quantifies the strength of the association between variables, but also intuitively shows the connectivity between variables. Uncolored nodes indicate that there is no significant potential association between the corresponding time-lagged variable pairs, while the color intensity of colored nodes accurately reflects the degree of association between variables through heat maps. The association between variables is inferred from the data through the SCGI model, and the attention coefficient is dynamically learned through RGAT, so as to deeply model the complex dependencies between variables. Figure 6 As shown, the causal diagram is based on Figure 5 The potential causal associations in the ,are obtained after applying the false causality elimination technique, and a ,network structure representing the true causal relationships between industrial ,process variables is constructed. Figure 6 Nodes var1-var13 represent key variables in the industrial process, and directed edges represent effective directed causal relationships. The historical data of the starting variables of the edges can significantly predict the ending variables, which is consistent with the definition of Granger causality. This causal network provides a reliable structural foundation for subsequent anomaly propagation analysis. Specifically: Figure 6 The arrows in the middle represent Granger causal relationships between variables discovered using the SCGI method. The variable at the beginning of the arrow is the influencing independent variable, and the variable at the end of the arrow is the affected dependent variable. The arrow in a Granger causal relationship should indicate "if historical information about variable A significantly improves the predictive power of variable B, then A → B."

[0204] Figure 7 Var1-var13 in the figure shows the variable contribution matrix calculated by the anomaly detection model based on the deep neural network. The contribution value of each variable reflects its deviation from the normal mode predicted by the model when the anomaly occurs. The sum of the contribution values ​​of all variables is normalized to 1, as shown in Figure 8 The anomaly contribution network shown here integrates attention coefficients and anomaly scores. It dynamically aggregates information from one- and two-hop neighbors through an attention mechanism. Combined with the directed edge structure of the causal graph, it forms a composite network encompassing both direct and indirect propagation paths between variables. This network not only preserves the absolute information of the original anomaly scores but also reflects the influence strength between variables through edge weights.

[0205] In the outlier contribution rate, Figure 7 Among the variables with the highest causal relationship, var3, var1, var7 and var15 are at the top. Figure 6 In the variability analysis, var1 is located upstream relative to var3 and var7, and its abnormal contribution ratio is still the highest. Combined with causal path analysis, these findings suggest that var1 may

[0206] It is the key driving factor of the anomaly, and its abnormal fluctuation plays an important predictive role in the deviation of the entire system. It is worth noting that although var15 ranks fourth in the abnormal contribution, its upstream position relative to var7 deserves extra attention. Combined with the actual physical meaning, var1 represents the flow rate of the inlet solution of the copper removal process, while var7 represents the redox potential. This chain effect stems from the abnormal fluctuation of the material flow rate in the copper removal process, which causes the redox potential of the solution to deviate from the normal range. At the same time, the failure to optimize the operating parameters in a timely manner further aggravates the attenuation of the copper removal efficiency. Due to the continuity of the metallurgical production process, the flow anomaly in the copper removal process will be transmitted through the system to the cobalt removal process, resulting in abnormal fluctuations in key control parameters such as the redox potential in the process, which will ultimately have a significant negative impact on the technical performance of the cobalt removal operation.

[0207] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed as a preferred embodiment as above, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments using the technical contents disclosed above without departing from the scope of the technical solution of the present invention. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A predictive reconstruction framework for interpretable anomaly monitoring in complex industrial processes: Causal-aware spatiotemporal networks, characterized by: The following steps are involved: S01: Adopting the sequential causal graph reasoning framework model, organically combining the gated causal mechanism with the prediction function to construct graph data E(V) and causal graph A two-stage training strategy is proposed to accelerate the convergence of the model by combining prior knowledge; S02: First, a time-frequency fusion network (TFFN) method based on a temporal convolutional network (TCN) and a Fourier analysis network (FAN) is established to extract and fuse the high-level temporal features S(·) of industrial process data. The fusion ratio of time-frequency features is automatically adjusted according to the data characteristics to ensure that the model can fully capture the data information in the time and frequency domains. Next: S(·) obtained through time-frequency feature learning is combined with E(V) representing the spatial structure to form spatiotemporal domain data; Then: introduce the residual graph attention network (RGAT) to convert the graph data E(V) into graph structure data G(S(V),E(V)); Finally: formed by stacking time and space Construct a prediction head module that takes the features extracted in the spatiotemporal representation learning phase as input to accurately predict the variables of the industrial process; S03: Use variational autoencoder (VAE) to reconstruct prediction errors, learn the error pattern of normal data, and provide anomaly judgment AD(V) for anomaly detection, and combine causal graphs to analyze the causal relationship between data And locate the cause of the anomaly based on the anomaly score.

2. A causal-aware spatiotemporal network for predictive reconstruction of explainable anomaly monitoring in complex industrial processes according to claim 1, characterized in that: The gated causal mechanism is organically combined with the prediction function to form an architecture that integrates the causal gating mechanism and LSTM as follows: For the data X collected by m sensors during t∈[t0,t1], X=[X1(t),X2(t),…,X m (t)], for each target variable X k , all variables X1,X2,…,X m are included in the prediction task, through the learnable gating vector g=[g1,g2,…,g m ], select the prediction X k Variables with reference value serve as potential causal variables.

3. A causal-aware spatiotemporal network for predictive reconstruction of explainable anomaly monitoring in complex industrial processes according to claim 2, characterized in that: Causal threshold g i As a multiplicative weight, the input is the window history observation value X t-w:t-1 =[X t-w ,X t-w+1 ,…,X t-1 ], where X t =[X 1t ,X 2t ,…,X mt ], the output Z generated by the causal gating mechanism t-w:t-1,i (i=1,2,…,m) is given by the following formula: Z t-w:t-1,i =g i ·X i,t-w:t-1 #; Among them, g i represents the learnable parameter in the causal gating mechanism, which quantifies the variable X i For the target variable X k The causal contribution of is learned by back propagation and gradient descent, and its update rule is:

4. A causal-aware spatiotemporal network for predictive reconstruction of explainable anomaly monitoring in complex industrial processes according to claim 3, characterized in that: LSTM outputs the m windowed historical observation sequence Z of the causal gating mechanism t-w:t-1,i (i=1,2,…,m) are processed to predict the target variable X k The current value of X kt , the input window size is w, the feature dimension is m, the batch size is N, and the structure of the input data is The input with time step t is About LSTM architecture: h i =LSTM(x i )i=1,2,…,m#; When calculating each gated unit and cell state, all input variables are simultaneously involved in the operation and their information is integrated with each other. Subsequently, all dimensional features are merged: h=Concatenate(h1,h2…h m )#; Finally, output the prediction results: y=Dense(h)#; Define the vertex set of the graph as Construct an n×n causal correlation matrix E, where element E ij Representative node X i With node X j the causal relationship between definition: Among them, E ij =1 means variable X i With variable X j There is a directed causal relationship between them; About the two-stage training strategy: Phase 1: Initial phase (training iterations < αP): In the early stages of training, the model has limited understanding of the data; let P denote the total number of training iterations, α∈(0,1), and when the learned gating vector g conflicts with the prior knowledge, the corresponding gating value g is reset. i , to limit the search space of the model to the area consistent with the prior knowledge, which can reduce invalid exploration and speed up convergence; Phase 2: Advanced phase (number of training iterations ≥ αP): As training progresses, data-driven learning becomes dominant, previous knowledge-based adjustments are no longer performed, and the causal gating coefficients are continuously updated only through data.

5. A causal-aware spatiotemporal network for predictive reconstruction of explainable anomaly monitoring in complex industrial processes according to claim 4, characterized in that: In step S02, the spatiotemporal feature fusion network is constructed based on TFFN-RGAT. TFFN is a hybrid architecture that integrates TCN and FAN. RGAT integrates the causal spatial association learned by the causal structure reasoning framework and the temporal features extracted by the time-frequency feature fusion network.

6. A causal-aware spatiotemporal network for predictive reconstruction of explainable anomaly monitoring in complex industrial processes according to claim 5, characterized in that: The extracted and fused industrial process variable node S i The features are applied to the Kth th layer, its feature fusion output can be expressed as: Among them, S represents the set of all industrial nodes; E i Represents the neighborhood spatial relationship set of node i; Represents the initial features S of the industrial variable node i (t); W k Indicates the Kth th The weight matrix of the hidden layer from the input of the encoder layer feature map to the hidden layer; σ(·) represents the activation function; It is K th The relative correlation attention coefficient between neighboring node j and node i is obtained by the following expression: in, It represents the correlation coefficient between node i and its neighboring node j, which can be calculated by the following formula: in, It is K th The trainable parameters of the encoder layer; LeakyReLU represents the activation function.

7. The predictive reconstruction framework causal-aware spatiotemporal network for interpretable anomaly monitoring of complex industrial processes according to claim 1 is characterized by: In step S02, the residual structure is introduced as follows: in, is the updated node feature, is the original node feature.

8. The predictive reconstruction framework causal-aware spatiotemporal network for interpretable anomaly monitoring of complex industrial processes according to claim 7 is characterized in that: Given that the features obtained by time-frequency domain fusion have multi-channel characteristics, the method of interleaving spatiotemporal convolutional layers is adopted. The definition of feature update is: in, and represents the features extracted and fused by the first layer of network, Represents the features extracted by the second layer, and the output state is processed by the multi-layer perceptron (MLP) Make a prediction: Where W mlp represents the learnable parameters of the MLP.

9. A causal-aware spatiotemporal network for predictive reconstruction of explainable anomaly monitoring in complex industrial processes according to claim 8, characterized in that: About using variational autoencoder (VAE) to reconstruct the prediction error: For each variable X i (t), calculate the absolute prediction error, the calculation formula is: set up For E i Normalized error matrix of the normalized error. In VAE, the error matrix output by the encoder is mapped to the mean μ and logarithmic variance logσ in the space 2 ; To implement backpropagation training, we use the reparameterization technique to sample from the latent distribution and obtain its representation z: Here, ∈ is a random vector sampled from a standard normal distribution and ⊙ represents element-wise multiplication.

10. A causal-aware spatiotemporal network for predictive reconstruction of explainable anomaly monitoring in complex industrial processes according to claim 9, characterized in that: The decoder of the encoder represents the reconstruction of the VAE: The training process of VAE is completed by an N(0,1) Gaussian encoder. The loss function L consists of two parts: one is the reconstruction loss, which is used to measure the difference between the input error matrix and the reconstruction error matrix; the other is the KL divergence, which is used to regularize the potential space distribution. The loss function is defined as follows: Among them L MSE represents the mean square error (MSE) loss, KL(·) represents the KL divergence, N(μ,σ 2 ) means the mean is μ and the variance is σ 2 Normal distribution, N(0,I) represents the standard normal distribution; using the reconstructed standardized error, calculate the L1-norm distance between the denoised error of the current timestamp and the original standardized error, and use it as the final anomaly score, assuming is the raw standardized error, is the normalized error after reconstruction, then the anomaly score S can be expressed as: L1 emphasizes the sum of the absolute values ​​of each element in the error vector, making the abnormal information more prominent and facilitating accurate identification of abnormal situations. Since the final anomaly score is a linear combination of the reconstruction errors, the contribution percentage of each single variable to the final anomaly score can be further determined by sorting the single variables that contribute the most to the anomaly score. Assume that the anomaly score S is the reconstruction error of n single variables e1, e2, ..., e n The linear combination of , that is: where w i is the weight of each single variable, the percentage C of single variable i's contribution to the anomaly score i The calculation formula is:

Citation Information

Cited By

  • Performance index prediction method and device based on industrial process monitoring, equipment and medium

    CN121093793A

  • Performance index prediction methods, devices, equipment, and media based on industrial process monitoring

    CN121093793B

  • Internet of Things data intelligent analysis method based on deep learning

    CN121388502A

  • Internet of Things cluster data analysis method based on big data and artificial intelligence

    CN121434866A

  • An internet of things cluster data analysis method based on big data and artificial intelligence

    CN121434866B