Graph neural network fault diagnosis method for process industry
By constructing a graph neural network model with process topology and time feature embedding in the chemical process, the problem of lack of knowledge fusion and timing information capture in the existing technology is solved, and more efficient fault diagnosis effect and multiple fault type distinction capabilities are achieved.
Patent Information
- Application Number
- CN202510364571.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
AI Technical Summary
The existing fault diagnosis method based on graph neural network lacks effective integrated chemical process knowledge and data-driven methods when processing multi-sensor signals in the chemical process, and it is difficult to capture the ability to distinguish between timing information and multiple fault types, resulting in a reduced credibility of diagnostic results.
A graph neural network fault diagnosis method for process industry is proposed. By constructing a graph neural network model with process topology and time feature embedding, integrating prior knowledge of chemical process and multi-sensor time series data, using graph convolutional neural network and knowledge to enhance graph embedding layer, extracting rich feature representations and performing fault diagnosis.
This method effectively integrates process knowledge and data characteristics, improves the interpretability and robustness of fault diagnosis, can more accurately capture the time dependence and complex dynamic faults of multi-sensor signals, and improves the ability to distinguish multiple fault types.
Smart Images

Figure CN120217104A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of fault diagnosis, and particularly relates to a graph neural network fault diagnosis method for process industry. Background Technique
[0002] Since the last century, fault diagnosis technology has been widely applied to chemical production processes and made important contributions to production safety. To accurately determine the type of fault, a large number of academic researches and industrial applications have been carried out. Traditionally, these methods are divided into model-based methods, knowledge-based methods, and data-driven methods.
[0003] With the development of science and technology, data-driven deep learning fault diagnosis methods have been widely studied. Graph neural networks (GNNs) have become a research hotspot in the field of fault diagnosis because they can effectively process unstructured data and have inherent advantages in processing multi-sensor data. They can reveal potential relationships in the data and utilize the dependencies between variables. Among them, graph convolutional neural networks (GCNs) and graph attention networks (GATs) are common variants of graph neural networks.
[0004] Graph convolutional neural network is the basic graph neural network, while graph attention network is extended based on graph convolutional neural network. The graph convolutional layer in graph convolutional neural network usually consists of 1 - 3 layers. See Figure 1 , as shown is a two-layer graph convolutional neural network. Each node is represented as a feature vector, and the topological relationship of the graph is represented by edges and their weights, reflecting the relationship between nodes. By aggregating information from adjacent nodes, the graph convolutional layer extracts rich feature representations to enhance the understanding of node context. The graph neural network acts on the graph G=(V, E, F), where V is the set of nodes, E is the set of edges representing the connections between nodes, and F is the feature matrix. The key components include the adjacency matrix A, the degree matrix D, and the feature matrix F. To incorporate self-loops, the adjacency matrix A is modified to and normalized using D to obtain The node feature calculation utilizes and a learnable weight matrix W, and applies an activation function (e.g., ReLU) to the non-linear transformation:
[0005]
[0006] where, and are the learnable parameters of the first layer and the second layer. Here, d, h, and f represent the input, hidden, and output feature dimensions respectively.
[0007] Multi-sensor signals not only exhibit correlations between variables but also time-varying correlations. However, graph neural networks have difficulties in capturing time dependencies, including the temporal relationships between variables and the autocorrelations within variables. In addition, existing graph neural network-based fault diagnosis methods usually rely on data mining techniques to extract graph structures and build end-to-end black-box models to capture complex variable relationships. These methods face significant challenges, mainly due to the lack of domain knowledge support, making it difficult to interpret model behaviors and results. The heavy reliance on data quality may also lead the model to learn incorrect variable relationships, especially when the data is noisy or incomplete. Therefore, the existing graph neural network-based fault diagnosis methods mainly have the following problems:
[0008] (1) Lack of effective integration of chemical process knowledge and data-driven methods. The construction of the graph structure of current GNN models usually relies on the correlations between variables or is generated by implicit learning from data, ignoring the topological constraints and mechanism characteristics in chemical processes. This limitation makes it difficult for the model to intuitively interpret the physical meaning of the relationships between variables, resulting in a decrease in the credibility of diagnostic results in engineering practice and hindering its wide application.
[0009] (2) Insufficient utilization of temporal information of node features for dynamic processes. Fault diagnosis in industrial processes relies on accurately capturing the dynamic characteristics of variables changing over time. However, the existing GNN models have limited capabilities in mining temporal features and usually lack specially designed temporal information extraction mechanisms. This leads to the inability of node features to fully reflect the time dependencies and dynamic change laws of variables, weakening the model's ability to identify complex dynamic faults.
[0010] (3) Limited ability to distinguish multiple fault types. The variable relationships in industrial processes exhibit diversity and heterogeneity, and different fault types may correspond to completely different variable association patterns. However, most existing methods use a single graph structure for modeling and lack the ability to flexibly model multiple variable relationships, thus limiting the performance of the model in multi-fault classification tasks. Summary of the Invention
[0011] To overcome the technical defects existing in the prior art, the present invention proposes a graph neural network-based fault diagnosis method for process industries.
[0012] To solve the technical problems existing in the prior art, the technical solution of the present invention is as follows:
[0013] In the first aspect of the present invention, there is provided a graph neural network-based fault diagnosis method for process industries, including at least the following steps:
[0014] Step S1: Collect industrial data applicable to the fault diagnosis task and organize it into a time series data set;
[0015] Step S2: Prepare input data that meets the requirements for the graph neural network model;
[0016] Step S3: Construct a graph neural network model for process topology and time feature embedding, and perform supervised training on the constructed training set;
[0017] Step S4: Load the trained fault diagnosis model and perform tests on the test set.
[0018] In the second aspect of the present invention, there is provided a graph neural network fault diagnosis device for process industry, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the above-mentioned graph neural network fault diagnosis method for process industry.
[0019] In the third aspect of the present invention, there is provided a computer-readable storage medium storing a computer program for executing the above-mentioned graph neural network fault diagnosis method for process industry.
[0020] Compared with the prior art, the benefits of the present invention are:
[0021] (1) The present invention fully integrates the static structure information from process knowledge and the dynamic features extracted from multi-sensor time series data, thereby constructing a topological graph that contains both explicit relationships and implicit process relationships.
[0022] Compared with the prior art, the present invention effectively overcomes the problems of traditional methods that overly rely on prior knowledge or pure data-driven, and has better interpretability, robustness, and fault diagnosis effect.
[0023] (2) The present invention utilizes the time series data contained in multi-sensor signal data and maps it to node representations through a graph convolutional neural network, thereby promoting the transmission of time information across the entire graph structure and effectively capturing the time dependence of multi-sensor signals. Description of the Drawings
[0024] Figure 1 It is the structure diagram of the two-layer graph convolutional neural network of the present invention;
[0025] Figure 2 It is the process flow diagram of the Tennessee Eastman process of the present invention;
[0026] Figure 3 It is the process topology diagram of the Tennessee Eastman process of the present invention;
[0027] Figure 4 It is the topology graph construction module of the present invention;
[0028] Figure 5 Structural diagram of the knowledge-enhanced graph embedding layer of the present invention;
[0029] Figure 6 Structural diagram of the graph neural network for process topology and time feature embedding of the present invention. Detailed implementation manners
[0030] To more clearly illustrate the technical solution of the present invention, the technical solution provided by the present invention will be further described below in conjunction with the accompanying drawings.
[0031] The present invention adopts a graph neural network (TTGCN) for process topology and time feature embedding. The fault diagnoser based on TTGCN takes process data and prior knowledge containing the process as inputs and the fault category as the output. A parallel structure is adopted in the network structure, which not only makes full use of process data to analyze historical data and mine potential rules, but also uses prior knowledge to constrain model construction. The two complement each other. Process data makes up for the deficiency of static prior knowledge and enhances the adaptability of the model under dynamic working conditions. Prior knowledge provides theoretical support and physical meaning explanation for model construction. In addition, one-dimensional convolution is combined with position embedding to extract key time series information from process data, and a learnable dynamic adjacency matrix is used to mine the time series features of variables, significantly improving the model's ability to capture dynamic features. The fault diagnosis performance of the model is improved and the interpretability is enhanced.
[0032] The specific content of the first aspect of the present invention is as follows:
[0033] Step S1: Collect industrial data applicable to the fault diagnosis task and organize it into a time series data set.
[0034] In some implementations, step S1 specifically includes:
[0035] Step S11: Collect the data set D = {X, Y} = {(x (m,t,i) , y m )} for the fault diagnosis method, where i ∈ 1, 2,..., N, N is the total number of variables, t ∈ 1, 2, 3,..., T, T is the length representing the time series data, that is, the total number of time steps, m ∈ 1, 2, 3,..., M, M represents the total number of fault types included in the data set; x (m,t,i) is the sample value of variable i at the t-th time step in the m-th fault type, and y m is the label of the m-th fault.
[0036] Step S12: Preprocess the collected data set;
[0037] Further, step S12 includes:
[0038] Step S121: Window the collected sequence data to obtain a series of shorter subsequences, where is expressed as follows:
[0039]
[0040] is the subsequence of the process variable corresponding to the m-th type of fault at the j-th time step, is the label of the m-th type of fault at the j-th time step, j ∈ seq_len, seq_len + 1, …, n, and seq_len is the window size;
[0041] Step S122: Randomly divide the subsequences in an 8:2 ratio to obtain a training set and a test set
[0042] Step S2: Prepare input data that meets the requirements for the graph neural network model.
[0043] In some embodiments, Step S2 specifically includes:
[0044] Step S21: Randomly initialize an adjacency matrix with all diagonal elements being 0 and the rest being 1 and sparsely represent it as a list of edge indices, one list containing the source node IDs and the other list containing the target node IDs.
[0045] Step S23: Combine the training set and the test set with edge_index respectively to form the data required for the graph neural network model and
[0046] Step S3: Construct a graph neural network model with process topology and time feature embedding, and perform supervised training on the constructed training set.
[0047] In some embodiments, Step S3 specifically includes:
[0048] Step S31: Construct a graph neural network model with process topology and time feature embedding
[0049] The graph neural network model designed in the present invention for process topology and time feature embedding mainly includes four parts from data input to the output of the fault diagnosis result, namely the topology graph construction module, the time feature embedding module, the knowledge-enhanced graph embedding layer, and the classification layer.
[0050] Furthermore, Step S31 includes:
[0051] Step S311: The topology graph construction module includes a process topology layer and a data perception layer, and its input is the process flow chart in Step S23.
[0052] Furthermore, Step S311 includes:
[0053] Step S3111: The process topology layer extracts the prior knowledge contained in the process flow chart and constructs a graph structure G that includes the process flow. K .
[0054] Since the pure data-driven fault diagnosis method lacks the support of domain knowledge, the variable relationships learned by the model may not conform to actual industrial production, resulting in unsatisfactory model training results. Therefore, the process topology layer is used to obtain knowledge such as material flow and chemical characteristics in the process flow chart to support the model training process, help the model understand the data, and enhance the interpretability of the model.
[0055] The prior knowledge in the process flow chart is extracted into the graph structure G according to the following steps K :
[0056] (1) Determine each basic operation unit, process variable, and operation variable in the chemical process;
[0057] (2) Determine the forward flow direction of the unit and each variable according to the process flow chart;
[0058] (3) Determine the control feedback loop according to mass conservation and energy conservation, and add a control unit to the flow chart;
[0059] (4) Create a directed edge according to the forward flow direction, and at the same time create a directed edge for the control loop, pointing from the control variable to the controlled variable or the control unit, and construct a unit flow graph;
[0060] (5) Simplify the unit flow graph, delete the basic operation unit, flow unit, and control unit, and retain the process variable and operation variable;
[0061] (6) Abstract the process variable and operation variable into nodes and construct them into the graph structure G K .
[0062] Then, the Pearson correlation coefficient is used to analyze the correlation degree between variables as the weight information of the edges between the nodes in G K in the graph.
[0063] The Pearson correlation coefficient is a statistical measure that measures the strength and direction of the linear relationship between two variables. The correlation value ranges from -1 to +1, where -1 represents a perfect negative correlation, +1 represents a perfect positive correlation, and 0 represents no correlation. The Pearson correlation coefficient is usually abbreviated as r, and the calculation formula is as follows:
[0064]
[0065] For the above formula, since the correlation coefficients need to be calculated among various variables therefore, i and j can respectively represent a variable. i s and j s represent the values of the variable at time s, and respectively represent the mean value of the variable, and l represents the length of the data set D.
[0066] The correlation coefficients among multiple variables can be calculated according to the formula Taking the correlation coefficient as the weight information between variables, that is, the attribute characteristics of the edge. The correlation matrix B is composed of the correlation coefficients calculated by Pearson, where n is the number of variables. Due to the relationship between variables in the graph structure G K if there is an edge between variables, the value is set to 1, and if there is no edge, the value is set to 0. Therefore, an adjacency matrix of 0 and 1 can be obtained Then for the graph structure G K the formula for calculating the weight information of the edges is as follows:
[0067]
[0068] A KW represents the weighted adjacency matrix obtained from the prior knowledge of the process, which contains the attribute characteristics of the edges, where ⊙ represents the product of the elements at the corresponding positions of the matrix. |B| represents taking the absolute value of each element in the matrix B. Since the direction of the edges in G K has been determined, therefore, only the weight information of the nodes needs to be known.
[0069] Step S3112: The data perception layer uses the sparse attention mechanism to extract the relationship between the edges in the process data, calculate the attribute characteristics of the edges, and is responsible for updating the index information and weight information between the edges in the process data.
[0070] The data perception layer inputs data selects any two variables x i and x j , and uses the graph-based attention mechanism to calculate the relationship between the two variables. The calculation formula is as follows:
[0071]
[0072] where σ(·) represents the activation function LeakReLU, represents the trainable parameter in the attention mechanism. CONCAT represents concatenating the variable features along the direction of the feature dimension.
[0073] The correlation between nodes with high scores is strong, and the correlation between nodes with low scores is weak. Therefore, to reduce the density of the graph structure and sparsify the edges of the graph structure, the sparsemax(·) function is used. Through the sparsification function, a threshold can be determined according to the constraint conditions to sparsify the weights of the edges. The edges with strong correlation are retained, and the edges with weak correlation are discarded. The weight of the edges with strong correlation is equal to the calculated correlation strength
[0074] According to the data perception layer, each time the sparse attention mechanism is used, a graph structure G can be extracted from the process data D , G D can represent the node relationship and the attribute characteristics of its edges with a weighted adjacency matrix A DW By using the sparse graph attention mechanism multiple times, multiple graph structures can be extracted That is where s ∈ {1, 2, …, n} represents the attribute of the s-th type of edge. Among these extracted graph structures, their nodes are the same, but the relationships between the edges of the nodes are different. There are multiple types of edge relationships, and finally a homogeneous heterogeneous graph G is formed H .
[0075] Step S3113: The process topology layer proposes a graph structure G based on the process flow chart K , which contains the explicit relationships in the chemical process, while the homogeneous heterogeneous graph G with multiple types of edges obtained by the data perception layer H contains the implicit relationships in the chemical process. Therefore, the graph structure G based on the process flow chart K and the homogeneous heterogeneous graph G extracted from the data are fused into a weighted directed homogeneous heterogeneous topology graph G H . KDW .
[0076] Step S312: The time feature embedding module includes a one-dimensional convolutional layer, position embedding, and feature fusion layer. Its input is the
[0077] In , the N-dimensional subsequence variables at the same moment are embedded into a vector of size q m : X t-L,t →X e , where The calculation formula is as follows
[0078]
[0079] where X t-L,t is normalized to 0 and 1 to obtain Then, one-dimensional convolution is used to Projected onto X e , with the parameter λ as a weight factor to improve the influence of the positional embedding and the one-dimensional convolutional embedding. Represents the positional embedding of the input X. The training process of the feature fusion layer includes two trainable parameter matrices, and Subsequently, relying on the two parameter matrices, an adaptive adjacency matrix A can be obtained adp . The adaptive adjacency matrix A adp Can learn the implicit relationships between nodes. The calculation formula is as follows:
[0080] A adp = Attention(P2(P1) T )
[0081] Use the attention function to determine the weights between different nodes. After obtaining A adp and the embedded feature vector X e , a feature fusion layer composed of residual-based graph convolutional layers is used to extract node features, and the calculation method is as follows:
[0082] H out1 = ReLU(A adp X e )
[0083]
[0084] Where H out1 Is the output of the first layer of graph convolutional layer, and H out2 Contains the extracted deep features, which contain the node features X N Regarding the timing information.
[0085] Step S313: The knowledge-enhanced graph embedding layer is composed of multiple layers of graph convolutional layers, responsible for fusing the G KDW Extracted by the topological graph construction module and the node features X N Of the time feature embedding module, and extracting deep fault features. Among them, there are various data-based edge features and knowledge-based edge features in G KDW .
[0086] First, fuse the edge attribute features of each type based on data and the edge attribute features extracted based on prior knowledge and send them into a type of graph convolutional layer to extract fault features. Therefore, various types of fault features can be extracted, and finally, after global pooling and local pooling, the fused and distinguishable fault features are obtained.
[0087]
[0088] Where A kAs a mathematical representation of the process knowledge topology structure, it is derived from the graph structure adjacency matrix of process knowledge extraction, A s As a mathematical representation of the process data topology structure, it is derived from the graph structure adjacency matrix extracted from process data. D represents the degree matrix of (A k +A s ), represents the normalized adjacency matrix. X N refers to the node features extracted by the node feature embedding module, H (i+1,s) represents the deep - level fault features obtained by combining the edge attributes of the s - th type after the i - th layer of convolution with the edge attributes and node features extracted from the process prior knowledge, X s , passing X s through the global average pooling layer and the global max pooling layer to obtain r (s,mean) and r (s,max) .
[0089] Step S314: The input data of the classification layer is r (s,mean) and r (s,max) . To maintain the independence of different edge - type representations, a node only aggregates the information of adjacent nodes with the same edge type. Therefore, one way to obtain a complete representation is to concatenate the vectors on all sub - graphs into a vector, and then concatenate the r (s,mean) and r (s,max) after multiple graph convolutions, and the result of the concatenation is result. The calculation formula is as follows:
[0090] result = CONCAT{r (1,max) ,r (1,mean) ,...,r (n,max) ,r (n,mean)} i ∈ {1,2,...,n}
[0091] Then, pass result through a two - layer fully - connected layer with dropout and the SoftMax layer to output the final classification result.
[0092] Step S32: Train the model on the constructed training set Set the hyper - parameters: window size seq_len, window step size window_size, number of iterations epoch, batch size batch_size, learning rate lr, number of graph convolution layers num in the knowledge - enhanced graph embedding layer, number of types of edges to be learned num_edge, and dropout rate.
[0093] Furthermore, step S32 includes:
[0094] Step S321: The training set
[0095] The process variable subsequence in is input into the graph neural network for process topology and time feature embedding and the corresponding fault diagnosis result is output to determine the fault type. When diagnosing the sample at time step j, the network inputs the process variable subsequence {(x (m,j-k+1,i) , x (m,j-k+2,i) , …, x (m,j,i) , edge_index)}, and the network outputs the diagnosis result of the corresponding subsequence's true fault label y j . The calculation is completed by the TTGCN network, and its calculation method is as follows:
[0096]
[0097] Among them, is the fault diagnosis result of the true fault label y j .
[0098] Step S322: Calculate the objective function and use the gradient descent algorithm to solve the model parameters that minimize the objective function, and store the process topology and time feature embedding network model and the optimal hyperparameters. The calculation method of the objective function is as follows:
[0099]
[0100] Among them represents the calculation for all samples in the training set, and m is the number of samples in the training set. Among them, C is the total number of categories, y j is the one-hot encoding of the true category, is the category probability predicted by the model, and satisfies the probability normalization constraint
[0101] Step S4: Load the trained model and test it on the divided test set . After completely traversing the test set samples, count the model evaluation metrics Accuracy and F1 score.
[0102] The calculation methods of the above evaluation metrics are as follows:
[0103]
[0104] Among them, TP refers to being predicted as positive and actually being positive; FP refers to being predicted as positive and actually being negative; TN refers to being predicted as negative and actually being negative; FN refers to being predicted as negative and actually being positive; Precision refers to the precision rate; Recall represents the recall rate.
[0105] Example: This example is applied to the industrial data set Tennessee Eastman process to achieve accurate diagnosis of fault types and effectively avoid safety accidents. The Tennessee Eastman process is a simulation model generated based on a real chemical industry production process. It is used to evaluate the real industrial process of fault diagnosis methods in process control. The process includes five units: reactor, condenser, compressor, separator and stripper. Figure 2 , which is a process flow chart of the Tennessee Eastman process. The Tennessee Eastman process involves many variables, and there is a strong correlation and coupling relationship between the variables. The process variables include 41 measured variables and 11 effective control variables, with a total of 52 variables. The Tennessee Eastman process has a total of 21 preset faults. This embodiment uses a total of 18 types of faults from two fault data sets for experimental research, excluding the 3rd, 9th, and 15th faults. Therefore, this embodiment has 1460 normal data, and each of the 18 types of faults has 1280 data. This embodiment divides the two fault data sets into a ratio of 8:2 in chronological order as training sets and test sets, respectively, samples by sliding windows to increase the number of data sets, and then mixes the two training sets and test sets to form the final training set and test set. This embodiment specifically includes the following steps under the example Tennessee Eastman process:
[0106] Step S1: Collect industrial data suitable for fault diagnosis tasks and organize them into a time series data set D;
[0107] Step S11: Collect the data set D = {X, Y} = {(x (m,t,i) ,y m )}, where i∈1,2,…,N, N=52 is the total number of variables, t∈1,2,3,…,T, T is the length of the time series data, that is, the total number of time steps, m∈1,2,3,…,M, M=18 is the total number of fault types contained in the data set; x (m,t,i) is the sample value of variable i at the tth time step in the mth fault type, y m is the label of the mth type of fault, represented by one-hot encoding;
[0108] Step S12: preprocessing the collected data set; wherein step S12 includes:
[0109] Step S121: Window processing is performed on the collected sequence data, with a window size of seq_len=64 and a window step of window_size=1, to obtain a series of shorter subsequences. in It is expressed as follows:
[0110]
[0111] is the subsequence of process variables corresponding to the m-th type of fault at the j-th time step, and is the label of the m-th type of fault at the j-th time step.
[0112] Step S122: Randomly divide the subsequence in an 8:2 ratio to obtain a training set and a test set where the training set contains 17,206 samples and the test set contains 2,506 samples.
[0113] Step S2: Prepare input data that meets the requirements for the graph neural network model.
[0114] Step S21: Randomly initialize an adjacency matrix with all diagonal elements being 0 and the rest being 1 and its sparse representation edge_index, one list containing source node IDs ∈ {0, 1, …, 51} and another list containing target node IDs ∈ {0, 1, …, 51}.
[0115] Step S23: Combine the training set and the test set with edge_index respectively to form the data required by the graph neural network model and
[0116] Step S3: Construct a graph neural network model for process topology and time feature embedding, and perform supervised training on the constructed training set;
[0117] Step S3 further includes the following steps:
[0118] Step S31: Construct a graph neural network model for process topology and time feature embedding
[0119] In this embodiment, the designed graph neural network model for process topology and time feature embedding mainly includes four parts from data input to fault diagnosis result output, namely the topology graph construction module, the time feature embedding module, the knowledge-enhanced graph embedding layer, and the classification layer.
[0120] Step S311: The topology graph construction module includes a process topology layer and a data perception layer, and its input is the Tennessee Eastman process flow chart in Step S23.
[0121] Step S3111: The process topology layer extracts the prior knowledge contained in the Tennessee Eastman process flow chart and constructs a graph structure G that includes the Tennessee Eastman process flow chart K .
[0122] Extract the prior knowledge in the process flow diagram into the graph structure G according to the following steps K :
[0123] (1) Determine each basic operation unit, process variable, and operation variable in the chemical process;
[0124] (2) Determine the forward flow direction of the unit and each variable according to the process flow diagram;
[0125] (3) Determine the control feedback loop according to mass conservation and energy conservation, and add control units to the process flow diagram, such as the feedback loop of the bottom flow unit of the separator, the feedback loop of the bottom flow unit of the stripper, and the feedback loop of the A feed;
[0126] (4) Create directed edges according to the forward flow direction, and at the same time create directed edges for the control loop, from the control variable to the controlled variable or the control unit, to construct the unit flow diagram;
[0127] (5) Simplify the unit flow diagram, delete the basic operation unit, flow unit, and control unit, and retain the process variable and operation variable;
[0128] (6) Abstract the process variable and operation variable into nodes to construct the graph structure G K .
[0129] Then, use the Pearson correlation coefficient to analyze the correlation degree between variables as the weight information of the edges between nodes in G K The Pearson correlation coefficient is a statistical measure that measures the strength and direction of the linear relationship between two variables. The correlation value ranges from -1 to +1, where -1 represents a perfect negative correlation, +1 represents a perfect positive correlation, and 0 represents no correlation. The Pearson correlation coefficient is usually abbreviated as r, and the calculation formula is as follows:
[0130]
[0131] For the above formula, since the correlation coefficients need to be calculated between variables Therefore, i and j can each represent a variable. i s and j s represent the values of the variable at time s, and respectively represent the mean values of the variables, and l represents the length of the data set D.
[0132] According to the formula, the correlation coefficients between multiple variables can be calculated The correlation coefficient As the weight information between variables, that is, the attribute characteristics of edges. The correlation coefficients calculated by Pearson form the correlation matrix B, where Due to the graph structure G K For the relationship between variables in it, if there is an edge between variables, the value is set to 1, and if there is no edge, the value is set to 0. Therefore, an adjacency matrix of 0 and 1 can be obtained Then G K The calculation formula for the weight information of the edges in it is as follows:
[0133]
[0134] A KW represents the weighted adjacency matrix obtained from the process prior knowledge, which contains the attribute characteristics of the edges. Here, ⊙ represents the product of the corresponding elements of the matrix. |B| represents taking the absolute value of each element in the matrix B. Since the direction of the edges in G K has been determined, therefore, only the weight information of the nodes needs to be known.
[0135] Finally, the complete topological graph structure of the Tennessee Eastman process is represented by G TE and it consists of 52 nodes and 76 edges. See Figure 3 for the topological graph of the Tennessee Eastman process shown.
[0136] Step S3112: The data perception layer uses the sparse attention mechanism to extract the relationship between the edges in the process data, calculate the attribute characteristics of the edges, and is responsible for updating the index information and weight information between the edges in the process data. The input data of this layer selects any two variables x i and x j , and uses the graph-based attention mechanism to calculate the relationship between the two variables. The calculation formula is as follows:
[0137]
[0138] where σ(·) represents the activation function LeakReLU, represents the trainable parameters in the attention mechanism. CONCAT represents concatenating the variable features along the direction of the feature dimension. The correlation between nodes with high scores is strong, and the correlation between nodes with low scores is weak. Therefore, in order to reduce the density of the graph structure and sparsify the edges of the graph structure, that is, the sparsemax(·) function is used. Through the sparsification function, a threshold can be determined according to the constraint conditions to sparsify the weights of the edges. The edges with strong correlation are retained, and the edges with weak correlation are discarded. The weight of the edges with strong correlation is equal to the calculated correlation strength
[0139] According to the data perception layer, each time the sparse attention mechanism is used, a graph structure G can be extracted from the process data D , G D can represent the node relationship and the attribute characteristics of its edges with a weighted adjacency matrix A DW In this embodiment, using the sparse graph attention mechanism twice can extract two graph structures That is where s ∈ {1, 2}. Among these extracted graph structures, their nodes are the same, but the relationships between the nodes are different, and there are various types of edge relationships, finally forming a homogeneous and heterogeneous graph G H .
[0140] Step S3113: The process topology layer proposes a graph structure G based on the Tennessee Eastman process flow chart K , which contains the explicit relationships in the chemical process, while the homogeneous and heterogeneous graph G with various types of edges obtained by the data perception layer H implies the implicit relationships in the chemical process. Therefore, the graph structure G based on the Tennessee Eastman process flow chart K and the homogeneous and heterogeneous graph G extracted based on the data H are fused into a weighted directed homogeneous and heterogeneous topological graph G KDW . See Figure 4 , which shows the topological graph construction module
[0141] Step S322: The time feature embedding module includes a one-dimensional convolutional layer, position embedding, and feature fusion layer. Its input is the
[0142] in step S122 Embed the N-dimensional subsequence variables at the same moment in into a vector of size q m : X t-L,t →X e , where The calculation formula is as follows
[0143]
[0144] Among them, X t-L,t is normalized to 0 and 1 to obtain Then use one-dimensional convolution to project
[0145] to X e , and the parameter λ is used as a weight factor to improve the influence of position embedding and one-dimensional convolution embedding represents the position embedding of the input X. The training process of the feature fusion layer includes two trainable parameter matrices and Subsequently, relying on two parameter matrices, an adaptive adjacency matrix A can be obtained. adp The adaptive adjacency matrix A adp can learn the implicit relationships between nodes. The calculation formula is as follows:
[0146] A adp = Attention(P2(P1) T )
[0147] The attention function is used to determine the weights between different nodes. After obtaining A adp and the embedded feature vector X e , a feature fusion layer composed of residual-based graph convolutional layers is adopted to extract node features, and the calculation method is as follows:
[0148] H out1 = ReLU(A adp X e )
[0149]
[0150] where H out1 is the output of the first layer of graph convolutional layer, and H out2 contains the extracted deep features, which contain the node features X N about the temporal information.
[0151] Step S313: The knowledge-enhanced graph embedding layer is composed of two layers of graph convolutional layers, which are responsible for fusing the G KDW extracted by the topological graph construction module and the node features X N in the time feature embedding module, and extracting deep fault features. Among them, there are two types of data-based edge features and one type of knowledge-based edge feature in G KDW . First, the edge attribute features of each type based on data and the edge attribute features extracted based on prior knowledge are fused and sent into a type of graph convolutional layer to extract fault features. Therefore, two types of fault features can be extracted, and finally, the fused and distinguishable fault features are obtained through global pooling and local pooling. See Figure 5 , which shows the structure diagram of the knowledge-enhanced graph embedding layer.
[0152]
[0153] s ∈ {1, 2}
[0154] where A k , as the mathematical representation of the process knowledge topological structure, is derived from the graph structure adjacency matrix extracted from process knowledge, and A s , as the mathematical representation of the process data topological structure, is derived from the graph structure adjacency matrix extracted from process data, and D represents the degree in the graph structure (Ak +A s )'s degree matrix represents the normalized adjacency matrix. X N refers to the node features extracted by the node feature embedding module, H (i+1,s) represents the deep fault features obtained by combining the edge attributes of the s-th type after the i-th layer of convolution with the edge attributes and node features extracted from the process prior knowledge X s , taking X s through the global average pooling layer and the global max pooling layer to obtain r (s,mean) and r (s,max) .
[0155] Step S314: The input data of the classification layer is r (s,mean) and r (s,max) . To maintain the independence of different edge type representations, the node only aggregates the information of adjacent nodes with the same edge type. Therefore, one way to obtain a complete representation is to concatenate the vectors on all subgraphs into a vector, and then concatenate r (s,mean) and r (s,max) after two graph convolutions. The concatenation result is result. The calculation formula is as follows:
[0156] result = CONCAT{r (1,max) , r (1,mean) , r (2,max) , r (2,mean)} i ∈ {1, 2}
[0157] Then pass result through two fully connected layers with dropout and the SoftMax layer to output the final classification result.
[0158] Step S32: Train the model on the constructed training set
[0159] Construct the model according to the graph neural network model structure based on process topology and time feature embedding. See Figure 6 . The figure shows the graph neural network structure diagram of process topology and time feature embedding. Set the hyperparameters window size seq_len = 64, window step size window_size = 1, number of iterations epoch = 300, batch size batch_size = 60, learning rate lr = 0.0005, number of graph convolution layers num = 2 in the knowledge-enhanced graph embedding layer, number of edge types to be learned num_edge = 2, and dropout = 0.3.
[0160] Step S321: Take the training set The process variable subsequence is input into the graph neural network with process topology and time feature embedding and the corresponding fault diagnosis result is output to determine the fault type. When diagnosing the sample at the j-th time step, the network inputs the process variable subsequence {(x (m,j-k+1,i) , x (m,j-k+2,i) , …, x (m,j,i) , edge_index)}, and the network outputs the diagnosis result of the corresponding subsequence's true fault label y j . The calculation is completed by the TTGCN network, and its calculation method is as follows:
[0161]
[0162] Among them, is the fault diagnosis result of the true fault label y j .
[0163] Step S322: Calculate the objective function and use the gradient descent algorithm to solve the model parameters that minimize the objective function, and store the process topology and time feature embedding graph neural network model and the optimal hyperparameters. The calculation method of the objective function is as follows:
[0164]
[0165] Among them represents the calculation for all samples in the training set, and m is the number of samples in the training set. Among them, C is the total number of categories, y j is the one-hot encoding of the true category, is the category probability predicted by the model, and satisfies the probability normalization constraint
[0166] Step S4: Use the graph neural network model with process topology and time feature embedding to test on the Tennessee Eastman process test set .
[0167] Step S4 includes the following steps:
[0168] Step S41: Create the stored graph neural network model with process topology and time feature embedding, and load the optimal hyperparameters.
[0169] Step S42: Input the test set into the graph neural network model with process topology and time feature embedding, and the model outputs the fault diagnosis result of the test set in the manner described in Step S32.
[0170] Step S43: Calculate the test set evaluation metrics Accuracy and F1 score, and the calculation methods are as follows:
[0171]
[0172] Among them, TP refers to being predicted as positive and actually being positive; FP refers to being predicted as positive and actually being negative; TN refers to being predicted as negative and actually being negative; FN refers to being predicted as negative and actually being positive; Precision refers to precision rate; Recall represents recall rate.
[0173] The embodiment of the present application also discloses a graph neural network fault diagnosis device for process industry, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, it implements the above-mentioned graph neural network fault diagnosis method for process industry.
[0174] The embodiment of the present application also discloses a computer-readable storage medium, the storage medium stores a computer program, and the computer program is used to execute the above-mentioned graph neural network fault diagnosis method for process industry.
[0175] Those of ordinary skill in the art can understand that all or part of the processes in implementing the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0176] The processor in this application may include one or more processing cores. By running or executing instructions, programs, code sets, or instruction sets stored in the memory, the processor calls the data stored in the memory and executes various functions of this application and processes data. The processor may be at least one of an Application Specific Integrated Circuit (ASIC), a Digital Signal Processor (DSP), a Digital Signal Processing Device (DSPD), a Programmable Logic Device (PLD), a Field Programmable Gate Array (FPGA), a Central Processing Unit (CPU), a controller, a microcontroller, and a microprocessor. It can be understood that for different devices, the electronic devices for implementing the above processor functions may also be others, and the embodiments of this application do not make specific limitations.
[0177] Those skilled in the art will readily conceive of other embodiments of this application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include known common knowledge or conventional technical means in the technical field not disclosed in this application. The specification and embodiments are only regarded as exemplary.
[0178] It should also be noted that the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity, or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, commodity, or device.
[0179] The above are only the preferred embodiments of this application and are not intended to limit this application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of this application shall be included within the scope of protection of this application.
Claims
1. A graph neural network fault diagnosis method for process industry, characterized in that: The following steps are involved: Step S1. Collect industrial data suitable for fault diagnosis tasks and organize them into a time series data set; Step S2. Prepare input data that meets the requirements for the graph neural network model, including randomly initializing an adjacency matrix with all diagonal elements as 0 and all other values as 1, and sparsely representing it as a list with edge indexes, and combining the training set and the test set with the list to form the data required by the graph neural network model; Step S3: construct a graph neural network model for embedding process topology and time features, and perform supervised training on the constructed training set; the model includes: The topology graph construction module includes a process topology layer and a data perception layer. The process topology layer extracts prior knowledge from the process flow chart to construct a graph structure. The data perception layer uses a sparse attention mechanism to extract the relationship between edges in the process data, updates the index information and weight information of the edges, and fuses the graph structure based on the process flow chart and the homogeneous and heterogeneous graph based on data extraction into a weighted directed homogeneous and heterogeneous topology graph. The temporal feature embedding module, which includes a one-dimensional convolutional layer, a position embedding layer, and a feature fusion layer, is used to extract key temporal information from process data and mine the temporal features of variables using a learnable dynamic adjacency matrix; The knowledge-enhanced graph embedding layer consists of multiple layers of graph convolutional layers, which is responsible for fusing the edge features extracted by the topology graph construction module and the node features of the time feature embedding module to extract deep fault features. The classification layer is used to output the fault category according to the extracted fault features; Step S4: Load the trained fault diagnosis model and test it on the test set to calculate the model evaluation indicators and scores.
2. A graph neural network fault diagnosis method for process industry according to claim 1, characterized in that: The step S1 specifically includes: Collect a data set for fault diagnosis modeling, where the data set contains time series data of multiple variables and corresponding fault type labels; The collected data set is preprocessed, including windowing the sequence data to obtain a series of shorter subsequences, and the subsequences are randomly divided into training sets and test sets according to the proportion.
3. A graph neural network fault diagnosis method for process industry according to claim 1 or 2, characterized in that: In the topology graph construction module, the specific steps of extracting prior knowledge in the process flow graph and constructing the graph structure at the process topology layer include: Determine each basic operating unit, process variable and operating variable in the chemical process; Determine the forward flow direction of units and variables according to the process flow diagram; Determine the control feedback loop based on the conservation of mass and energy, and add a control unit to the process flow chart; Create directed edges according to the flow direction, and create directed edges for the control loop to build a unit flow graph; Simplify the unit flow diagram, delete the basic operation unit, flow unit and control unit, and retain the process variables and operation variables; The process variables and operation variables are abstracted into nodes and constructed into a graph structure.
4. A graph neural network fault diagnosis method for process industry according to claim 3, characterized in that: When the data perception layer uses the sparse attention mechanism to extract the relationship between edges, the two variables x are calculated by the following formula i and x j The relationship between: Where σ(·) represents the activation function LeakReLU, It represents the trainable parameters in the attention mechanism, and CONCAT represents the concatenation of variable features along the feature dimension.
5. The graph neural network fault diagnosis method for process industry according to claim 3 is characterized in that: In the knowledge-enhanced graph embedding layer, edge attribute features based on each type of data and edge attribute features extracted based on prior knowledge are fused and sent to a type of graph convolution layer to extract fault features, and fused fault features with differentiation are obtained through global pooling and local pooling.
6. The graph neural network fault diagnosis method for process industry according to claim 1 is characterized in that: The classification layer outputs the fault category through the following steps: The fused fault features are input into two fully connected layers, and the probability distribution of fault categories is output; According to the probability distribution, the category with the highest probability is selected as the final fault diagnosis result.
7. A graph neural network fault diagnosis method for process industry according to claim 1 or 6, characterized in that: During the model training process, the cross entropy loss function is used as the objective function, and the model parameters are optimized by the gradient descent algorithm.
8. A graph neural network fault diagnosis method for process industry according to claim 1 or 6, characterized in that: In the test step, Accuracy and F1 scores are used as evaluation indicators to evaluate the fault diagnosis performance of the model.
9. A graph neural network fault diagnosis device for process industry, characterized by include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a graph neural network fault diagnosis method for process industry as described in any one of claims 1 to 8 is implemented.
10. A computer-readable storage medium, characterized in that The storage medium stores a computer program, and the computer program is used to execute the graph neural network fault diagnosis method for process industry as described in any one of claims 1 to 8.
Citation Information
Cited By
Motor fault diagnosis method and system based on hybrid graph neural network and path graph
CN120873764A