Industrial equipment fault detection method fusing complex relation and space-time dependence
By constructing a multidimensional adjacency matrix and performing weighted fusion, combining GCN and random GAT to extract spatiotemporal features, and adopting multiple node-level binary classifiers and a voting mechanism, the problems of the existing technology that cannot fully capture the complex multidimensional relationships between devices and the insufficient fusion of spatiotemporal features are solved, and industrial equipment fault detection with higher accuracy and robustness is achieved.
Patent Information
- Application Number
- CN202510628363.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-09-26
AI Technical Summary
Existing spatiotemporal graph neural networks cannot fully capture the complex multidimensional relationships between devices in industrial equipment fault detection. They lack an effective spatiotemporal feature fusion mechanism, and a single global classifier is difficult to perform personalized analysis, resulting in insufficient detection accuracy and robustness.
A multi-dimensional adjacency matrix is constructed and weightedly fused. Spatial features are extracted by combining GCN and random GAT. Temporal features are extracted through temporal convolution and multi-head attention mechanism. Multiple node-level binary classifiers and voting mechanism are used for comprehensive judgment. Spatiotemporal features are dynamically fused to generate graph-level features.
It improves the accuracy, stability and reliability of industrial equipment fault detection, reduces the risk of misjudgment, and enhances the ability to judge the overall status of industrial control systems.
Smart Images

Figure CN120705726A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial anomaly detection, and in particular to an industrial equipment fault detection method integrating complex relationships and spatiotemporal dependencies. Background Art
[0002] Accurate monitoring of industrial anomalies is crucial for ensuring production, not only reducing system response time but also providing proactive guidance for equipment maintenance and process operations. Modeling industrial big data as dynamic graphs with dynamic properties and structures can more accurately reflect the various state changes and interactions between elements in the industrial production process, better handle complex spatiotemporal relationships, and promptly identify abnormal fluctuations and potential failures in the production process.
[0003] However, existing spatiotemporal graph neural networks still have the following defects in industrial equipment fault detection applications: on the one hand, existing methods mostly use a single type of adjacency matrix to describe the relationship between nodes, which cannot fully capture the complex multidimensional relationships between industrial equipment, especially the various relationships such as distance dependence, topological order dependence and correlation dependence between equipment; on the other hand, existing spatiotemporal graph neural networks often focus on one aspect of temporal features or spatial features, lack an effective dynamic fusion mechanism, and find it difficult to balance and fully utilize information in both time and space dimensions; in addition, at the state discrimination level, existing technologies mostly use a single global classifier for discrimination, which cannot perform personalized analysis of the characteristics of different nodes, and lack an effective comprehensive judgment mechanism for multi-node classification results, which can easily lead to misjudgment. Especially in the face of complex and changeable industrial environments, detection accuracy and robustness are difficult to guarantee. Summary of the Invention
[0004] In response to the shortcomings of the prior art, the present invention aims to provide an industrial equipment fault detection method that integrates complex relationships and spatiotemporal dependencies, solving the technical problem of constructing a reasonable multidimensional adjacency matrix and weighted fusion of it when there are multidimensional relationships between devices, effectively extracting and fusing features in the spatiotemporal dimension, and accurately and efficiently performing equipment anomaly detection. By constructing multiple adjacency matrices (distance adjacency matrix based on Gaussian kernel, topological order matrix, node correlation matrix, etc.) and weighted fusion to form an enhanced adjacency matrix, the complex multidimensional relationships between industrial equipment are comprehensively and accurately described; a spatiotemporal feature extraction module is designed, spatial features are extracted in parallel using GCN and random GAT, temporal features are extracted through temporal convolution and multi-head attention mechanism, and graph-level features are generated by dynamically fusing spatiotemporal features with the help of a gating mechanism; a state discrimination layer composed of multiple node-level binary classifiers and a voting mechanism are used to comprehensively judge the classification results of each node, thereby enhancing the stability and reliability of the overall state judgment of the industrial control system and reducing the risk of misjudgment; a reasonable node anomaly probability threshold τ and a full-graph anomaly judgment threshold δ are set to more accurately judge the system state, thereby improving the accuracy, stability and reliability of anomaly detection in the industrial control system.
[0005] In order to solve the above technical problems, the technical solution provided by the present invention is:
[0006] A method for industrial equipment fault detection that integrates complex relationships and spatiotemporal dependencies includes the following steps:
[0007] S1. Obtain multidimensional time series data for equipment and construct a distance adjacency matrix, a topological order matrix, and a node correlation matrix based on a Gaussian kernel to capture the influencing factors of industrial equipment node relationships. This matrix is then weighted and fused to form an enhanced adjacency matrix that describes the multidimensional relationships between equipment.
[0008] S2. Construct a spatiotemporal feature extraction module, input the multidimensional time series data of the device and the enhanced adjacency matrix described in S1, extract the temporal and spatial features of the device nodes, and use a gating mechanism to dynamically adjust the contribution ratio of temporal and spatial features, dynamically fusing temporal and spatial features to generate high-dimensional graph-level features;
[0009] S3. Each device node is equipped with a binary classifier based on the Sigmoid activation function. Multiple node-level binary classifiers form a state discrimination layer. The high-dimensional graph-level features generated in S2 are input into this state discrimination layer, which independently determines whether the corresponding device node's state is abnormal or normal. A voting mechanism is used to combine the classification results of each device node to form an industrial equipment fault detection model that integrates complex relationships and spatiotemporal dependencies. This model then determines whether the overall state of the industrial control system is abnormal.
[0010] S4. Use precision, recall, and F1 score performance metrics to evaluate and optimize the detection capability of the model described in S3 by adjusting the binary classifier parameters and voting weight distribution, and combining cross-validation. The precision measures the proportion of samples predicted to be abnormal that are actually abnormal; the recall measures the proportion of samples that are correctly identified as abnormal; and the F1 score is the harmonic average of the precision and recall.
[0011] Preferably, in S1:
[0012] The characteristics of the factors affecting the relationship between industrial equipment nodes are:
[0013] The spatial correlation between devices is captured by the distance adjacency matrix of the Gaussian kernel;
[0014] Capture the direct interaction relationship between devices through the topological order matrix;
[0015] The node correlation matrix is used to capture the time series similarity, anomaly similarity, and statistical feature similarity of data between devices.
[0016] Preferably, the S1 is specifically:
[0017] S1.1 Constructing a distance adjacency matrix based on Gaussian kernel
[0018] The Gaussian kernel function is used to convert the physical distance between devices into the similarity between nodes, capturing the spatial correlation between devices. The specific definition is as follows:
[0019]
[0020] Among them, h i and h j is the spatial location feature of node i and node j; σ is the bandwidth parameter of the Gaussian kernel function, which controls the sensitivity of similarity to distance changes; α ij is the similarity between node i and node j;
[0021] S1.2 Constructing a topological order matrix
[0022] The topological order matrix is used to model a clear physical or logical connection structure, reflecting the direct interaction relationship between devices. It is defined as follows:
[0023]
[0024] S1.3 Constructing node correlation matrix
[0025] The node correlation matrix describes the similarity between nodes by analyzing the time series, abnormal data and statistical characteristics of the nodes; the node correlation matrix is subdivided into a time series similarity matrix, an abnormality similarity matrix, and a statistical feature similarity matrix;
[0026] S1.4 Weighted fusion of the Gaussian kernel-based distance adjacency matrix, topological order matrix and node correlation matrix to form an enhanced adjacency matrix.
[0027] Preferably, the step S1.3 of constructing a node correlation matrix is as follows:
[0028] S1.3.1 Constructing a time series similarity matrix
[0029] Dynamic Time Warping measures the similarity of the changing patterns of time series generated by industrial equipment by aligning them. It is defined as follows:
[0030] D ij =DTW(X i ,X j ) (3)
[0031] Among them, D ij is the time series similarity; X i and X j Represent the time series data of node i and node j respectively. DTW() represents the dynamic time warping algorithm, which calculates the alignment distance between sequences. The value of the DTW matrix is between 0 and 1. The smaller the value, the more similar the time series of the two nodes are, and the larger the value, the lower the similarity of the time series of the two nodes.
[0032] S1.3.2 Constructing anomaly similarity matrix
[0033] By analyzing historical anomaly data, quantifying the anomaly similarity between nodes, and capturing the law of anomaly propagation, the definition is as follows:
[0034]
[0035] Among them, A ij AnomalyCount is the abnormal similarity; ij AnomalyCount represents the number of common anomalies that occur between nodes i and j within a certain time window. i and AnomalyCount j are the total number of abnormalities of node i and node j respectively; A ij ∈[0,1], the higher the value, the higher the similarity between node i and node j in abnormal behavior;
[0036] S1.3.3 Constructing statistical feature similarity matrix
[0037] The statistical characteristics of a device can reflect its operating mode. The statistical characteristics similarity matrix is used to quantify the similarity of the statistical characteristics of the devices. It is defined as follows:
[0038]
[0039] Among them, S ij is the statistical feature similarity; and Represent the values of node i and node j on the kth statistical feature respectively, and M is the total number of statistical features. The statistical feature similarity matrix measures the differences between nodes on multiple statistical features. The smaller the difference, the higher the similarity.
[0040] The definition of weighted fusion of multiple adjacency matrices in S1.4 is as follows:
[0041]
[0042] in, Represents the enhanced adjacency matrix elements between node i and node j; each matrix (α ij 、T ij 、D ij 、A ij 、S ij ) are used to control the contribution of each matrix to the final adjacency matrix; w1 reflects the importance of spatial relationships, w2 reflects the influence of topological connections, w3 reflects the similarity weight of time series patterns, w4 reflects the importance of anomaly similarity, and w5 reflects the weight of statistical feature similarity.
[0043] Preferably, in S2:
[0044] Spatial feature extraction uses graph convolutional neural networks and random graph attention networks to extract spatial features of different dimensions in parallel, capturing local and global spatial relationships between devices in parallel;
[0045] Temporal feature extraction captures dependencies and dynamic behaviors in time series through temporal convolution and multi-head attention mechanisms;
[0046] Temporal features and spatial features are fused dynamically and weighted through a gating mechanism to generate high-dimensional graph-level features.
[0047] Preferably, the S2 is specifically:
[0048] S2.1 Spatial feature extraction
[0049] S2.1.1 Extracting spatial features using graph convolutional neural networks
[0050] GCN is used to capture local fixed spatial dependencies in industrial information graphs, adjacency matrix The normalized form of balances the feature propagation between nodes to avoid numerical explosion or disappearance, as shown in the following formula:
[0051]
[0052] in, is the node feature matrix of the l+1th layer; is the normalized augmented adjacency matrix, D t is the degree matrix, is the node feature matrix of the lth layer, W (l) is the learnable weight matrix;
[0053] The final output of GCN is:
[0054]
[0055] Among them, H (GCN,t) is the final output feature matrix of GCN; is the dimension space of the output feature; L represents the number of layers of GCN, and the output feature dimension is d out ;
[0056] S2.1.2 Spatial Feature Extraction by Random Graph Attention Network
[0057] The random graph attention network is introduced to capture dynamic non-uniform spatial dependencies. Random GAT uses a randomly generated attention matrix Combined with the graph structure mask function mask(·), the relationship weights between nodes are dynamically adjusted; the attention weights are generated as follows:
[0058]
[0059] in, is the inter-node weight matrix randomly generated for the i-th attention head at time step t; the mask function mask(·) is used to retain the edges existing in the industrial graph structure and filter out invalid edges; is the normalized dynamic weight matrix between nodes;
[0060] Calculate the learnable attention coefficient between nodes based on the random attention matrix To capture node v i and v j The dependency of is normalized with softmax to obtain the attention weight matrix between nodes:
[0061]
[0062] in, is node v i Neighborhood sets in industrial graphs;
[0063] Using the attention weight matrix between nodes Update node characteristics:
[0064]
[0065] Where σ(·) is the activation function and W is the learnable weight matrix; For node v j Initial feature representation of ; For node v i Updated features;
[0066] Introducing a multi-head mechanism to enhance learning stability and feature expression capabilities:
[0067]
[0068] Where K is the number of attention heads, Indicates concatenating the results of K heads, W k is the linear transformation weight matrix of the kth attention head;
[0069] The hidden state update formula of random GAT is:
[0070]
[0071] in, is the hidden state matrix of the i-th attention head at time step t; H t is the input feature matrix at time step t, and W is the learnable weight matrix;
[0072] After introducing the multi-head mechanism, the final characteristic matrix of random GAT is:
[0073]
[0074] Among them, H (R,t) is the final feature matrix of random GAT at time step t; is the weight matrix of the kth attention head;
[0075] S2.1.3 Spatial Feature Fusion
[0076] The features extracted by random GAT are concatenated with the features extracted by GCN in the feature dimension to form the final spatial feature representation:
[0077] H (spatial,t) =concat(H (GCN,t) ,H (r,t) ,dim=-1) (15)
[0078] The feature dimension after splicing is where d out is the output feature dimension of GCN, and D is the output feature dimension of GAT;
[0079] S2.2 Temporal feature extraction
[0080] The temporal feature extraction module is used to capture the dependencies and dynamic behavior changes in the node time series; the query matrix Q, key matrix K and value matrix V are calculated through temporal convolution, combined with the enhanced adjacency matrix Achieve effective modeling of time series characteristics;
[0081] First, calculate the attention weights:
[0082]
[0083] in, The query, key, and value matrices are calculated through temporal convolution, d head It is the single-head attention dimension;
[0084] Then perform softmax normalization on the fusion result to obtain the weighted attention weight:
[0085] P=softmax(S),P∈R B×N×T×T (17)
[0086] Finally, the output features are calculated by combining the value matrix V:
[0087]
[0088] The temporal features are further processed by the feedforward network FNN and batch normalization BN to extract nonlinear features and stabilize the training process:
[0089] H (temporal,t) =BN(FFN(H (temporal,t) )) (19)
[0090] S2.3 Fusion of temporal and spatial features
[0091] Spatiotemporal features are dynamically weighted and fused through a gating mechanism, thereby dynamically adjusting the contribution ratio of spatial and temporal features according to the characteristics of the industrial scenario;
[0092] First, the gate weight g is generated through two layers of fully connected layer networks and activation functions spatial and g temporal :
[0093] g spatial =σ(W s H spatial,t +b s ),g temporal =σ(W t H′ temporal,t +b t ) (20)
[0094] Among them, W s 、W tand b s 、b t It is a learnable parameter trained based on industrial data. σ represents the sigmoid activation function, which ensures that the weight value is between [0, 1].
[0095] Weight the features according to the gating weights:
[0096] H (fusion,t) =g spatial ⊙H (spatial,t) +g temporal ⊙H (temporal,t) (twenty one)
[0097] Among them, ⊙ represents element-wise multiplication;
[0098] The fused features are further subjected to nonlinear transformation and normalization to obtain the final fused features:
[0099] H (final,t) =BN(ReLU(FC(H (fusion,t) ))) (twenty two)
[0100] FC is the fully connected layer, BN represents batch normalization, and ReLU is the activation function; the final graph-level feature H (final,t) Used for system-wide status monitoring and anomaly detection.
[0101] Preferably, in S3:
[0102] The input of the state discrimination layer is the feature vector corresponding to the node, and the output is the probability value of whether the node is abnormal. The output formula of the classifier is as follows:
[0103]
[0104] in, is the classification output of the node; h i represents the feature vector of node i, W i and b i is the classifier parameter, σ is the Sigmoid activation function, which is used to map the classification results to the interval [0,1];
[0105] The voting mechanism is used to aggregate the output results of all node classifiers to determine the overall state of the graph;
[0106] First, the classification output of all nodes Perform binarization to determine whether the node anomaly probability threshold τ is exceeded. Then, calculate the proportion of abnormal nodes and compare it with the overall graph anomaly determination threshold δ:
[0107]
[0108] Wherein, I(·) represents an indicator function, and when the condition is true, I(·) = 1, otherwise I(·) = 0.
[0109] Preferably, the S4 is specifically:
[0110] The performance indicators precision, recall and F1 score are used for classification evaluation and optimization. The calculation formula is as follows:
[0111]
[0112] Among them, TP, FP, and FN represent the positive samples predicted by the model as positive, the negative samples predicted by the model as negative, and the positive samples predicted by the model as negative, respectively; Precision is the precision rate; Recall is the recall rate; and F1_Score is the F1 score.
[0113] The beneficial effects of the present invention are:
[0114] In terms of equipment relationship modeling, the present invention constructs multiple adjacency matrices and weightedly fuses them to form an enhanced adjacency matrix, which can comprehensively and accurately describe the complex multidimensional relationships between industrial equipment, effectively solving the difficult problems of equipment relationship modeling and multidimensional information fusion, and improving the accuracy and adaptability of anomaly detection.
[0115] Traditional methods for spatiotemporal feature extraction often fail to fully exploit the spatiotemporal dependencies of industrial equipment states. The present invention's spatiotemporal feature extraction module, through a unique structural design, simultaneously captures both spatial and temporal dependencies and utilizes a gating mechanism to dynamically adjust feature contribution. This allows for more precise feature extraction and fusion, resulting in more accurate anomaly detection results.
[0116] In terms of state discrimination, the present invention adopts a state discrimination layer and voting mechanism composed of multiple node-level binary classifiers. Compared with the single discrimination method of some existing technologies, it can better integrate the information of each node, enhance the stability and reliability of the overall state judgment of the industrial control system, reduce the risk of misjudgment, and effectively improve the overall performance of the anomaly detection system. BRIEF DESCRIPTION OF THE DRAWINGS
[0117] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:
[0118] Figure 1 This is a text flow chart of an industrial equipment fault detection method that integrates complex relationships and spatiotemporal dependencies.
[0119] Figure 2 A framework diagram of an industrial equipment fault detection method that integrates complex relationships and spatiotemporal dependencies DETAILED DESCRIPTION
[0120] The following describes preferred embodiments of the present invention in conjunction with the accompanying drawings. It should be understood that the following embodiments are provided for illustrative purposes only and are not intended to limit the scope of the present invention. Those skilled in the art may make various modifications and substitutions to the present invention without departing from the purpose and spirit of the present invention.
[0121] The present invention provides an industrial equipment fault detection method that integrates complex relationships and spatiotemporal dependencies, such as Figure 1 As shown in the figure, by constructing a variety of adjacency matrices (distance adjacency matrix based on Gaussian kernel, topological order matrix, node correlation matrix, etc.) and weighted fusion to form an enhanced adjacency matrix, the complex multi-dimensional relationship between industrial equipment is fully and accurately described, and the multi-dimensional time series data of industrial equipment (such as temperature, pressure, flow, etc.) is input into the spatiotemporal feature extraction module together; the spatiotemporal feature extraction module uses graph convolutional neural network (GCN) and random graph attention network (GAT) to extract spatial features in parallel, extracts time features through temporal convolution (tconv) and multi-head attention mechanism, and dynamically fuses spatiotemporal features with the help of gating mechanism to generate graph-level features; a state discrimination layer composed of multiple node-level binary classifiers and a voting mechanism are used to make a comprehensive judgment on the classification results of each node, enhance the stability and reliability of the overall state judgment of the industrial control system, and reduce the risk of misjudgment; reasonably set the node anomaly probability threshold τ and the full graph anomaly judgment threshold δ to more accurately judge the system state. As shown in the figure, the complex multi-dimensional relationship between industrial equipment is fully described, and the multi-dimensional time series data (such as temperature, pressure, flow, etc.) of industrial equipment are input into the spatiotemporal feature extraction module; the spatiotemporal feature extraction module uses graph convolutional neural network (GCN) and random graph attention network (GAT) to extract spatial features in parallel, extracts time features through temporal convolution (tconv) and multi-head attention mechanism, and dynamically fuses spatiotemporal features with the help of gating mechanism to generate graph-level features; a state discrimination layer and voting mechanism composed of multiple node-level binary classifiers are used to make a comprehensive judgment on the classification results of each node, enhance the stability and reliability of the overall state judgment of the industrial control system, and reduce the risk of misjudgment; reasonably set the node anomaly probability threshold τ and the full graph anomaly judgment threshold δ to more accurately judge the system state. Figure 2 As shown in the figure, it is a framework diagram of the industrial equipment fault detection method that integrates complex relationships and spatiotemporal dependencies.
[0122] The present invention provides an industrial equipment fault detection method that integrates complex relationships and spatiotemporal dependencies, specifically:
[0123] 1. Definition and integration of multi-dimensional adjacency matrix in industrial scenarios
[0124] In industrial scenarios, the relationships between devices are influenced by a variety of factors, including spatial location, topology, temporal behavior patterns, and operational characteristics. To accurately represent the complex relationships between nodes, we construct a Gaussian kernel-based distance adjacency matrix, a topological order matrix, and a node correlation matrix to capture the influencing factors of the relationships between devices. We then use weighted fusion to form an enhanced adjacency matrix, which comprehensively and accurately describes the complex, multidimensional relationships between industrial devices.
[0125] The following defines the construction of the Gaussian kernel-based distance adjacency matrix, topological order matrix, node correlation matrix, and the final weighted adjacency matrix:
[0126] 1.1 Constructing a distance adjacency matrix based on Gaussian kernel
[0127] In industrial scenarios, devices are usually arranged according to their physical locations. For example, some sensors or devices may be located in adjacent locations, and the distance between them reflects the spatial correlation between the devices.
[0128] The Gaussian kernel function is used to convert the physical distance into the similarity between nodes, capturing the spatial correlation between devices. The specific definition is as follows:
[0129]
[0130] Among them, h i and h j is the spatial location feature of node i and node j; σ is the bandwidth parameter of the Gaussian kernel function, which controls the sensitivity of similarity to distance changes; α ij is the similarity between node i and node j.
[0131] 1.2 Constructing a topological order matrix
[0132] In industrial systems, devices are typically connected via networks, pipelines, or control signals. These connections reflect the direct interactions between devices. The topological order matrix is used to model explicit physical or logical connection structures and is the most basic adjacency matrix in industrial scenarios. It is defined as follows:
[0133]
[0134] 1.3 Constructing the node correlation matrix
[0135] The node correlation matrix describes the similarity between nodes by analyzing the time series, abnormal data, and statistical characteristics of the nodes. The node correlation matrix is subdivided into the following types:
[0136] 1.3.1 Time Series Similarity Matrix
[0137] Industrial equipment often generates time series data (such as temperature, pressure, or flow). Even if the time series of two devices are offset on the time axis, they may still have similar change patterns. Dynamic Time Warping (DTW) measures this similarity by aligning the time series. The specific definition is as follows:
[0138] D ij =DTW(X i ,X j ) (3)
[0139] Among them, X i and X jwhere represents the time series data for node i and node j, respectively. DTW() represents the dynamic time warping algorithm, which calculates the alignment distance between the sequences. The values of the DTW matrix are typically between 0 and 1. Smaller values indicate greater similarity between the time series of the two nodes, while larger values indicate less similarity.
[0140] 1.3.2 Anomaly Similarity Matrix
[0141] In industrial scenarios, different nodes may have similarities in abnormal patterns, such as deviations in equipment operating status, abnormal sensor signals, or failure modes.
[0142] By analyzing historical anomaly data and quantifying the anomaly similarity between nodes, the model can better capture the laws of anomaly propagation.
[0143]
[0144] Among them, AnomalyCount ij AnomalyCount represents the number of common anomalies that occur between nodes i and j within a certain time window. i and AnomalyCount j are the total number of abnormalities of node i and node j respectively. ij ∈[0,1], a higher value indicates a higher similarity between node i and node j in abnormal behavior.
[0145] 1.3.3 Statistical feature similarity matrix
[0146] The statistical characteristics of a device (such as average daily temperature, maximum flow rate, etc.) can reflect its operating mode. If two devices have similar statistical characteristics, they may have the same function or working status. The statistical characteristic similarity matrix is used to quantify the similarity of the statistical characteristics of the devices. The specific definition is as follows:
[0147]
[0148] in, and Denote the values of node i and node j on the kth statistical feature, respectively, and M is the total number of statistical features. The statistical feature similarity matrix measures the differences between nodes on multiple statistical features. The smaller the difference, the higher the similarity.
[0149] 1.4 Weighted Fusion Adjacency Matrix
[0150] In industrial scenarios, the relationships between devices are influenced not only by the physical space but also by their operating status, temporal behavior patterns, and system characteristics. A single adjacency matrix cannot fully capture these complex relationships. The weighted fusion adjacency matrix combines the strengths of multiple adjacency matrices to dynamically express the multidimensional relationships between nodes, enabling more accurate capture of the propagation patterns of anomalies in industrial systems. The weighted matrix is defined as follows:
[0151]
[0152] in, Represents the enhanced adjacency matrix element between node i and node j. Each matrix (α ij 、T ij 、D ij 、A ij 、S ij ) are used to control the contribution of each matrix to the final adjacency matrix. Specifically, w1 reflects the importance of spatial relationships (such as the impact of device layout on failures), w2 reflects the influence of topological connectivity (such as coupling between directly physically connected devices), w3 reflects the similarity weight of time series patterns (such as the operational synergy between nodes), w4 reflects the importance of anomaly similarity (such as whether devices exhibit similar abnormal behavior), and w5 reflects the weight of statistical feature similarity (such as similar operating conditions or physical characteristics).
[0153] 2. Spatiotemporal feature extraction
[0154] A spatiotemporal feature extraction module was designed to extract temporal and spatial features, capturing the spatial and temporal dependencies of nodes in industrial information graphs. Using a gating mechanism, the module dynamically integrates temporal and spatial features to generate graph-level features for monitoring the operational status of the entire industrial control system. The module inputs multidimensional time series data and an enhanced adjacency matrix of devices, and outputs high-dimensional graph-level features that effectively reflect the overall status and abnormal behavior of the industrial system.
[0155] In industrial control systems, the states of devices are not only affected by the individual devices themselves but also propagate across space and time through interactions between adjacent devices. Therefore, the normal or abnormal state of a node can affect the state of the entire system through interactions with other nodes. To fully leverage these characteristics, the spatiotemporal feature extraction module is implemented using three components: the spatial feature extraction component uses GCN (graph convolutional network) and random GAT (graph attention network) to capture local and global spatial relationships between devices in parallel; the temporal feature extraction component uses temporal convolution and multi-head attention mechanisms to capture dependencies and dynamic behaviors in time series; and finally, feature fusion dynamically weights the spatial and temporal features through a gating mechanism to generate the final graph-level features.
[0156] 2.1 Spatial Feature Extraction
[0157] The spatial feature extraction module aims to extract spatial dependencies between nodes from the enhanced adjacency matrix, including local dependencies and dynamic interactions. By extracting spatial features of different dimensions in parallel using GCN and random GAT, the module is able to take into account both topological structures and complex relationships.
[0158] 2.1.1 GCN Extracts Spatial Features
[0159] GCN is mainly used to capture local fixed spatial dependencies in industrial information graphs. The normalized form of can balance the feature propagation between nodes and avoid numerical explosion or disappearance.
[0160]
[0161] in, is the normalized augmented adjacency matrix, D t is the degree matrix, is the node feature matrix of the lth layer, W (l) is the learnable weight matrix.
[0162] The final output of GCN is:
[0163]
[0164] Among them, L represents the number of layers of GCN, and the output feature dimension is d out .
[0165] 2.1.2 Random GAT Extraction of Spatial Features
[0166] In industrial control systems, the spatial dependencies between devices or sensor nodes can be dynamically affected by a variety of factors, including real-time changes in industrial processes, the heterogeneity of sensor data, and the collaborative relationships between devices. Traditional fixed adjacency matrices can be difficult to fully express these complex relationships, so a random graph attention network (Random GAT) is introduced to capture dynamic, heterogeneous spatial dependencies.
[0167] The core idea of random GAT is to generate a randomly generated attention matrix Combined with the graph structure mask function mask(·), the relationship weights between nodes are dynamically adjusted. The attention weights are generated as follows:
[0168]
[0169] in, is the inter-node weight matrix randomly generated for the i-th attention head at time step t; the mask function mask(·) is used to retain the edges existing in the industrial graph structure (such as the equipment connections determined by the industrial process) and filter out invalid edges. is the normalized dynamic weight matrix between nodes.
[0170] Based on the random attention matrix, further calculate the learnable attention coefficient between nodes To capture node v i and v j The dependency of is normalized with softmax to obtain the attention weight matrix between nodes:
[0171]
[0172] in, is node v i Neighborhood collection in industrial graphs.
[0173] Using the attention weight matrix between nodes Update node characteristics:
[0174]
[0175] Among them, σ(·) is the activation function, For node v i Updated features.
[0176] In order to improve the stability of learning and adapt to the complexity of industrial data, a multi-head mechanism is introduced. The multi-head mechanism further enhances the stability of learning and the ability to express features:
[0177]
[0178] Where K is the number of attention heads, Indicates concatenating the results of K heads, W k is the linear transformation weight matrix of the kth attention head.
[0179] The hidden state update formula of random GAT is:
[0180]
[0181] Among them, H t is the input feature matrix at time step t, and W is the learnable weight matrix.
[0182] After introducing the multi-head mechanism, the final characteristic matrix of random GAT is:
[0183]
[0184] in, is the weight matrix of the kth attention head.
[0185] 2.1.3 Spatial Feature Fusion
[0186] The features extracted by random GAT are concatenated with the features extracted by GCN in the feature dimension to form the final spatial feature representation:
[0187] H (spatal,t) =concat(H (GCN,t) ,H (R,t) ,dim=-1) (15)
[0188] The feature dimension after splicing is where d out is the output feature dimension of GCN, and D is the output feature dimension of GAT.
[0189] 2.2 Temporal feature extraction
[0190] The temporal feature extraction module focuses on capturing the dependencies and dynamic behavior changes in the node time series. The query matrix Q, key matrix K and value matrix V are calculated through temporal convolution, combined with the enhanced adjacency matrix Enables effective modeling of time series characteristics.
[0191] First, calculate the attention weights:
[0192]
[0193] in, The query, key, and value matrices are calculated through temporal convolution, d head is the single-head attention dimension.
[0194] Then perform softmax normalization on the fusion result to obtain the weighted attention weight:
[0195] P=softmax(S),P∈R B×N×T×T (17)
[0196] Finally, the output features are calculated by combining the value matrix V:
[0197]
[0198] The temporal features are further processed by a feedforward network (FNN) and batch normalization (BN) to extract nonlinear features and stabilize the training process:
[0199] H (temporal,t) =BN(FFN(H (temporal,t) )) (19)
[0200] 2.3 Feature Fusion
[0201] Spatiotemporal features are dynamically weighted and fused through a gating mechanism, thereby dynamically adjusting the contribution ratio of spatial and temporal features according to the characteristics of the industrial scenario.
[0202] First, the gate weight g is generated through two layers of fully connected layer networks and activation functions spatial and g temporal :
[0203] g spatial =σ(W s H spatial,t +b s ),g temporal =σ(W t H' temporal,t +b t ) (20)
[0204] Among them, W s 、W t and b s 、b t It is a learnable parameter trained based on industrial data. σ represents the sigmoid activation function, which ensures that the weight value is between [0, 1].
[0205] Weight the features according to the gating weights:
[0206] H (fusion,t) =g spatial ⊙H (spatial,t) +g temporal ⊙H (temporal,t) (twenty one)
[0207] Here, ⊙ represents element-wise multiplication.
[0208] The fused features are further subjected to nonlinear transformation and normalization to obtain the final fused features:
[0209] H (final,t) =BN(ReLU(FC(H (fusion,t) ))) (twenty two)
[0210] Here, FC is the fully connected layer, BN represents batch normalization, and ReLU is the activation function. The final graph-level feature H (final,t) Used for system-wide status monitoring and anomaly detection.
[0211] 3. Comprehensive judgment of the state discrimination layer
[0212] After completing the spatiotemporal feature extraction, the graph-level feature H finalThe input is sent to the state discrimination layer. This discrimination layer consists of multiple node-level binary classifiers, each of which independently determines the state of the corresponding node (abnormal or normal). The classification results of each node are combined through a voting mechanism to detect anomalies in the overall state of the industrial control system.
[0213] For each node, the state discrimination layer contains an independent binary classifier. The input is the feature vector corresponding to the node, and the output is the probability value of whether the node is abnormal. The output formula of the classifier is as follows:
[0214]
[0215] Among them, h i represents the feature vector of node i, W i and b i is the classifier parameter, and σ is the Sigmoid activation function, which is used to map the classification results to the interval [0,1].
[0216] The voting mechanism is used to aggregate the output of all node classifiers to determine the overall state of the graph. First, the classification output of all nodes is Perform binarization to determine whether the node anomaly probability threshold τ is exceeded. Then, calculate the proportion of abnormal nodes and compare it with the full graph anomaly determination threshold δ:
[0217]
[0218] Wherein, I(·) represents an indicator function, and when the condition is true, I(·) = 1, otherwise I(·) = 0.
[0219] 4. Classification Performance Evaluation and Optimization
[0220] In the performance evaluation and optimization steps, three key performance indicators, precision, recall, and F1-Score, are used to evaluate the classifier's effectiveness. The calculation formula is as follows:
[0221]
[0222] TP, FP, and FN represent the positive samples predicted by the model as positive, the negative samples predicted by the model as negative, and the positive samples predicted by the model as negative, respectively. Precision measures the proportion of samples predicted as abnormal that are actually abnormal; recall measures the proportion of abnormal samples that are correctly identified; and the F1 score, the harmonic average of precision and recall, provides a comprehensive assessment of the classifier's overall performance. During the optimization process, the system adjusts classifier parameters, optimizes threshold settings, and vote weight distribution, and combines cross-validation to evaluate the model's performance on different datasets to improve its generalization and robustness.
[0223] Different nodes have different roles and importance in the task, so each node requires a different feature extraction method. For each node, a customized feature extraction submodule is created to obtain more accurate graph-level features that reflect the specific operational status of the node.
[0224] The contents not described in detail in this specification belong to the prior art known to those skilled in the art.
[0225] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A method for industrial equipment fault detection that integrates complex relationships and spatiotemporal dependencies, characterized in that: The following steps are involved: S1. Obtain multidimensional time series data for equipment and construct a distance adjacency matrix, a topological order matrix, and a node correlation matrix based on a Gaussian kernel to capture the influencing factors of industrial equipment node relationships. This matrix is then weighted and fused to form an enhanced adjacency matrix that describes the multidimensional relationships between equipment. S2. Construct a spatiotemporal feature extraction module, input the multidimensional time series data of the device and the enhanced adjacency matrix described in S1, extract the temporal and spatial features of the device nodes, and use a gating mechanism to dynamically adjust the contribution ratio of temporal and spatial features, dynamically fusing temporal and spatial features to generate high-dimensional graph-level features; S3. Each device node is equipped with a binary classifier based on the Sigmoid activation function. Multiple node-level binary classifiers form a state discrimination layer. The high-dimensional graph-level features generated in S2 are input into this state discrimination layer, which independently determines whether the corresponding device node's state is abnormal or normal. A voting mechanism is used to combine the classification results of each device node to form an industrial equipment fault detection model that integrates complex relationships and spatiotemporal dependencies. This model then determines whether the overall state of the industrial control system is abnormal. S4. Use precision, recall, and F1 score performance metrics to evaluate and optimize the detection capability of the model described in S3 by adjusting the binary classifier parameters and voting weight distribution, and combining cross-validation. The precision measures the proportion of samples predicted to be abnormal that are actually abnormal; the recall measures the proportion of samples that are correctly identified as abnormal; and the F1 score is the harmonic average of the precision and recall.
2. The industrial equipment fault detection method integrating complex relationships and spatiotemporal dependencies according to claim 1 is characterized in that: In S1: The characteristics of the factors affecting the relationship between industrial equipment nodes are: The spatial correlation between devices is captured by the distance adjacency matrix of the Gaussian kernel; Capture the direct interaction relationship between devices through the topological order matrix; The node correlation matrix is used to capture the time series similarity, anomaly similarity, and statistical feature similarity of data between devices.
3. The industrial equipment fault detection method integrating complex relationships and spatiotemporal dependencies according to claim 1 is characterized in that: The S1 is specifically: S1.1 Constructing a distance adjacency matrix based on Gaussian kernel The Gaussian kernel function is used to convert the physical distance between devices into the similarity between nodes, capturing the spatial correlation between devices. The specific definition is as follows: Among them, h i and h j is the spatial location feature of node i and node j; σ is the bandwidth parameter of the Gaussian kernel function, which controls the sensitivity of similarity to distance changes; α ij is the similarity between node i and node j; S1.2 Constructing a topological order matrix The topological order matrix is used to model a clear physical or logical connection structure, reflecting the direct interaction relationship between devices. It is defined as follows: S1.3 Constructing node correlation matrix The node correlation matrix describes the similarity between nodes by analyzing the time series, abnormal data and statistical characteristics of the nodes; the node correlation matrix is subdivided into a time series similarity matrix, an abnormality similarity matrix, and a statistical feature similarity matrix; S1.4 Weighted fusion of the Gaussian kernel-based distance adjacency matrix, topological order matrix and node correlation matrix to form an enhanced adjacency matrix.
4. The industrial equipment fault detection method integrating complex relationships and spatiotemporal dependencies according to claim 3 is characterized in that: The node correlation matrix constructed in S1.3 is specifically as follows: S1.3.1 Constructing a time series similarity matrix Dynamic Time Warping measures the similarity of the changing patterns of time series generated by industrial equipment by aligning them. It is defined as follows: D ij =DTW(X i ,X j ) (3) Among them, D ij is the time series similarity; X i and X j Represent the time series data of node i and node j respectively. DTW() represents the dynamic time warping algorithm, which calculates the alignment distance between sequences. The value of the DTW matrix is between 0 and 1. The smaller the value, the more similar the time series of the two nodes are, and the larger the value, the lower the similarity of the time series of the two nodes. S1.3.2 Constructing anomaly similarity matrix By analyzing historical anomaly data, quantifying the anomaly similarity between nodes, and capturing the law of anomaly propagation, the definition is as follows: Among them, A ij AnomalyCount is the abnormal similarity; ij AnomalyCount represents the number of common anomalies that occur between nodes i and j within a certain time window. i and AnomalyCount j are the total number of abnormalities of node i and node j respectively; A ij ∈[0,1], the higher the value, the higher the similarity between node i and node j in abnormal behavior; S1.3.3 Constructing statistical feature similarity matrix The statistical characteristics of a device can reflect its operating mode. The statistical characteristics similarity matrix is used to quantify the similarity of the statistical characteristics of the devices. It is defined as follows: Among them, S ij is the statistical feature similarity; and Represent the values of node i and node j on the kth statistical feature respectively, and M is the total number of statistical features. The statistical feature similarity matrix measures the differences between nodes on multiple statistical features. The smaller the difference, the higher the similarity. The definition of weighted fusion of multiple adjacency matrices in S1.4 is as follows: in, Represents the enhanced adjacency matrix elements between node i and node j; each matrix (α ij 、T ij 、D ij 、A ij 、S ij ) are used to control the contribution of each matrix to the final adjacency matrix; w1 reflects the importance of spatial relationships, w2 reflects the influence of topological connections, w3 reflects the similarity weight of time series patterns, w4 reflects the importance of anomaly similarity, and w5 reflects the weight of statistical feature similarity.
5. The industrial equipment fault detection method integrating complex relationships and spatiotemporal dependencies according to claim 4 is characterized in that: In S2: Spatial feature extraction uses graph convolutional neural networks and random graph attention networks to extract spatial features of different dimensions in parallel, capturing local and global spatial relationships between devices in parallel; Temporal feature extraction captures dependencies and dynamic behaviors in time series through temporal convolution and multi-head attention mechanisms; Temporal features and spatial features are fused dynamically and weighted through a gating mechanism to generate high-dimensional graph-level features.
6. The industrial equipment fault detection method integrating complex relationships and spatiotemporal dependencies according to claim 5 is characterized in that: The S2 is specifically: S2.1 Spatial feature extraction S2.1.1 Extracting spatial features using graph convolutional neural networks GCN is used to capture local fixed spatial dependencies in industrial information graphs, adjacency matrix The normalized form of balances the feature propagation between nodes to avoid numerical explosion or disappearance, as shown in the following formula: in, is the node feature matrix of the l+1th layer; is the normalized augmented adjacency matrix, D t is the degree matrix, is the node feature matrix of the lth layer, W (l) is the learnable weight matrix; The final output of GCN is: Among them, H (GCN,t) is the final output feature matrix of GCN; is the dimension space of the output feature; L represents the number of layers of GCN, and the output feature dimension is d out ; S2.1.2 Spatial Feature Extraction by Random Graph Attention Network The random graph attention network is introduced to capture dynamic non-uniform spatial dependencies. Random GAT uses a randomly generated attention matrix Combined with the graph structure mask function mask(·), the relationship weights between nodes are dynamically adjusted; the attention weights are generated as follows: in, is the inter-node weight matrix randomly generated for the i-th attention head at time step t; the mask function mask(·) is used to retain the edges existing in the industrial graph structure and filter out invalid edges; is the normalized dynamic weight matrix between nodes; Calculate the learnable attention coefficient between nodes based on the random attention matrix To capture node v i and v j The dependency of is normalized with softmax to obtain the attention weight matrix between nodes: in, is node v i Neighborhood sets in industrial graphs; Using the attention weight matrix between nodes Update node characteristics: Where σ(·) is the activation function and W is the learnable weight matrix; For node v j Initial feature representation of ; For node v i Updated features; Introducing a multi-head mechanism to enhance learning stability and feature expression capabilities: Where K is the number of attention heads, Indicates concatenating the results of K heads, W k is the linear transformation weight matrix of the kth attention head; The hidden state update formula of random GAT is: in, is the hidden state matrix of the i-th attention head at time step t; H t is the input feature matrix at time step t, and W is the learnable weight matrix; After introducing the multi-head mechanism, the final characteristic matrix of random GAT is: Among them, H (R,t) is the final feature matrix of random GAT at time step t; is the weight matrix of the kth attention head; S2.1.3 Spatial Feature Fusion The features extracted by random GAT are concatenated with the features extracted by GCN in the feature dimension to form the final spatial feature representation: H (spatial,t) =concat(H (GCN,t) ,H (R,t) ,dim=-1) (15) The feature dimension after splicing is where d out is the output feature dimension of GCN, and D is the output feature dimension of GAT; S2.2 Temporal feature extraction The temporal feature extraction module is used to capture the dependencies and dynamic behavior changes in the node time series; the query matrix Q, key matrix K and value matrix V are calculated through temporal convolution, combined with the enhanced adjacency matrix Achieve effective modeling of time series characteristics; First, calculate the attention weights: in, The query, key, and value matrices are calculated through temporal convolution, d head It is the single-head attention dimension; Then perform softmax normalization on the fusion result to obtain the weighted attention weight: P=softmax(S),P∈R B×N×T×T (17) Finally, the output features are calculated by combining the value matrix V: The temporal features are further processed by the feedforward network FNN and batch normalization BN to extract nonlinear features and stabilize the training process: H (temporal,t) =BN(FFN(H (temporal,t) )) (19) S2.3 Fusion of temporal and spatial features Spatiotemporal features are dynamically weighted and fused through a gating mechanism, thereby dynamically adjusting the contribution ratio of spatial and temporal features according to the characteristics of the industrial scenario; First, the gate weight g is generated through two layers of fully connected layer networks and activation functions spatial and g temporal : g spatial =σ(W s H spatial,t +b s ),g temporal =σ(W t H′ temporal,t +b t ) (20) Among them, W s 、W t and b s 、b t It is a learnable parameter trained based on industrial data. σ represents the sigmoid activation function, which ensures that the weight value is between [0, 1]. Weight the features according to the gating weights: H (fusion,t) =g spatial ⊙H (spatial,t) +g temporal ⊙H (temporal,t) (21) Among them, ⊙ represents element-wise multiplication; The fused features are further subjected to nonlinear transformation and normalization to obtain the final fused features: H (final,t) =BN(ReLU(FC(H (fusion,t) ))) (22) FC is the fully connected layer, BN represents batch normalization, and ReLU is the activation function; the final graph-level feature H (final,t) Used for system-wide status monitoring and anomaly detection.
7. The industrial equipment fault detection method integrating complex relationships and spatiotemporal dependencies according to claim 6 is characterized in that: In the S3: The input of the state discrimination layer is the feature vector corresponding to the node, and the output is the probability value of whether the node is abnormal. The output formula of the classifier is as follows: in, is the classification output of the node; h i represents the feature vector of node i, W i and b i is the classifier parameter, σ is the Sigmoid activation function, which is used to map the classification results to the interval [0,1]; The voting mechanism is used to aggregate the output results of all node classifiers to determine the overall state of the graph; First, the classification output of all nodes Perform binarization to determine whether the node anomaly probability threshold τ is exceeded. Then, calculate the proportion of abnormal nodes and compare it with the overall graph anomaly determination threshold δ: Wherein, I(·) represents an indicator function, and when the condition is true, I(·) = 1, otherwise I(·) = 0.
8. The industrial equipment fault detection method integrating complex relationships and spatiotemporal dependencies according to claim 7 is characterized in that: The S4 is specifically: The performance indicators precision, recall and F1 score are used for classification evaluation and optimization. The calculation formula is as follows: Among them, TP, FP, and FN represent the positive samples predicted by the model as positive, the negative samples predicted by the model as negative, and the positive samples predicted by the model as negative, respectively; Precision is the precision rate; Recall is the recall rate; and F1_Score is the F1 score.
Citation Information
Cited By
Multi-head attention model conversion method and device, storage medium and electronic equipment
CN121009924A
Methods, devices, storage media, and electronic devices for converting multi-head attention models
CN121009924B
Abnormality detection method and electronic equipment
CN121233438A
Industrial production fault intelligent diagnosis method and system based on multi-modal data
CN121365206A
New energy station data anomaly detection method and system based on cloud edge collaborative architecture
CN121412797A