Fault detection method of multi-dimensional data reconstruction mode based on transformer

By combining the space-time joint feature reconstruction method of Transformer and TCN, the problem that traditional fault detection methods are difficult to capture joint abnormal signals in multi-testing point and multi-dimensional data processing is solved, achieving high-precision fault diagnosis and prediction, and improving the accuracy and robustness of detection.

CN120030312AInactive Publication Date: 2025-05-23CHINA YANGTZE POWER

Patent Information

Application Number
CN202510497961.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional fault detection methods are difficult to capture the combined abnormal signals between multiple measurement points and multi-dimensional industrial monitoring data, and cannot comprehensively model the global state of complex industrial equipment. They also model independently in the time and space dimensions, ignoring the dynamic correlation between measurement points and space-time coupling characteristics, affecting the accuracy of abnormal signal recognition.

Method used

A fault detection method based on Transformer's multidimensional data reconstruction method is proposed, combined with causal convolutional neural network (TCN) and Transformer structure, a joint spatial feature reconstruction method is proposed. Efficiently extract temporal and spatial correlation features under a single architecture, and accurately capture abnormal signals in multi-test point data by capturing time characteristics through TCN capture and processing spatiotemporal relationships through Transformer.

Benefits of technology

It realizes high-precision fault diagnosis in multi-dimensional data scenarios, improves the accuracy of fault prediction and detection, significantly reduces false detection and missed detection rates, and can effectively identify coordinated abnormal signals between multiple measurement points in a complex multi-dimensional data environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030312A_ABST
    Figure CN120030312A_ABST
Patent Text Reader

Abstract

The invention discloses a fault detection method based on a transformer multi-dimensional data reconstruction mode. The fault detection method comprises the following steps: S1, preprocessing and embedding input data; s2, time sequence feature extraction: in order to effectively extract time dependence features in time sequence data, a TCN module is used; s3, extracting spatial features, and inputting the data processed by the TCN block into a Transformer encoder; s4, decoding and target sequence generation: in the process of generating reconstructed data, the Transform decoder takes the output of the encoder as the input, and further learning and reconstruction are carried out in combination with the target sequence; s5, reconstruction data is generated and output, the decoded data is mapped through a full connection layer, and an original input data structure is recovered; and S6, model training and optimization. According to the method, the causal convolutional neural network and the Transform structure are combined, the space-time joint feature reconstruction method is provided, time features and space correlation features can be efficiently extracted under a single framework, abnormal signals in multi-measurement-point data can be accurately captured, and the limitation of a traditional method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of industrial big data computing, and in particular to a fault detection method based on a transformer-based multi-dimensional data reconstruction method. Background Art

[0002] In the field of industrial equipment status detection and fault diagnosis, with the rapid development of sensor technology and data acquisition systems, industrial equipment can monitor multi-dimensional data (such as temperature, pressure, flow, vibration, etc.) of a large number of measurement points in real time. Traditional single-measurement-point or single-data-stream analysis methods have limitations and are difficult to capture coordinated abnormal signals between multiple measurement points in complex equipment systems. For example, statistical analysis methods or threshold rule methods can only detect single-point anomalies and cannot effectively identify potential failure modes caused by joint changes in multiple measurement points. This limitation makes it difficult for traditional methods to achieve high-precision fault diagnosis in multi-dimensional data scenarios, and they are easily affected by false detections and missed detections.

[0003] Existing fault detection methods have many technical defects when processing multi-point and multi-dimensional industrial monitoring data: traditional statistical analysis, threshold rules or single time series models are difficult to capture joint abnormal signals between multiple measuring points, resulting in the inability to comprehensively model the global state of complex industrial equipment; at the same time, these methods independently model in the time and space dimensions, ignoring the dynamic correlation and spatiotemporal coupling characteristics between measuring points, which seriously affects the accuracy of abnormal signal identification.

[0004] In addition, the current deployment of industrial systems usually requires monitoring a large number of measurement points on the same device. In order to improve computing efficiency and deployment convenience, an efficient method is urgently needed to complete multi-measurement point data analysis, spatiotemporal joint feature extraction, and abnormal signal identification under a unified architecture. However, methods based on traditional time series models (such as RNN or LSTM) often face problems such as long training time and large cumulative errors, and cannot meet the needs of real-time monitoring in industrial environments.

[0005] Therefore, developing an advanced method that takes into account joint detection of multi-measurement point data, efficient deployment and real-time anomaly capture has important engineering value and practical application significance. Summary of the invention

[0006] The technical problem to be solved by the present invention is to provide a fault detection method based on transformer-based multi-dimensional data reconstruction. By combining the causal convolutional neural network (TCN) and the Transformer structure, a spatiotemporal joint feature reconstruction method is proposed, which can efficiently extract temporal features and spatial correlation features under a single architecture, accurately capture abnormal signals in multi-measurement point data, and solve the limitations of traditional methods.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is: a fault detection method based on a transformer-based multidimensional data reconstruction method, comprising the following steps: S1, input data preprocessing and embedding; S2, time series feature extraction, in order to effectively extract the time-dependent features in the time series data, the TCN module is used; S3, spatial feature extraction, the data processed by the TCN block will be input into the Transformer encoder; S4, decoding and target sequence generation. In the process of generating reconstructed data, the Transformer decoder takes the encoder output as input and combines it with the target sequence for further learning and reconstruction; S5, reconstructed data generation and output, the decoded data is mapped through the fully connected layer to restore the original input data structure; S6. Model training and optimization.

[0008] Preferably, in step S1, the input data is time series data. , where B is the batch size, T is the time step, and C is the number of nodes.

[0009] Preferably, in order to map the input data into the hidden space, a fully connected layer is used for embedding: ; in is the mapping matrix, , is the result of mapping to the hidden space, is the dimension of the latent space, is the offset.

[0010] Preferably, in step S2, each TCN module includes two convolutional layers, using a ReLU activation function and Dropout regularization.

[0011] Preferably, the TCN block captures the temporal characteristics of the data through local convolution operations: ; Each TCN block operates with a custom convolution kernel size and dilation factor to increase the receptive field and capture long-term dependencies.

[0012] Preferably, in step S3, the Transformer encoder extracts spatiotemporal features through a multi-layer self-attention mechanism and a feedforward neural network, and outputs an encoded hidden layer representation.

[0013] Preferably, the encoder formula is as follows: ; in, It is the encoded spatiotemporal feature representation.

[0014] Preferably, in step S4, the decoder generates the corresponding output through a self-attention mechanism: ; in, is the decoded spatiotemporal feature representation.

[0015] Preferably, step S5 is expressed as: ; in, is the reconstructed data, which is close to the original input data X. The reconstruction error is used for fault detection. The reconstruction error of the healthy state is small, while the reconstruction error of the fault state is large. is the weight matrix of the fully connected layer, used for dimension conversion, is the bias term of the fully connected layer.

[0016] Preferably, in step S6, during the training process, the model uses mean square error as the loss function of the reconstruction error, and the optimizer uses the Adam optimization algorithm.

[0017] The present invention provides a fault detection method based on transformer-based multidimensional data reconstruction, which has the following beneficial effects: 1. Combining the multi-dimensional spatiotemporal feature extraction capabilities of Transformer and TCN: The present invention innovatively combines the powerful time series modeling capabilities of Transformer with the causal convolution characteristics of TCN, making full use of the advantages of Transformer in processing long time series dependencies and the ability of TCN to efficiently capture local time features. This combined method can handle long-term dependencies and short-term dependencies at the same time, improving the accuracy and robustness of time series data modeling. Compared with the methods in the prior art that only rely on a single model (such as using only LSTM or Transformer), the multi-model joint structure of the present invention can not only more accurately capture the spatiotemporal relationships at different time scales, but also achieve higher fault prediction and detection accuracy in more complex multidimensional data environments. This cross-model synergy has unique advantages in processing multi-point data of industrial equipment, and can improve the overall detection accuracy while ensuring high efficiency.

[0018] 2. TCN block’s local collaborative detection capability for multi-dimensional data: The present invention designs multiple TCN blocks to process different time scales of input data, adds a dilation factor in the causal convolution operation, and effectively improves the receptive field of the model, enabling it to simultaneously process local features and global dependencies in time series. Unlike existing traditional convolutional neural networks or long short-term memory networks (LSTM), the TCN module does not rely on recursive or fixed-step calculation methods, but instead performs parallel processing through causal convolution, which greatly improves computational efficiency. In multidimensional data, the local feature extraction capability of TCN can effectively capture the collaborative information between each measuring point, and is particularly suitable for application in industrial equipment fault detection tasks where there are complex spatiotemporal correlations between multiple measuring points. This innovation can effectively simplify the redundant calculations of traditional models in time series data processing, thereby improving the efficiency of model training and reasoning.

[0019] 3. Multi-dimensional data local collaborative detection and joint fault identification capabilities: The present invention utilizes the advantages of Transformer and TCN, and proposes an innovative method that can perform local collaborative detection and joint fault identification on multi-dimensional data. In the collaborative analysis process of multi-measurement point data, Transformer can capture the long-term spatiotemporal dependencies between measurement points through the self-attention mechanism, while TCN is responsible for extracting the time characteristics of each measurement point. Through this local collaborative detection mechanism, the present invention can improve the accuracy of fault detection when processing large-scale data, especially when there are multiple potential faults in the equipment at the same time, it can better distinguish and predict the fault type. The prior art usually treats the fault detection of each measurement point as an independent task, and cannot effectively consider the synergy between measurement points. In contrast, the solution of the present invention can utilize the joint information between the measurement points, significantly improve the joint diagnosis capability of multi-point faults, and effectively reduce the fault missed detection rate. This innovation provides a stronger fault diagnosis capability for the intelligent operation and maintenance and fault warning system of industrial equipment, and has great practical application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The present invention will be further described below in conjunction with the accompanying drawings and embodiments: Figure 1 is a flow chart of the method of the present invention; Figure 2 A comparison diagram of the reconstructed data and the real data during the boot process of the present invention; Figure 3 It is a graph showing the percentage error between the reconstructed data and the real data during the booting process of the present invention; Figure 4 It is a comparison diagram of the reconstructed data and the real data of the present invention under stable working conditions; Figure 5 This is a graph showing the percentage errors between the reconstructed data and the real data under stable working conditions of the present invention. DETAILED DESCRIPTION

[0021] like Figure 1 As shown, the fault detection method based on the transformer-based multi-dimensional data reconstruction method includes the following steps: S1, input data preprocessing and embedding; S2. Time series feature extraction. In order to effectively extract the time-dependent features in time series data, the TCN (causal convolution) module is used; S3, spatial feature extraction, the data processed by the TCN block will be input into the Transformer encoder; S4, decoding and target sequence generation. In the process of generating reconstructed data, the Transformer decoder takes the encoder output as input and combines it with the target sequence for further learning and reconstruction; S5, reconstructed data generation and output, the decoded data is mapped through the fully connected layer to restore the original input data structure; S6. Model training and optimization.

[0022] Preferably, in step S1, the input data is time series data. , where B is the batch size, T is the time step, and C is the number of nodes (feature channels). Clarifying the format of input time series data makes data processing and model building clearer and more standardized, helps to accurately understand and process input data, and provides a clear definition of data dimensions for subsequent data embedding, feature extraction and other operations.

[0023] Preferably, in order to map the input data into the hidden space, a fully connected layer is used for embedding: ; in is the mapping matrix, , is the result of mapping to the hidden space. The fully connected layer is used for data embedding, and the input data is mapped to the hidden space, which provides a suitable data representation for subsequent feature extraction and model training. This method is simple, direct and universal.

[0024] Preferably, in step S2, each TCN module includes two convolutional layers, using ReLU activation function and Dropout regularization; the ReLU activation function can increase the nonlinear expression ability of the model, so that the model can learn more complex features; Dropout regularization helps prevent overfitting of the model and improve the generalization ability of the model.

[0025] Preferably, the TCN block captures the temporal characteristics of the data through local convolution operations: ; Each TCN block operates through a custom convolution kernel size and dilation factor to increase the receptive field and capture long-term dependencies. Through local convolution operations and custom convolution kernel sizes and dilation factors, the TCN block can effectively capture the temporal characteristics of the data, expand the receptive field, better handle long-term and short-term dependencies in time series, and improve the ability to analyze time series data.

[0026] Preferably, in step S3, the Transformer encoder extracts spatiotemporal features through a multi-layer self-attention mechanism and a feedforward neural network, and outputs an encoded hidden layer representation. The Transformer encoder uses a multi-layer self-attention mechanism and a feedforward neural network to extract spatiotemporal features. The self-attention mechanism can automatically learn the association between data at different positions and effectively capture the spatial features and long-term temporal dependencies in the data; the feedforward neural network further processes and enhances the features to improve the expressiveness of the features.

[0027] Preferably, the encoder formula is as follows: ; in, It is the encoded spatiotemporal feature representation.

[0028] Preferably, in step S4, the decoder generates the corresponding output through a self-attention mechanism: ; in, is the decoded spatiotemporal feature representation.

[0029] Preferably, step S5 is expressed as: ; in, is the reconstructed data, which is close to the original input data X. The reconstruction error is used for fault detection. The reconstruction error of the healthy state is small, while the reconstruction error of the faulty state is large.

[0030] Preferably, in step S6, during the training process, the model uses mean square error (MSE) as the loss function of the reconstruction error, and the optimizer uses the Adam optimization algorithm. Using mean square error (MSE) as the loss function can intuitively measure the difference between the reconstructed data and the original data, which is convenient for model training and optimization; the Adam optimization algorithm has the characteristics of adaptive learning rate, can converge quickly during the training process, and improve the training efficiency of the model.

[0031] like Figure 2The figure shows the comparison of the reconstructed data and the real data during the startup of the 16-tile temperature of Unit 4 of a large power plant. The blue line in the figure is the real data and the red line is the reconstructed data. The horizontal axis represents the time step and the vertical axis represents the normalized tile temperature, which is in the range of [0,1]. It can be seen that the reconstructed data is very close to the real data.

[0032] like Figure 3 The figure shows the error percentage between the reconstructed data and the real data during the startup process of the 16-watt unit 4 of a large power plant. The horizontal axis represents the number of time steps and the vertical axis represents the percentage. It can be seen from the figure that even during the startup process, the maximum error percentage of the reconstructed data is only about 10%.

[0033] like Figure 4 The figure shows the comparison between the reconstructed data and the real data of the stable temperature condition of 16 tiles in Unit 4 of a large power plant. The blue line in the figure is the real data and the red line is the reconstructed data. The horizontal axis represents the time step and the vertical axis represents the normalized tile temperature, which is in the range of [0,1]. It can be seen that the reconstructed data is very close to the real data.

[0034] like Figure 5 The figure shows the error percentage between the reconstructed data and the real data of the 16-watt unit 4 of a large power plant under stable temperature conditions. The horizontal axis represents the number of time steps and the vertical axis represents the percentage. It can be seen from the figure that even under stable conditions, the error percentage of the reconstructed data is less than 5%.

[0035] The present invention proposes a spatiotemporal joint feature extraction and reconstruction method based on the combination of causal convolutional neural network (TCN) and Transformer, which is suitable for fault detection, prediction and reconstruction of multi-dimensional time series data. This method can effectively capture the spatiotemporal dependency in time series data by combining TCN to extract time features and Transformer to process spatiotemporal relationships, and is particularly suitable for industrial data with multiple measurement points and long-term dependencies. The present invention proposes a spatiotemporal joint feature reconstruction method by combining causal convolutional neural network (TCN) and Transformer structure, which can efficiently extract time features and spatial correlation features under a single architecture, accurately capture abnormal signals in multi-measurement point data, and solve the limitations of traditional methods.

[0036] The above embodiments are only preferred technical solutions of the present invention and should not be regarded as limiting the present invention. The protection scope of the present invention shall be the technical solutions recorded in the claims, including equivalent replacement solutions of the technical features in the technical solutions recorded in the claims. That is, equivalent replacement improvements within this scope are also within the protection scope of the present invention.

Claims

1. A fault detection method based on transformer-based multidimensional data reconstruction, characterized in that: The following steps are involved: S1, input data preprocessing and embedding; S2, time series feature extraction, in order to effectively extract the time-dependent features in the time series data, the TCN module is used; S3, spatial feature extraction, the data processed by the TCN block will be input into the Transformer encoder; S4, decoding and target sequence generation. In the process of generating reconstructed data, the Transformer decoder takes the encoder output as input and combines it with the target sequence for further learning and reconstruction; S5, reconstructed data generation and output, the decoded data is mapped through the fully connected layer to restore the original input data structure; S6. Model training and optimization.

2. The fault detection method based on transformer-based multidimensional data reconstruction method according to claim 1, characterized in that: In step S1, the input data is time series data. , where B is the batch size, T is the time step, and C is the number of nodes.

3. The fault detection method based on transformer-based multidimensional data reconstruction method according to claim 2, characterized in that: In order to map the input data to the hidden space, a fully connected layer is used for embedding: ; in is the mapping matrix, , is the result of mapping to the hidden space, is the dimension of the latent space, is the offset.

4. The fault detection method based on transformer-based multidimensional data reconstruction method according to claim 1, characterized in that: In step S2, each TCN module includes two convolutional layers, using ReLU activation function and Dropout regularization.

5. The fault detection method based on transformer-based multidimensional data reconstruction method according to claim 1, characterized in that: The TCN block captures the temporal characteristics of the data through local convolution operations: ; Each TCN block operates with a custom convolution kernel size and dilation factor to increase the receptive field and capture long-term dependencies.

6. The fault detection method based on transformer-based multidimensional data reconstruction method according to claim 1, characterized in that: In step S3, the Transformer encoder extracts spatiotemporal features through a multi-layer self-attention mechanism and a feedforward neural network, and outputs an encoded hidden layer representation.

7. The fault detection method based on transformer-based multidimensional data reconstruction method according to claim 6, characterized in that: The encoder formula is as follows: ; in, It is the encoded spatiotemporal feature representation.

8. The fault detection method based on transformer-based multidimensional data reconstruction method according to claim 1, characterized in that: In step S4, the decoder generates the corresponding output through the self-attention mechanism: ; in, is the decoded spatiotemporal feature representation.

9. The fault detection method based on transformer-based multidimensional data reconstruction method according to claim 1, characterized in that: The step S5 is expressed as: ; in, is the reconstructed data, which is close to the original input data X. The reconstruction error is used for fault detection. The reconstruction error of the healthy state is small, while the reconstruction error of the fault state is large. is the weight matrix of the fully connected layer, used for dimension conversion, is the bias term of the fully connected layer.

10. The fault detection method based on transformer-based multidimensional data reconstruction method according to claim 1, characterized in that: In the training process of step S6, the model uses mean square error as the loss function of reconstruction error, and the optimizer uses the Adam optimization algorithm.

Citation Information

Patent Citations

  • Software system multi-index anomaly detection method and system

    CN117724935A

Cited By

  • A bridge health monitoring abnormal data reconstruction method, system, equipment and product

    CN120524402B

  • Bridge data reconstruction method based on multi-scale space-time fusion and uncertainty perception

    CN121959459A

  • Bridge data reconstruction method based on multi-scale spatio-temporal fusion and uncertainty perception

    CN121959459B