Power generation equipment fault diagnosis method and system based on multi-modal data fusion

By constructing a multimodal data stream and utilizing a joint dual-channel network and a multi-head self-attention mechanism for feature analysis and fusion, the problem of difficult integration of multimodal data is solved, and the accuracy of fault diagnosis of power generation equipment is improved.

CN122020353APending Publication Date: 2026-05-12LIAONING DATANG INT NEW ENERGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
LIAONING DATANG INT NEW ENERGY CO LTD
Filing Date
2025-12-15
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing monitoring systems cannot effectively integrate multimodal data, resulting in low accuracy in diagnosing power generation equipment faults.

Method used

By dynamically collecting data on the operation of power generation equipment, a multimodal data stream is constructed. Feature analysis and fusion are performed using a joint dual-channel network and a multi-head self-attention mechanism to generate a multimodal fusion feature vector. Finally, a fully connected classification is performed to identify faults.

Benefits of technology

It improved the accuracy of fault diagnosis for power generation equipment and achieved deep fusion and effective integration of multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122020353A_ABST
    Figure CN122020353A_ABST
Patent Text Reader

Abstract

The invention discloses a power generation equipment fault diagnosis method and system based on multi-modal data fusion, and relates to the technical field of equipment diagnosis. The method comprises the following steps: carrying out operation dynamic acquisition on power generation equipment to obtain multi-modal heterogeneous data, and carrying out space-time alignment on the multi-modal heterogeneous data to construct a multi-modal data stream; constructing a joint two-channel network, and performing feature analysis on the multi-modal data flow through the joint two-channel network to generate two-dimensional feature parameters; activating a multi-head self-attention mechanism to fuse the two-dimensional feature parameters, calculating a multi-modal feature association weight, and generating a multi-modal fusion feature vector; and backtracking to the power generation equipment based on the multi-modal fusion feature vector to carry out full connection classification, carrying out fault identification according to a classification result, and generating a fault diagnosis result. The technical problem that in the prior art, due to the fact that multi-modal data are difficult to integrate effectively, the fault diagnosis accuracy is low is solved, and the technical effect of improving the fault diagnosis accuracy is achieved through deep fusion of the multi-modal data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of equipment diagnostic technology, and specifically to a method and system for fault diagnosis of power generation equipment based on multimodal data fusion. Background Technology

[0002] During operation, power generation equipment generates various types of monitoring data, including electrical quantities, mechanical quantities, heat, environmental quantities, and image and acoustic data. These monitoring information collectively reflect the equipment's operating status and potential fault characteristics. However, existing monitoring systems typically collect data from different modalities independently. These modalities differ significantly in sampling frequency, data format, spatiotemporal resolution, and noise characteristics, making direct fusion analysis difficult. Traditional fault diagnosis methods mostly rely on feature extraction based on single-modal or simply pieced-together data, failing to effectively capture the correlations between multimodal data. This results in insufficient feature representation capabilities, making accurate identification of equipment faults under complex operating conditions challenging. Summary of the Invention

[0003] This application provides a method and system for fault diagnosis of power generation equipment based on multimodal data fusion, which solves the technical problem that multimodal data is difficult to integrate effectively in the prior art, resulting in low accuracy of fault diagnosis.

[0004] The first aspect of this application provides a method for fault diagnosis of power generation equipment based on multimodal data fusion, the method comprising: Dynamic data acquisition of power generation equipment is performed to obtain multimodal heterogeneous data. This multimodal heterogeneous data is then spatiotemporally aligned to construct a multimodal data stream. A joint dual-channel network is constructed, and feature analysis is performed on the multimodal data stream through this network to generate two-dimensional feature parameters. A multi-head self-attention mechanism is activated to fuse the two-dimensional feature parameters, calculate the multimodal feature association weights, and generate a multimodal fused feature vector. Based on the multimodal fused feature vector, a fully connected classification is performed on the power generation equipment. Fault identification is then performed based on the classification results, and a fault diagnosis result is generated.

[0005] A second aspect of this application provides a fault diagnosis system for power generation equipment based on multimodal data fusion, the system comprising: Data processing module: dynamically collects data on the operation of the power generation equipment to obtain multimodal heterogeneous data, aligns the multimodal heterogeneous data spatiotemporally, and constructs a multimodal data stream; Feature analysis module: constructs a joint dual-channel network, performs feature analysis on the multimodal data stream through the joint dual-channel network, and generates two-dimensional feature parameters; Feature fusion module: activates a multi-head self-attention mechanism to fuse the two-dimensional feature parameters, calculates the multimodal feature association weights, and generates a multimodal fused feature vector; Fault identification module: backtracks to the power generation equipment based on the multimodal fused feature vector to perform fully connected classification, identifies faults based on the classification results, and generates fault diagnosis results.

[0006] One or more technical solutions provided in this application have at least the following technical effects or advantages: First, dynamic data acquisition of the power generation equipment is performed to obtain multimodal heterogeneous data. This data is then spatiotemporally aligned to construct a multimodal data stream. Next, a joint dual-channel network is built to perform feature analysis on the multimodal data stream, generating two-dimensional feature parameters. Then, a multi-head self-attention mechanism is activated to fuse the two-dimensional feature parameters, calculate the multimodal feature association weights, and generate a multimodal fused feature vector. Finally, based on the multimodal fused feature vector, a fully connected classification is performed on the power generation equipment. Fault identification is then performed based on the classification results, generating fault diagnosis results. This approach solves the technical problem of low fault diagnosis accuracy caused by the difficulty in effectively integrating multimodal data in existing technologies. Through deep fusion of multimodal data, the technical effect of improving fault diagnosis accuracy is achieved. Attached Figure Description

[0007] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0008] Figure 1 A schematic diagram of the process for a power generation equipment fault diagnosis method based on multimodal data fusion provided in this application embodiment; Figure 2 This is a schematic diagram of the structure of a power generation equipment fault diagnosis system based on multimodal data fusion, provided in an embodiment of this application.

[0009] Figure labeling: Data processing module 11, feature analysis module 12, feature fusion module 13, fault identification module 14. Detailed Implementation

[0010] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.

[0011] Example 1, as Figure 1 As shown, this application provides a fault diagnosis method for power generation equipment based on multimodal data fusion, wherein the method includes: Dynamic data collection is performed on the power generation equipment to obtain multimodal heterogeneous data. The multimodal heterogeneous data is then spatiotemporally aligned to construct a multimodal data stream.

[0012] During the operation of the power generation equipment, vibration signals, thermal image sequences, acoustic spectrum data, and multi-dimensional operating condition parameters are continuously collected by vibration sensors, infrared thermal imaging devices, acoustic acquisition units, and operating condition monitoring modules deployed on the equipment body and its key components. During the acquisition process, a unified clock signal output by the system is used as the time alignment reference to synchronously map the original heterogeneous data with different sampling frequencies, formats, and sources to a unified time axis. The modal data is sliced ​​according to the preset time window length, and the sampling step size difference is eliminated by interpolation, downsampling, or upsampling. Furthermore, regional interest extraction is performed on the thermal image sequence, time offset compensation is performed on the acoustic data, and time-frequency transformation is performed on the vibration signal to ensure that all modes have an alignable temporal structure. Finally, the processed vibration modal data, thermal imaging modal data, acoustic modal data, and operating condition modal data are combined and recombined in chronological order to form a multimodal data stream with a unified timestamp and a unified data structure.

[0013] Furthermore, the method involves dynamically acquiring data from the power generation equipment to obtain multimodal heterogeneous data, and then spatiotemporally aligning this multimodal heterogeneous data to construct a multimodal data stream. A clock signal is introduced and used as a reference parameter. Based on this reference parameter, the multimodal heterogeneous data of the power generation equipment is decomposed according to multiple sampling frequencies to obtain a first signal set, a second signal set, and a third signal set. Vibration signals are extracted from the first signal set and subjected to short-time Fourier transform to generate a time-spectrum image, which serves as the first modal time-series data. Thermal image sequences are extracted from the second signal set, and regional interest extraction is performed on key components of the power generation equipment according to the thermal image sequences to obtain local thermal image sequences, which serve as the second modal time-series data. Acoustic spectrum data extracted from the third signal set is subjected to mutual information analysis with the first modal time-series data. Time offset compensation is performed on the acoustic spectrum data according to the mutual information parameters to obtain a fourth signal set. The first modal time-series data, the second modal time-series data, the third signal set, and the fourth signal set are reassembled along a time axis to generate the multimodal data stream.

[0014] The system introduces a unified clock signal, which serves as the reference parameter for multimodal data alignment. Based on this reference parameter, the raw heterogeneous data from vibration sensors, infrared thermal imaging devices, acoustic acquisition units, and operational condition monitoring modules are decomposed and sliced ​​according to their respective sampling frequencies to obtain a first signal set, a second signal set, and a third signal set. A short-time Fourier transform is performed on the vibration signals in the first signal set to convert the time-domain vibration data into a time-spectrum graph, which is then used as the first-mode time-series data capable of characterizing the evolution of vibration features. The thermal image sequence in the second signal set is processed by traversing the key components of the equipment according to a pre-defined component list. The system extracts regions of interest based on thermal imaging temperature distribution to obtain local thermal image sequences, which are then used as the second modal time-series data. Acoustic spectrum data is extracted from the third signal set and mutual information analysis is performed between it and the first modal time-series data. The time dependency between the acoustic signal and the vibration mode is calculated based on the mutual information parameters. Time offset compensation is applied to the acoustic spectrum data to obtain a fourth signal set with time alignment capability. Finally, the first modal time-series data, the second modal time-series data, the third signal set, and the fourth signal set are reassembled according to a unified time axis order to construct a structurally consistent and time-synchronized multimodal data stream.

[0015] A joint dual-channel network is constructed, and feature analysis is performed on the multimodal data stream through the joint dual-channel network to generate two-dimensional feature parameters.

[0016] Based on multimodal data streams, and considering the differences in spatial distribution and temporal variation of different modal features, a joint dual-channel network structure is constructed by introducing convolutional neural networks (CNNs) and long short-term memory (LSTM) networks. The CNNs are used to extract spatial features from vibration spectrograms and thermal imaging sequences, while the LSTM networks are used to learn the temporal dependencies between acoustic spectrum data and operating parameters. The corresponding first-modal temporal data, second-modal temporal data, third signal set, and fourth signal set from the multimodal data stream are input into the CNN and LSTM channels, respectively. Deep representation features of each modality are extracted through methods such as convolutional kernel sliding calculation, feature map mapping, and time-step iterative calculation of gated recurrent units. Based on the extracted feature vectors, spatial and temporal feature vectors are constructed according to the characteristics of the convolutional and temporal channels. The spatial and temporal feature vectors are then dimensionally divided and structurally organized to form a dual-dimensional feature parameter containing both spatial and temporal domain features, which is used for subsequent fusion calculations using a multi-head self-attention mechanism.

[0017] Furthermore, a joint dual-channel network is constructed, and feature analysis is performed on the multimodal data stream through the joint dual-channel network to generate two-dimensional feature parameters. The method includes: A joint dual-channel network is constructed using a convolutional neural network (CNN) and a long short-term memory (LSTM) network. The CNN includes a first CNN network and a second CNN network, and the LSTM network includes a first LSTM network and a second LSTM network. The time-spectrum image of the first modality time-series data is synchronized to the first CNN network to extract the deep spatial spectrum features of the vibration signal as the first feature vector. The second modality time-series data is synchronized to the second CNN network to extract the abnormal hotspot features of the temperature distribution of the infrared thermal image as the second feature vector. The acoustic spectrum data of the third signal set is synchronized to the first LSTM network for acoustic feature learning to obtain the short-time change parameters of the acoustic features as the third feature vector. Based on the fourth signal set, multi-dimensional operating condition parameters are extracted and synchronized to the second LSTM network for operating condition learning and analysis to obtain the long-period temporal dependence of the learned operating condition parameters as the fourth feature vector. The first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are dimensionally divided to construct the dual-dimensional feature parameters.

[0018] To address the differences in data characteristics among vibration modes, thermal imaging modes, acoustic modes, and operating condition modes in multimodal data streams, a joint dual-channel network is constructed using a convolutional neural network (CNN) and a long short-term memory (LSTM) network. The CNN comprises a first CNN and a second CNN for extracting vibration frequency domain and thermal image spatial information, while the LSTM comprises a first LSTM and a second LSTM for learning acoustic frequency variations and operating condition evolution time-series information. The time-spectrum map from the first modality's time-series data is input into the first CNN network, where deep spatial spectral features of the vibration signal in different frequency bands are extracted through convolutional kernel sliding, feature map mapping, and pooling operations, serving as the first feature vector. The second modality's time-series data is input into the second CNN network, where spatial convolution and local region aggregation are used to extract... The temperature characteristics of abnormal hot spots in the thermal image distribution are used as the second feature vector. The acoustic spectrum data from the third signal set are synchronized to the first LSTM network. The short-time change patterns of the acoustic signals are learned and analyzed using the gating structure of the LSTM to obtain the short-time dynamic parameters of the acoustic features, which are used as the third feature vector. Based on the fourth signal set, the multi-dimensional operating condition parameters are synchronized to the second LSTM network. The long-period dependence of the operating condition parameters is obtained through iterative calculation with time steps, forming a fourth feature vector that can characterize the evolution trend of the equipment's operating state. Finally, the first, second, third, and fourth feature vectors are divided into spatial and temporal domains and their feature structures are organized to construct a two-dimensional feature parameter that includes spatial and temporal features.

[0019] Furthermore, the construction process of the first CNN network, the second CNN network, the first LSTM network, and the second LSTM network includes the following methods: The first CNN network consists of at least two convolutional blocks and is used to extract frequency domain spatial features from the time-spectrum graph of the first modality time-series data, from local to global. The second CNN network consists of at least two convolutional blocks and is used to extract the spatial distribution morphology features of the temperature field from the image sequence of the second modality time-series data. The first LSTM network contains at least one LSTM unit and is used to process the acoustic spectrum data of the third signal set for short-term dynamic change capture. The second LSTM network contains at least one LSTM unit and is used to process the multi-dimensional operating condition parameters of the fourth signal set for long-term evolution analysis.

[0020] The first CNN network consists of at least two convolutional blocks, each containing a convolutional layer, a nonlinear activation layer, and a pooling layer. It progressively expands the receptive field through multiple convolutional kernels, extracting local window region features and converging global frequency domain features from the time-series data of the first modality to obtain frequency domain spatial features from local to global perspectives. The second CNN network also consists of at least two convolutional blocks. It extracts morphological and structural features and abnormal hotspot region features from the temperature field distribution by performing frame-by-frame convolution, local region weighting, and spatial feature aggregation on the image sequence of the second modality's time-series data. This is used to characterize the spatial development pattern of abnormal surface temperatures on equipment. The LSTM network contains at least one LSTM unit layer. It performs time-dimensional analysis on the acoustic spectrum data of the third signal set through input gates, forget gates, and output gates to capture the short-term dynamic changes of acoustic features between adjacent time steps, which is used to identify transient acoustic anomalies or short-period signal drift. The second LSTM network contains at least one LSTM unit layer. It performs iterative learning on the multi-dimensional operating condition parameters of the fourth signal set and uses the long-range dependency capability of the LSTM unit to extract the evolution trend and state change pattern of the operating condition parameters over a long period, so as to construct a time-dependent expression that can reflect the periodicity and slow-change characteristics of equipment operation.

[0021] The multi-head self-attention mechanism is activated to fuse the two-dimensional feature parameters, calculate the multimodal feature association weights, and generate a multimodal fused feature vector.

[0022] Furthermore, the method involves activating a multi-head self-attention mechanism to fuse the two-dimensional feature parameters, calculating multimodal feature association weights, and generating a multimodal fused feature vector. The first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are concatenated to form an initial fusion feature matrix. This initial fusion feature matrix is ​​then synchronized to a multi-head self-attention layer for linear transformation calculation, constructing a multi-head self-attention mechanism containing multiple attention heads. Normalization analysis is performed on the multiple attention heads to determine the attention weight distribution data. Based on the attention weight distribution data, the multiple attention heads are weighted and summed to obtain the multimodal feature association weights of the multiple attention heads. Based on the multimodal feature association weights, the first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are linearly projected to obtain the multimodal fusion feature vector.

[0023] Specifically, the first, second, third, and fourth feature vectors are sequentially concatenated along the feature dimension to form an initial fusion feature matrix with a unified feature dimension. This initial fusion feature matrix is ​​then input into a multi-head self-attention layer, where a query matrix, key matrix, and value matrix are generated through linear mapping to construct a multi-head self-attention mechanism containing multiple attention heads. Within each attention head, the association score between different modal features is calculated based on the similarity between the query matrix and the key matrix, and the association score is normalized to determine the attention weight distribution data. This attention weight distribution characterizes the implicit association strength of each modal feature with the diagnostic target. The value matrices of each attention head are weighted and summed according to the determined attention weight distribution to obtain the multimodal feature association weights corresponding to each attention head. Furthermore, linear projection processing is performed on the first, second, third, and fourth feature vectors based on the multimodal feature association weights, enabling weighted fusion of different modal features in the same embedding space, thereby obtaining a multimodal fusion feature vector for subsequent classification and diagnostic inference.

[0024] Furthermore, the first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are concatenated to form an initial fused feature matrix. The method includes: The first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are sequentially concatenated along the feature dimension to determine the target row vector; based on the target row vector, a two-dimensional matrix is ​​generated, and the two-dimensional matrix is ​​used as the initial fusion feature matrix, wherein each row or column of the initial fusion feature matrix represents an independent feature vector.

[0025] Specifically, the first, second, third, and fourth feature vectors obtained from the convolutional neural network and the long short-term memory network are sequentially concatenated along the feature dimension according to a preset modal order to form a target row vector containing all modal information. Based on the target row vector, a two-dimensional matrix is ​​generated by feature dimension copying, expansion, or matrix mapping, so that each row or column of the two-dimensional matrix represents the combination relationship of different modal features in a consistent structure. The two-dimensional matrix is ​​used as the initial fusion feature matrix. Each row or column of the initial fusion feature matrix represents an independent feature vector or feature combination unit, which is used to provide an input structure that can be processed and compared simultaneously by the multi-head self-attention mechanism.

[0026] Based on the multimodal fusion feature vector, the system backtracks to the power generation equipment for fully connected classification, identifies faults based on the classification results, and generates fault diagnosis results.

[0027] The multimodal fusion feature vector output by the multi-head self-attention mechanism is input into a pre-constructed deep neural network classifier. This classifier consists of multiple fully connected layers sequentially. Through linear mapping, activation function transformation, and feature compression operations, the fusion feature vector is reconstructed layer by layer to form a high-dimensional feature representation capable of distinguishing different operating states. In the final output layer, the output of the last fully connected layer is mapped to a dimension consistent with the number of target fault categories to obtain the classification result of the power generation equipment. Based on the classification result, the fusion feature vector is state-determined to generate equipment state parameters characterizing the equipment's operational health. These parameters include at least fault state parameters and health state parameters. Further, state comparison analysis is performed based on the probability distribution data of the equipment state parameters. When the fault state probability exceeds a preset threshold, the power generation equipment is determined to have a corresponding fault mode, thereby generating a fault diagnosis result. If the fault state probability does not reach the threshold but the health state probability is low, anomaly detection is triggered and a corresponding alarm result is generated to support subsequent maintenance decisions, achieving intelligent diagnosis and reliable identification of the power generation equipment's operating state.

[0028] Furthermore, based on the multimodal fusion feature vector, a fully connected classification is performed back to the power generation equipment. Fault identification is then performed based on the classification results, and fault diagnosis results are generated. The method includes: A deep neural network classifier is constructed, comprising multiple fully connected layers. The multimodal fusion feature vector is synchronized to the multiple fully connected layers for classification to obtain a classification result. Based on the classification result, the power generation equipment is subjected to state analysis to obtain equipment state parameters, including fault state parameters and health state parameters. Based on the fault state parameters and the health state parameters, distribution calculations are performed to obtain probability distribution data of the fault state parameters and the health state parameters. Based on the probability distribution data of the fault state parameters and the health state parameters, fault identification is performed to generate the fault diagnosis result.

[0029] Specifically, a deep neural network classifier is constructed, consisting of multiple fully connected layers stacked sequentially. Each fully connected layer compresses, abstracts, and reconstructs the input features through linear mapping and activation functions. The multimodal fusion feature vector, after being fused using a multi-head self-attention mechanism, is used as input and sequentially synchronized to the multiple fully connected layers. The final classification output is obtained through layer-by-layer feature propagation, yielding a classification result used to determine the operating status of the power generation equipment. Based on the classification result, the operating characteristics of the power generation equipment are analyzed to generate corresponding equipment status parameters. These parameters include fault status parameters describing the potential for equipment anomalies and health status parameters describing the normal operation of the equipment. Probability distribution models are constructed based on the fault status parameters and health status parameters, respectively, and their distributions are calculated to obtain the corresponding probability distribution data for the fault status parameters and health status parameters. Further, a status determination is performed based on the probability distribution data. When the fault status probability is higher than a set threshold and significantly exceeds the health status probability, the corresponding fault mode is identified and a fault diagnosis result is generated. When the health status probability is high, the equipment is determined to be in normal operating condition, thereby achieving intelligent identification and diagnostic output of the power generation equipment's operating status.

[0030] Furthermore, the method for synchronizing the multimodal fused feature vector to the multiple fully connected layers for classification to obtain classification results includes: The plurality of fully connected layers includes at least a first fully connected layer, a second fully connected layer, and an output layer; the first fully connected layer is used to receive the multimodal fusion feature vector, perform linear transformation, and generate a first output result; the second fully connected layer receives the first output result of the first fully connected layer, performs feature compression, and generates a second output result; the output layer maps the second output result generated by the second fully connected layer to a dimension with the same number of fault categories to obtain the classification result.

[0031] The plurality of fully connected layers include at least a first fully connected layer, a second fully connected layer, and an output layer. The first fully connected layer receives the multimodal fusion feature vector obtained through a multi-head self-attention mechanism, performs linear transformation and nonlinear activation on the input features, and generates a first output result, which is used to initially complete the dimensionality reduction and structural reconstruction of the high-dimensional fusion features. The second fully connected layer receives the first output result, enhances the feature abstraction capability through further linear transformation and feature compression operations, and generates a second output result, giving it stronger class discrimination characteristics. The output layer performs a final mapping on the second output result generated by the second fully connected layer, transforming the processing result to an output dimension corresponding to the preset number of fault categories, and obtains the classification probability through Softmax or other activation functions, thereby forming a classification result used to characterize the current operating state of the power generation equipment.

[0032] Furthermore, the method for identifying faults and generating fault diagnosis results based on the probability distribution data of the fault state parameters and the probability distribution data of the health state parameters includes: A probability threshold is set based on the probability distribution data of the health status parameters. The probability distribution data of the fault status parameters is compared with the probability threshold. When the probability distribution data of the fault status parameters exceeds the probability threshold, a fault is identified in the power generation equipment, and a fault mode is determined. When all fault status parameters are below the probability threshold and the health status probability is below a preset health threshold, an anomaly detection alarm is triggered. A confidence analysis is performed based on the fault mode to generate a first confidence level. A confidence analysis is performed based on the anomaly detection alarm to generate a second confidence level. Fault identification is performed based on the first confidence level, the second confidence level, the fault mode, and the anomaly detection alarm to generate the fault diagnosis result.

[0033] A probability threshold is set based on the probability distribution data of health status parameters to distinguish between normal and potentially abnormal states, and this threshold is used as a reference for judgment. The probability distribution data of fault status parameters is compared with the probability threshold. When the probability distribution data of fault status parameters exceeds the probability threshold, it is determined that the power generation equipment may have corresponding fault symptoms, and the specific fault mode is determined based on the classification results. When all fault status parameters are below the probability threshold and the probability distribution of health status parameters is below the preset health threshold, it indicates that although the equipment has not reached the fault threshold, its health has shown a downward trend. At this time, an anomaly detection alarm is triggered to indicate that the equipment has potential abnormalities. The system identifies fault patterns and performs confidence analysis based on fault modes. It then fits the fault state probability to the classification output to generate a first confidence level characterizing the diagnostic reliability. Based on the anomaly detection alarm, it performs confidence analysis on the health status fluctuation amplitude and risk accumulation index to generate a second confidence level. Finally, it combines the first and second confidence levels to comprehensively determine the fault mode identification results and anomaly detection results. When the fault confidence level is dominant, it outputs the specific fault mode; when the anomaly confidence level is dominant, it outputs the anomaly alarm mode. This generates a multi-modal fusion-driven final fault diagnosis result, achieving accurate identification and reliable representation of the power generation equipment's operating status.

[0034] In summary, the embodiments of this application have at least the following technical effects: First, dynamic data acquisition of the power generation equipment is performed to obtain multimodal heterogeneous data. This data is then spatiotemporally aligned to construct a multimodal data stream. Next, a joint dual-channel network is built to perform feature analysis on the multimodal data stream, generating two-dimensional feature parameters. Then, a multi-head self-attention mechanism is activated to fuse the two-dimensional feature parameters, calculate the multimodal feature association weights, and generate a multimodal fused feature vector. Finally, based on the multimodal fused feature vector, a fully connected classification is performed on the power generation equipment. Fault identification is then performed based on the classification results, generating fault diagnosis results. This approach solves the technical problem of low fault diagnosis accuracy caused by the difficulty in effectively integrating multimodal data in existing technologies. Through deep fusion of multimodal data, the technical effect of improving fault diagnosis accuracy is achieved.

[0035] Example 2, based on the same inventive concept as the power generation equipment fault diagnosis method based on multimodal data fusion in the foregoing examples, such as... Figure 2 As shown, this application provides a fault diagnosis system for power generation equipment based on multimodal data fusion, wherein the system includes: Data processing module 11: dynamically collects data on the operation of the power generation equipment to obtain multimodal heterogeneous data, aligns the multimodal heterogeneous data in time and space, and constructs a multimodal data stream; Feature analysis module 12: constructs a joint dual-channel network, performs feature analysis on the multimodal data stream through the joint dual-channel network, and generates two-dimensional feature parameters; Feature fusion module 13: activates a multi-head self-attention mechanism to fuse the two-dimensional feature parameters, calculates the multimodal feature association weights, and generates a multimodal fused feature vector; Fault identification module 14: backtracks to the power generation equipment based on the multimodal fused feature vector to perform fully connected classification, identifies faults based on the classification results, and generates fault diagnosis results.

[0036] Furthermore, the data processing module 11 is used to perform the following methods: A clock signal is introduced and used as a reference parameter. Based on this reference parameter, the multimodal heterogeneous data of the power generation equipment is decomposed according to multiple sampling frequencies to obtain a first signal set, a second signal set, and a third signal set. Vibration signals are extracted from the first signal set and subjected to short-time Fourier transform to generate a time-spectrum image, which serves as the first modal time-series data. Thermal image sequences are extracted from the second signal set, and regional interest extraction is performed on key components of the power generation equipment according to the thermal image sequences to obtain local thermal image sequences, which serve as the second modal time-series data. Acoustic spectrum data extracted from the third signal set is subjected to mutual information analysis with the first modal time-series data. Time offset compensation is performed on the acoustic spectrum data according to the mutual information parameters to obtain a fourth signal set. The first modal time-series data, the second modal time-series data, the third signal set, and the fourth signal set are reassembled along a time axis to generate the multimodal data stream.

[0037] Furthermore, the feature analysis module 12 is used to perform the following methods: A joint dual-channel network is constructed using a convolutional neural network (CNN) and a long short-term memory (LSTM) network. The CNN includes a first CNN network and a second CNN network, and the LSTM network includes a first LSTM network and a second LSTM network. The time-spectrum image of the first modality time-series data is synchronized to the first CNN network to extract the deep spatial spectrum features of the vibration signal as the first feature vector. The second modality time-series data is synchronized to the second CNN network to extract the abnormal hotspot features of the temperature distribution of the infrared thermal image as the second feature vector. The acoustic spectrum data of the third signal set is synchronized to the first LSTM network for acoustic feature learning to obtain the short-time change parameters of the acoustic features as the third feature vector. Based on the fourth signal set, multi-dimensional operating condition parameters are extracted and synchronized to the second LSTM network for operating condition learning and analysis to obtain the long-period temporal dependence of the learned operating condition parameters as the fourth feature vector. The first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are dimensionally divided to construct the dual-dimensional feature parameters.

[0038] Furthermore, the feature analysis module 12 is used to perform the following methods: The first CNN network consists of at least two convolutional blocks and is used to extract frequency domain spatial features from the time-spectrum graph of the first modality time-series data, from local to global. The second CNN network consists of at least two convolutional blocks and is used to extract the spatial distribution morphology features of the temperature field from the image sequence of the second modality time-series data. The first LSTM network contains at least one LSTM unit and is used to process the acoustic spectrum data of the third signal set for short-term dynamic change capture. The second LSTM network contains at least one LSTM unit and is used to process the multi-dimensional operating condition parameters of the fourth signal set for long-term evolution analysis.

[0039] Furthermore, the feature fusion module 13 is used to perform the following method: The first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are concatenated to form an initial fusion feature matrix. This initial fusion feature matrix is ​​then synchronized to a multi-head self-attention layer for linear transformation calculation, constructing a multi-head self-attention mechanism containing multiple attention heads. Normalization analysis is performed on the multiple attention heads to determine the attention weight distribution data. Based on the attention weight distribution data, the multiple attention heads are weighted and summed to obtain the multimodal feature association weights of the multiple attention heads. Based on the multimodal feature association weights, the first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are linearly projected to obtain the multimodal fusion feature vector.

[0040] Furthermore, the feature fusion module 13 is used to perform the following method: The first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are sequentially concatenated along the feature dimension to determine the target row vector; based on the target row vector, a two-dimensional matrix is ​​generated, and the two-dimensional matrix is ​​used as the initial fusion feature matrix, wherein each row or column of the initial fusion feature matrix represents an independent feature vector.

[0041] Furthermore, the fault identification module 14 is used to perform the following method: A deep neural network classifier is constructed, comprising multiple fully connected layers. The multimodal fusion feature vector is synchronized to the multiple fully connected layers for classification to obtain a classification result. Based on the classification result, the power generation equipment is subjected to state analysis to obtain equipment state parameters, including fault state parameters and health state parameters. Based on the fault state parameters and the health state parameters, distribution calculations are performed to obtain probability distribution data of the fault state parameters and the health state parameters. Based on the probability distribution data of the fault state parameters and the health state parameters, fault identification is performed to generate the fault diagnosis result.

[0042] Furthermore, the fault identification module 14 is used to perform the following method: The plurality of fully connected layers includes at least a first fully connected layer, a second fully connected layer, and an output layer; the first fully connected layer is used to receive the multimodal fusion feature vector, perform linear transformation, and generate a first output result; the second fully connected layer receives the first output result of the first fully connected layer, performs feature compression, and generates a second output result; the output layer maps the second output result generated by the second fully connected layer to a dimension with the same number of fault categories to obtain the classification result.

[0043] Furthermore, the fault identification module 14 is used to perform the following method: A probability threshold is set based on the probability distribution data of the health status parameters. The probability distribution data of the fault status parameters is compared with the probability threshold. When the probability distribution data of the fault status parameters exceeds the probability threshold, a fault is identified in the power generation equipment, and a fault mode is determined. When all fault status parameters are below the probability threshold and the health status probability is below a preset health threshold, an anomaly detection alarm is triggered. A confidence analysis is performed based on the fault mode to generate a first confidence level. A confidence analysis is performed based on the anomaly detection alarm to generate a second confidence level. Fault identification is performed based on the first confidence level, the second confidence level, the fault mode, and the anomaly detection alarm to generate the fault diagnosis result.

[0044] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.

Claims

1. A fault diagnosis method for power generation equipment based on multimodal data fusion, characterized in that, The method includes: Dynamic data acquisition is performed on the power generation equipment to obtain multimodal heterogeneous data. The multimodal heterogeneous data is then spatiotemporally aligned to construct a multimodal data stream. A joint dual-channel network is constructed, and feature analysis is performed on the multimodal data stream through the joint dual-channel network to generate two-dimensional feature parameters; The multi-head self-attention mechanism is activated to fuse the two-dimensional feature parameters, calculate the multimodal feature association weights, and generate a multimodal fused feature vector. Based on the multimodal fusion feature vector, the system backtracks to the power generation equipment for fully connected classification, identifies faults based on the classification results, and generates fault diagnosis results.

2. The fault diagnosis method for power generation equipment based on multimodal data fusion as described in claim 1, characterized in that, The method involves dynamically acquiring multimodal heterogeneous data from power generation equipment, spatiotemporally aligning this multimodal heterogeneous data, and constructing a multimodal data stream. A clock signal is introduced and used as a reference parameter. Based on the reference parameter, the multimodal heterogeneous data of the power generation equipment is decomposed according to multiple sampling frequencies to obtain a first signal set, a second signal set, and a third signal set. Based on the first signal set, the vibration signal is extracted and subjected to short-time Fourier transform to generate a time spectrum, which serves as the first mode time series data. Based on the second signal set, thermal image sequences are extracted, and regional interest extraction is performed by traversing the key components of the power generation equipment according to the thermal image sequences to obtain local thermal image sequences, which are used as second mode time series data. Based on the third signal set, the acoustic spectrum data is extracted and the first modal time series data are subjected to mutual information analysis. The acoustic spectrum data is then compensated for time offset according to the mutual information parameters to obtain the fourth signal set. The first modal timing data, the second modal timing data, the third signal set, and the fourth signal set are recombined according to the time axis to generate the multimodal data stream.

3. The fault diagnosis method for power generation equipment based on multimodal data fusion as described in claim 2, characterized in that, Constructing a joint dual-channel network, and performing feature analysis on the multimodal data stream through the joint dual-channel network to generate two-dimensional feature parameters, the method includes: A joint dual-channel network is constructed using a convolutional neural network and a long short-term memory network. The convolutional neural network includes a first CNN network and a second CNN network, and the long short-term memory network includes a first LSTM network and a second LSTM network. The time spectrum of the first modal time series data is synchronized to the first CNN network, and the deep spatial spectrum features of the vibration signal are extracted as the first feature vector. The second modality time series data is synchronized to the second CNN network, and the abnormal hot spot features of the temperature distribution of the infrared thermal image are extracted as the second feature vector; The acoustic spectrum data of the third signal set is synchronized to the first LSTM network for acoustic feature learning to obtain short-time change parameters of acoustic features, which are used as the third feature vector. Based on the fourth signal set, multi-dimensional operating condition parameters are extracted and synchronized to the second LSTM network for operating condition learning and analysis to obtain the long-term temporal dependency relationship of the learned operating condition parameters, which is used as the fourth feature vector. The first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are divided into dimensions to construct the two-dimensional feature parameters.

4. The fault diagnosis method for power generation equipment based on multimodal data fusion as described in claim 3, characterized in that, The construction process of the first CNN network, the second CNN network, the first LSTM network, and the second LSTM network includes the following methods: The first CNN network consists of at least two convolutional blocks and is used to extract frequency domain spatial features from the time spectrum of the first modality time series data from the local to the global perspective. The second CNN network consists of at least two convolutional blocks and is used to extract the spatial distribution morphological features of the temperature field from the image sequence of the second modality time series data. The first LSTM network contains at least one layer of LSTM units and is used to process the acoustic spectrum data of the third signal set to capture short-term dynamic changes. The second LSTM network contains at least one layer of LSTM units and is used to process the multi-dimensional operating condition parameters of the fourth signal set for long-term evolution analysis.

5. The fault diagnosis method for power generation equipment based on multimodal data fusion as described in claim 3, characterized in that, The method includes: activating a multi-head self-attention mechanism to fuse the two-dimensional feature parameters, calculating multimodal feature association weights, and generating a multimodal fused feature vector. The first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are concatenated to form an initial fusion feature matrix; The initial fused feature matrix is ​​synchronized to a multi-head self-attention layer for linear transformation calculation to construct a multi-head self-attention mechanism, which includes multiple attention heads. Normalization analysis is performed on all the attention heads to determine the attention weight distribution data; The multiple attention heads are weighted and summed based on the attention weight distribution data to obtain the multimodal feature association weights of the multiple attention heads; Based on the multimodal feature association weights, the first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are linearly projected to obtain the multimodal fusion feature vector.

6. The fault diagnosis method for power generation equipment based on multimodal data fusion as described in claim 5, characterized in that, The method involves concatenating the first feature vector, the second feature vector, the third feature vector, and the fourth feature vector to form an initial fused feature matrix. The first feature vector, the second feature vector, the third feature vector, and the fourth feature vector are sequentially concatenated along the feature dimension direction to determine the target row vector; Based on the target row vector, a two-dimensional matrix is ​​generated by expansion. The two-dimensional matrix is ​​used as the initial fusion feature matrix, wherein each row or column of the initial fusion feature matrix represents an independent feature vector.

7. The fault diagnosis method for power generation equipment based on multimodal data fusion as described in claim 1, characterized in that, Based on the multimodal fusion feature vector, a fully connected classification is performed back to the power generation equipment. Fault identification is then performed based on the classification results, and fault diagnosis results are generated. The method includes: Construct a deep neural network classifier, which contains multiple fully connected layers; The multimodal fused feature vectors are synchronized to the multiple fully connected layers for classification to obtain classification results; Based on the classification results, the power generation equipment is subjected to state analysis to obtain equipment state parameters, which include fault state parameters and health state parameters. Based on the fault state parameters and the health state parameters, a distribution calculation is performed to obtain the probability distribution data of the fault state parameters and the probability distribution data of the health state parameters. Fault identification is performed based on the probability distribution data of the fault state parameters and the probability distribution data of the health state parameters, and the fault diagnosis result is generated.

8. The fault diagnosis method for power generation equipment based on multimodal data fusion as described in claim 7, characterized in that, The method involves synchronizing the multimodal fused feature vector to the multiple fully connected layers for classification to obtain classification results. The plurality of fully connected layers includes at least a first fully connected layer, a second fully connected layer, and an output layer; The first fully connected layer is used to receive the multimodal fusion feature vector, perform a linear transformation, and generate a first output result; The second fully connected layer receives the first output result from the first fully connected layer, performs feature compression, and generates a second output result; The output layer maps the second output result generated by the second fully connected layer to a dimension equal to the number of fault categories to obtain the classification result.

9. The fault diagnosis method for power generation equipment based on multimodal data fusion as described in claim 7, characterized in that, Fault identification is performed based on the probability distribution data of the fault state parameters and the probability distribution data of the health state parameters, and the fault diagnosis result is generated. The method includes: A probability threshold is set based on the probability distribution data of the health status parameters, and the probability distribution data of the fault status parameters is compared with the probability threshold for determination. When the probability distribution data of the fault state parameters exceeds the probability threshold, it is determined that the power generation equipment has a fault, and fault identification is performed to determine the fault mode. When all the fault status parameters are below the probability threshold and the health status probability is below the preset health threshold, an anomaly detection alarm is triggered. Based on the aforementioned failure modes, a confidence analysis is performed to generate a first confidence level. Based on the anomaly detection alarm, a confidence analysis is performed to generate a second confidence level; Based on the first confidence level, the second confidence level, the fault mode, and the anomaly detection alarm, fault identification is performed, and the fault diagnosis result is generated.

10. A fault diagnosis system for power generation equipment based on multimodal data fusion, characterized in that, The system is used to implement the power generation equipment fault diagnosis method based on multimodal data fusion as described in any one of claims 1-9, the system comprising: Data processing module: performs dynamic data acquisition on the power generation equipment to obtain multimodal heterogeneous data, performs spatiotemporal alignment on the multimodal heterogeneous data, and constructs a multimodal data stream; Feature analysis module: Constructs a joint dual-channel network, performs feature analysis on the multimodal data stream through the joint dual-channel network, and generates two-dimensional feature parameters; Feature fusion module: Activates multi-head self-attention mechanism to fuse the two-dimensional feature parameters, calculates multimodal feature association weights, and generates multimodal fusion feature vector; Fault identification module: Based on the multimodal fusion feature vector, backtrack to the power generation equipment for fully connected classification, identify faults based on the classification results, and generate fault diagnosis results.