Power equipment fault diagnosis method and system based on multi-modal data fusion

Through the attention mechanism and BiLSTM model, the multimodal data of power equipment is characterized by fusion and training, which solves the problem of high computational complexity in high-dimensional modal feature processing of cross attention mechanisms, and realizes efficient power equipment fault diagnosis.

CN120541370APending Publication Date: 2025-08-26GUANGZHOU CITY UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510613002.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the existing power equipment fault diagnosis method based on multimodal information fusion, the cross attention mechanism has a high computational complexity when dealing with high-dimensional modal features, resulting in excessive computing resource consumption.

Method used

The attention mechanism is used to weighted fusion of multimodal features, combined with the Bidirectional Long Short-term Memory Network (BiLSTM) model, feature extraction and training is extracted and trained by collecting multimodal data from power equipment, and the attention mechanism is used to reduce complex correlation calculations between cross-modal features, focusing only on the attention weight allocation within each mode.

Benefits of technology

It significantly reduces the computational complexity, improves processing efficiency, improves the accuracy of power equipment fault diagnosis and the ability to handle high-dimensional modal features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541370A_ABST
    Figure CN120541370A_ABST
Patent Text Reader

Abstract

The invention discloses a power equipment fault diagnosis method and system based on multi-modal data fusion, and the method comprises the following steps: collecting the multi-modal data of the operation of power equipment, carrying out the preprocessing of the multi-modal data, obtaining the preprocessed multi-modal data, and dividing the preprocessed multi-modal data into a training data set and a test data set; performing feature extraction on the training data set to obtain multi-modal features; carrying out weighted fusion on the multi-modal features by utilizing an attention mechanism; a BiLSTM model is constructed; using the fused multi-modal features to train a BiLSTM (Bidirectional Long Short Term Memory) model; and inputting real-time operation data of the power equipment into the trained BiLSTM model to carry out fault diagnosis on the power equipment. The problem that in an existing fault diagnosis method based on multi-modal information fusion, although a cross attention mechanism can capture data association of different modals, when a cross-modal attention weight matrix is calculated, processing of high-dimensional modal features can cause calculation complexity to be increased, and a large number of calculation resources are consumed is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power equipment fault diagnosis, and in particular to a power equipment fault diagnosis method and system based on multimodal data fusion. Background Art

[0002] In recent years, with the development of modern power systems, the operating environment of power equipment has become increasingly complex. Rapid grid load growth, the integration of new energy sources, complex and changing electricity consumption patterns, and the coexistence of aging power equipment have led to a continuous increase in power equipment failure rates. To ensure the stable operation of the power system, power equipment fault diagnosis has become a crucial technical means to ensure power supply security. Currently, to improve the accuracy of power equipment fault diagnosis, most power equipment fault diagnosis technologies rely on the analysis of multimodal data.

[0003] Prior art discloses a fault diagnosis method based on multimodal information fusion. This method acquires multimodal data from power equipment and then uses a cross-attention mechanism to achieve deep fusion of this multimodal data, providing strong support for subsequent fault diagnosis. However, while the cross-attention mechanism can effectively capture the correlation information between different modal data when processing multimodal data fusion, it requires the calculation of a cross-modal attention weight matrix. This process is relatively efficient when processing low-dimensional modal features, but the computational complexity increases when dealing with high-dimensional modal features, resulting in a significant loss of computing resources. Summary of the Invention

[0004] In response to the above-mentioned defects, the present invention proposes a method and system for fault diagnosis of power equipment based on multimodal data fusion, with the aim of solving the problem that in the existing fault diagnosis method based on multimodal information fusion, multimodal data is fused using a cross-attention mechanism. Although the cross-attention mechanism can capture the association between different modal data, when calculating the cross-modal attention weight matrix, processing high-dimensional modal features will lead to increased computational complexity and consume a large amount of computing resources.

[0005] To achieve this object, the present invention adopts the following technical solutions:

[0006] A method for diagnosing faults in power equipment based on multimodal data fusion includes the following steps:

[0007] Step S1: collecting multimodal data of the operation of the power equipment, wherein the multimodal data of the operation of the power equipment includes thermal infrared imaging images, sound data of the equipment, and vibration data of the equipment;

[0008] Step S2: preprocessing the multimodal data of the power equipment operation to obtain the preprocessed multimodal data, and dividing the preprocessed multimodal data into a training data set and a test data set;

[0009] Step S3: extract features from the training data set to obtain multimodal features;

[0010] Step S4: Use the attention mechanism to perform weighted fusion on the multimodal features to obtain the fused multimodal features;

[0011] Step S5: construct a bidirectional long short-term memory network BiLSTM model;

[0012] Step S6: Use the fused multimodal features to train the BiLSTM model to obtain a trained BiLSTM model;

[0013] Step S7: Collecting real-time data of the power equipment operation and inputting it into the trained BiLSTM model to perform fault diagnosis of the power equipment and output the fault prediction value of the power equipment;

[0014] Step S8: Calculate the cross entropy loss function value according to the test data set and the fault prediction value of the power equipment, and evaluate the prediction accuracy of the trained BiLSTM model based on the cross entropy loss function value.

[0015] Preferably, in step S2, the multimodal data of the operation of the power equipment is preprocessed to obtain the preprocessed multimodal data, which specifically includes the following sub-steps:

[0016] Step S21: using Gaussian filtering to remove noise from the thermal infrared imaging image, and using wavelet transform to remove noise from the sound data and vibration data of the device respectively;

[0017] Step S22: using linear interpolation to fill in missing values ​​in the denoised thermal infrared imaging image, the device sound data, and the device vibration data;

[0018] Step S23: normalizing the padded thermal infrared imaging image, the device sound data, and the device vibration data to obtain normalized thermal infrared imaging image, the device sound data, and the device vibration data.

[0019] Preferably, step S3 specifically includes the following sub-steps:

[0020] The convolutional neural network (CNN) algorithm is used to extract features from the normalized thermal infrared imaging image to obtain the temperature distribution features. The mathematical expression of the specific feature extraction is as follows:

[0021] F image =CNN(X image );

[0022] Among them, F image Represents the temperature distribution characteristics, X imagerepresents the normalized thermal infrared imaging image;

[0023] Short-time Fourier transform (STFT) is used to extract features from the normalized device sound data to obtain time-frequency features. The mathematical expression for the specific feature extraction is as follows:

[0024]

[0025] Among them, F sound (τ, ω) represents the time-frequency characteristics, x sound (n) represents the normalized device sound data, w(n) represents the window function, τ represents time, and ω represents frequency;

[0026] Fast Fourier transform (FFT) is used to extract features from the normalized vibration data of the equipment to obtain frequency domain features. The mathematical expression for the specific feature extraction is as follows:

[0027]

[0028] Among them, F vibration (k) represents the kth frequency domain feature, x vibration (n) represents the normalized vibration data of the device, and N represents the total number of normalized vibration data of the device.

[0029] Preferably, step S4 specifically includes the following sub-steps:

[0030] Use the attention mechanism to transform the temperature distribution feature F image , time-frequency features F sound (τ, ω) and frequency domain features F vibration (k) Fusion is performed to obtain the fused multimodal feature F concat , where F concat The specific mathematical expression is as follows:

[0031] F concat =[F image , F sound (τ,ω),F vibration (k)];

[0032] The mathematical expression of the attention mechanism is as follows:

[0033]

[0034] Where Q represents the query matrix; K represents the key matrix; V represents the value matrix; d k represents the dimension of the key; softmax(x) represents the normalization function; Attention(x) represents the function of the attention mechanism.

[0035] Preferably, step S6 specifically includes the following sub-steps:

[0036] Step S61: The fused multimodal features F concat Input the BiLSTM model to process the multi-layer network structure and obtain the corresponding hidden state h t , where h t The specific mathematical expression is as follows:

[0037] h t =BiLSTM(F concat );

[0038] Step S62: h t Convert the result into a fault classification type to complete the training of the BiLSTM model. The specific conversion calculation formula is as follows:

[0039] y=softmax(W·h t +b);

[0040] Where y represents the fault classification type result, W represents the weight matrix, and b represents the bias term.

[0041] Preferably, in step S8, the cross entropy loss function value is calculated based on the test data set and the fault prediction value of the power equipment. The specific calculation formula is as follows:

[0042]

[0043] Among them, L represents the cross entropy loss function value, y i represents the true value of the i-th sample, represents the predicted value of the BiLSTM model for the i-th sample, and N' represents the total number of samples in the test dataset.

[0044] Preferably, the method further includes the following steps: optimizing the parameters of the BiLSTM model using an Adam optimizer, wherein the mathematical expression for optimizing the parameters of the BiLSTM model is as follows:

[0045]

[0046] Among them, θ t represents the parameters optimized in the tth round of the BiLSTM model, η represents the learning rate, m t represents the first-order momentum estimate, v t Represents the second-order momentum estimate, ∈ represents a positive real number, and its value range is 10 -8 ~10 -10 .

[0047] Another aspect of the present application provides a power equipment fault diagnosis system based on multimodal data fusion, the system comprising:

[0048] A first acquisition module is used to collect multimodal data of the operation of the power equipment, wherein the multimodal data of the operation of the power equipment includes thermal infrared imaging images, sound data of the equipment, and vibration data of the equipment;

[0049] A data preprocessing module is used to preprocess the multimodal data of the power equipment operation to obtain preprocessed multimodal data;

[0050] Data partitioning module, used to divide the preprocessed multimodal data into training data set and test data set;

[0051] Feature extraction module, used to extract features from the training data set to obtain multimodal features;

[0052] The feature fusion module is used to perform weighted fusion of multimodal features using the attention mechanism to obtain fused multimodal features;

[0053] Building module for constructing BiLSTM model;

[0054] The model training module is used to train the BiLSTM model using the fused multimodal features to obtain a trained BiLSTM model;

[0055] The second acquisition module is used to collect real-time data of power equipment operation;

[0056] The fault diagnosis module is used to input the real-time data of power equipment operation into the trained BiLSTM model to perform fault diagnosis on the power equipment and output the fault prediction value of the power equipment;

[0057] A calculation module, used to calculate a cross entropy loss function value based on a test data set and a fault prediction value of the power equipment;

[0058] The evaluation module is used to evaluate the prediction accuracy of the trained BiLSTM model based on the cross-entropy loss function value.

[0059] Preferably, the data preprocessing module includes:

[0060] A noise removal submodule is used to remove noise from thermal infrared imaging images using Gaussian filtering, and to remove noise from the device's sound data and vibration data using wavelet transform.

[0061] A missing value filling submodule is used to fill the missing values ​​in the denoised thermal infrared imaging image, the device sound data, and the device vibration data using a linear interpolation method;

[0062] A normalization processing submodule is used to perform normalization processing on the padded thermal infrared imaging image, the sound data of the device, and the vibration data of the device, respectively, to obtain normalized thermal infrared imaging image, the sound data of the device, and the vibration data of the device;

[0063] The feature extraction module includes:

[0064] The first feature extraction submodule is used to extract features from the normalized thermal infrared imaging image using a convolutional neural network (CNN) algorithm to obtain temperature distribution features. The specific mathematical expression for feature extraction is as follows:

[0065] F image =CNN(X image );

[0066] Among them, F image Represents the temperature distribution characteristics, X image represents the normalized thermal infrared imaging image;

[0067] The second feature extraction submodule is used to extract features from the normalized device sound data using short-time Fourier transform (STFT) to obtain time-frequency features. The mathematical expression for the specific feature extraction is as follows:

[0068]

[0069] Among them, F sound (τ, ω) represents the time-frequency characteristics, x sound (n) represents the normalized device sound data, w(n) represents the window function, τ represents time, and ω represents frequency;

[0070] The third feature extraction submodule is used to extract features from the normalized vibration data of the device using the fast Fourier transform (FFT) to obtain frequency domain features. The mathematical expression of the specific feature extraction is as follows:

[0071]

[0072] Among them, F vibration (k) represents the kth frequency domain feature, x vibration (n) represents the normalized vibration data of the device, and N represents the total number of normalized vibration data of the device;

[0073] The feature fusion module includes:

[0074] The feature fusion submodule is used to use the attention mechanism to integrate the temperature distribution feature F image , time-frequency features F sound (τ, ω) and frequency domain features F vibration(k) Fusion is performed to obtain the fused multimodal feature F concat , where F concat The specific mathematical expression is as follows:

[0075] F concat =[F image , F sound (τ,ω),F vibration (k)];

[0076] The mathematical expression of the attention mechanism is as follows:

[0077]

[0078] Where Q represents the query matrix; K represents the key matrix; V represents the value matrix; d k represents the dimension of the key; softmax(x) represents the normalization function; Attention(x) represents the function of the attention mechanism;

[0079] The model training module includes:

[0080] The multi-layer network structure processing submodule is used to transform the fused multimodal features E concat Input the BiLSTM model to process the multi-layer network structure and obtain the corresponding hidden state h t , where h t The specific mathematical expression is as follows:

[0081] h t =BiLSTM(F concat );

[0082] The conversion submodule is used to convert h t Convert the result into a fault classification type to complete the training of the BiLSTM model. The specific conversion calculation formula is as follows:

[0083] y=softmax(W·h t +b);

[0084] Where y represents the fault classification type result, W represents the weight matrix, and b represents the bias term.

[0085] Preferably, the method further includes: a model parameter optimization module for optimizing the parameters of the BiLSTM model using an Adam optimizer, wherein the mathematical expression for optimizing the parameters of the BiLSTM model is as follows:

[0086]

[0087] Among them, θ t represents the parameters optimized in the tth round of the BiLSTM model, η represents the learning rate, mt represents the first-order momentum estimate, v t Represents the second-order momentum estimate, ∈ represents a positive real number, and its value range is 10 -8 ~10 -10 .

[0088] The technical solutions provided by the embodiments of the present application may have the following beneficial effects:

[0089] This solution collects multimodal data from power equipment operation and extracts features. It then uses an attention mechanism to fuse the multimodal features. The fused multimodal features are then used to train a pre-built BiLSTM model. The trained BiLSTM model is then used to perform real-time power equipment fault diagnosis. This solution fuses multimodal features using an attention mechanism. Compared to the traditional cross-attention mechanism, the attention mechanism employed in this solution avoids the complex association calculations between cross-modal features and focuses solely on allocating attention weights within each modality. This reduces the number of feature associations that need to be calculated when processing high-dimensional modal features, significantly reducing computational complexity and improving processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1 The present invention is a flowchart of the steps of a power equipment fault diagnosis method based on multimodal data fusion. DETAILED DESCRIPTION

[0091] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and are not to be construed as limiting the present invention.

[0092] A method for diagnosing faults in power equipment based on multimodal data fusion includes the following steps:

[0093] Step S1: collecting multimodal data of the operation of the power equipment, wherein the multimodal data of the operation of the power equipment includes thermal infrared imaging images, sound data of the equipment, and vibration data of the equipment;

[0094] Step S2: preprocessing the multimodal data of the power equipment operation to obtain the preprocessed multimodal data, and dividing the preprocessed multimodal data into a training data set and a test data set;

[0095] Step S3: extract features from the training data set to obtain multimodal features;

[0096] Step S4: Use the attention mechanism to perform weighted fusion on the multimodal features to obtain the fused multimodal features;

[0097] Step S5: construct a bidirectional long short-term memory network BiLSTM model;

[0098] Step S6: Use the fused multimodal features to train the BiLSTM model to obtain a trained BiLSTM model;

[0099] Step S7: Collecting real-time data of the power equipment operation and inputting it into the trained BiLSTM model to perform fault diagnosis of the power equipment and output the fault prediction value of the power equipment;

[0100] Step S8: Calculate the cross entropy loss function value according to the test data set and the fault prediction value of the power equipment, and evaluate the prediction accuracy of the trained BiLSTM model based on the cross entropy loss function value.

[0101] This solution is a power equipment fault diagnosis method based on multimodal data fusion, such as Figure 1As shown, the first step is to collect multimodal data on the operation of power equipment. This multimodal data includes thermal infrared imaging images, equipment sound data, and equipment vibration data. In this embodiment, collecting this multimodal data provides a data foundation for subsequent diagnosis of power equipment faults. Thermal infrared imaging images can reflect abnormal temperature distribution of power equipment, equipment sound data can reflect wear or other abnormal conditions of internal mechanical components, and equipment vibration data can reflect abnormal vibration during operation. The second step is to preprocess the multimodal data to obtain preprocessed multimodal data and divide it into a training dataset and a test dataset. In this embodiment, preprocessing the multimodal data helps improve the quality of the multimodal data. Dividing the preprocessed multimodal data into a training dataset and a test dataset facilitates subsequent training of the BiLSTM model, while dividing the test dataset facilitates subsequent testing of the BiLSTM model. The third step is to extract features from the training dataset to obtain multimodal features. In this embodiment, extracting multimodal features helps capture characteristic changes in the multimodal data. The fourth step is to use an attention mechanism to perform weighted fusion of multimodal features to obtain fused multimodal features. In this embodiment, the attention mechanism, by learning the correlations between different modal features, can effectively reduce the interference of irrelevant features on the fusion results, thereby improving the accuracy of the fusion results. The fifth step is to construct a bidirectional long-short-term memory (BiLSTM) model. In this embodiment, since the operating data of power equipment is a time-varying sequence, the BiLSTM model, through its internal memory units and gating mechanism, can effectively capture the long-term and short-term temporal dependencies in this data, thereby more comprehensively determining whether the power equipment is in a fault state. The sixth step is to train the BiLSTM model using the fused multimodal features to obtain a trained BiLSTM model. In this embodiment, training the BiLSTM model helps improve the BiLSTM model's accuracy in identifying power equipment faults. The seventh step is to collect real-time power equipment operating data and input it into the trained BiLSTM model for power equipment fault diagnosis, outputting a fault prediction value. In this embodiment, using the BiLSTM model to predict the future operating status of power equipment facilitates operation and maintenance personnel to formulate maintenance plans in advance, thereby reducing operation and maintenance costs. The eighth step is to calculate the cross-entropy loss function value based on the test data set and the fault prediction value of the power equipment, and evaluate the prediction accuracy of the trained BiLSTM model based on the cross-entropy loss function value. In this embodiment, by calculating the cross-entropy loss function value, it is helpful to measure the difference between the prediction result and the true label, thereby better optimizing the BiLSTM model.

[0102] This solution collects multimodal data from power equipment operation and extracts features. It then uses an attention mechanism to fuse the multimodal features. The fused multimodal features are then used to train a pre-built BiLSTM model. The trained BiLSTM model is then used to perform real-time power equipment fault diagnosis. This solution fuses multimodal features using an attention mechanism. Compared to the traditional cross-attention mechanism, the attention mechanism employed in this solution avoids the complex association calculations between cross-modal features and focuses solely on allocating attention weights within each modality. This reduces the number of feature associations that need to be calculated when processing high-dimensional modal features, significantly reducing computational complexity and improving processing efficiency.

[0103] Preferably, in step S2, the multimodal data of the operation of the power equipment is preprocessed to obtain the preprocessed multimodal data, which specifically includes the following sub-steps:

[0104] Step S21: using Gaussian filtering to remove noise from the thermal infrared imaging image, and using wavelet transform to remove noise from the sound data and vibration data of the device respectively;

[0105] Step S22: using linear interpolation to fill in missing values ​​in the denoised thermal infrared imaging image, the device sound data, and the device vibration data;

[0106] Step S23: normalizing the padded thermal infrared imaging image, the device sound data, and the device vibration data to obtain normalized thermal infrared imaging image, the device sound data, and the device vibration data.

[0107] In this embodiment, in step S21, a Gaussian filter algorithm is used to remove noise interference from the thermal infrared image. Wavelet transform technology is applied to the device's sound data and vibration data to perform denoising to improve data purity. In step S22, missing values ​​in the data are filled using linear interpolation to ensure data integrity and consistency. In step S23, the padded thermal infrared image, device sound data, and device vibration data are normalized to unify the data scale and facilitate subsequent processing by the BiLSTM model.

[0108] Preferably, step S3 specifically includes the following sub-steps:

[0109] The convolutional neural network (CNN) algorithm is used to extract features from the normalized thermal infrared imaging image to obtain the temperature distribution features. The mathematical expression of the specific feature extraction is as follows:

[0110] Fimage =CNN(X image );

[0111] Among them, F image Represents the temperature distribution characteristics, X image represents the normalized thermal infrared imaging image;

[0112] Short-time Fourier transform (STFT) is used to extract features from the normalized device sound data to obtain time-frequency features. The mathematical expression for the specific feature extraction is as follows:

[0113]

[0114] Among them, F sound (τ, ω) represents the time-frequency characteristics, x sound (n) represents the normalized device sound data, w(n) represents the window function, τ represents time, and ω represents frequency;

[0115] Fast Fourier transform (FFT) is used to extract features from the normalized vibration data of the equipment to obtain frequency domain features. The mathematical expression for the specific feature extraction is as follows:

[0116]

[0117] Among them, F vibration (k) represents the kth frequency domain feature, x vibration (n) represents the normalized vibration data of the device, and N represents the total number of normalized vibration data of the device.

[0118] In this embodiment, by using the convolutional neural network (CNN) algorithm to extract features from the normalized thermal infrared imaging image, CNN can effectively extract key thermal features, such as abnormally hot areas of the equipment, and reduce the impact of background noise on the detection results through multi-layer convolution and pooling operations. By using the short-time Fourier transform (STFT) to extract features from the normalized sound data of the equipment, STFT divides the sound data into short-time segments through a sliding window, and performs Fourier transform on each segment, it can extract the features of the sound signal in the time and frequency dimensions. By using the fast Fourier transform (FFT) to extract features from the normalized vibration data of the equipment, FFT quickly converts the time domain vibration data into the frequency domain, and can intuitively extract the frequency components of the equipment vibration data, revealing the core characteristics of the equipment's operating status.

[0119] Preferably, step S4 specifically includes the following sub-steps:

[0120] Use the attention mechanism to transform the temperature distribution feature F image , time-frequency features F sound(τ, ω) and frequency domain features F vibration (k) Fusion is performed to obtain the fused multimodal feature F concat , where F concat The specific mathematical expression is as follows:

[0121] F concat =[F image , F sound (τ,ω),F vibration (k)];

[0122] The mathematical expression of the attention mechanism is as follows:

[0123]

[0124] Where Q represents the query matrix; K represents the key matrix; V represents the value matrix; d k represents the dimension of the key; softmax(x) represents the normalization function; Attention(x) represents the function of the attention mechanism.

[0125] In this embodiment, the attention mechanism can dynamically adjust the weights of temperature distribution features, time-frequency features, and frequency domain features in the fusion process. It can identify the correlation patterns between different features so that the unified features obtained by fusion can better reflect the intrinsic connection between the features, thereby improving the quality of feature fusion.

[0126] Preferably, step S6 specifically includes the following sub-steps:

[0127] Step S61: The fused multimodal features F concat Input the BiLSTM model to process the multi-layer network structure and obtain the corresponding hidden state h t , where h t The specific mathematical expression is as follows:

[0128] h t =BiLSTM(F concat );

[0129] Step S62: h t Convert the result into a fault classification type to complete the training of the BiLSTM model. The specific conversion calculation formula is as follows:

[0130] y=softmax(W·h t +b);

[0131] Where y represents the fault classification type result, W represents the weight matrix, and b represents the bias term.

[0132] In this embodiment, in step S61, by inputting the fused multimodal features into the BiLSTM model for multi-layer network structure processing, the BiLSTM model can capture the contextual dependencies in the fused multimodal features, thereby more accurately judging the operating status and fault development trend of the power equipment. t The data is converted into fault classification type results, so that the BiLSTM model can accurately classify the faults of power equipment according to the input multimodal feature data.

[0133] Preferably, in step S8, the cross entropy loss function value is calculated based on the test data set and the fault prediction value of the power equipment. The specific calculation formula is as follows:

[0134]

[0135] Among them, L represents the cross entropy loss function value, y i represents the true value of the i-th sample, represents the predicted value of the BiLSTM model for the i-th sample, and N' represents the total number of samples in the test dataset.

[0136] In this embodiment, by calculating the cross entropy loss function, it can help evaluate the accuracy of the BiLSTM model in predicting power equipment faults.

[0137] Preferably, the method further includes the following steps: optimizing the parameters of the BiLSTM model using an Adam optimizer, wherein the mathematical expression for optimizing the parameters of the BiLSTM model is as follows:

[0138]

[0139] Among them, θ t represents the parameters optimized in the tth round of the BiLSTM model, η represents the learning rate, m t represents the first-order momentum estimate, v t Represents the second-order momentum estimate, ∈ represents a positive real number, and its value range is 10 -8 ~10 -10 .

[0140] In this embodiment, the Adam optimizer is an optimization algorithm applied to deep learning models. By using the Adam optimizer to optimize the parameters of the BiLSTM model, the Adam optimizer can quickly and effectively reduce the loss function value due to its adaptive learning rate and momentum characteristics.

[0141] Another aspect of the present application provides a power equipment fault diagnosis system based on multimodal data fusion, the system comprising:

[0142] A first acquisition module is used to collect multimodal data of the operation of the power equipment, wherein the multimodal data of the operation of the power equipment includes thermal infrared imaging images, sound data of the equipment, and vibration data of the equipment;

[0143] A data preprocessing module is used to preprocess the multimodal data of the power equipment operation to obtain preprocessed multimodal data;

[0144] Data partitioning module, used to divide the preprocessed multimodal data into training data set and test data set;

[0145] Feature extraction module, used to extract features from the training data set to obtain multimodal features;

[0146] The feature fusion module is used to perform weighted fusion of multimodal features using the attention mechanism to obtain fused multimodal features;

[0147] Building module for constructing BiLSTM model;

[0148] The model training module is used to train the BiLSTM model using the fused multimodal features to obtain a trained BiLSTM model;

[0149] The second acquisition module is used to collect real-time data of power equipment operation;

[0150] The fault diagnosis module is used to input the real-time data of power equipment operation into the trained BiLSTM model to perform fault diagnosis on the power equipment and output the fault prediction value of the power equipment;

[0151] A calculation module, used to calculate a cross entropy loss function value based on a test data set and a fault prediction value of the power equipment;

[0152] The evaluation module is used to evaluate the prediction accuracy of the trained BiLSTM model based on the cross-entropy loss function value.

[0153] This solution is based on a multimodal data fusion power equipment fault diagnosis system. Through the cooperation of a first acquisition module, a data preprocessing module, a data partitioning module, a feature extraction module, a feature fusion module, a construction module, a model training module, a second acquisition module, a fault diagnosis module, a calculation module, and an evaluation module, the fault diagnosis of power equipment is achieved. In this solution, multimodal features are fused through an attention mechanism. Compared with the traditional cross-attention mechanism, the attention mechanism adopted in this solution avoids the complex correlation calculations between cross-modal features and focuses only on the allocation of attention weights within each modality. This reduces the number of feature correlations that need to be calculated when processing high-dimensional modal features, thereby significantly reducing the computational complexity and improving processing efficiency.

[0154] Preferably, the data preprocessing module includes:

[0155] A noise removal submodule is used to remove noise from thermal infrared imaging images using Gaussian filtering, and to remove noise from the device's sound data and vibration data using wavelet transform.

[0156] A missing value filling submodule is used to fill the missing values ​​in the denoised thermal infrared imaging image, the device sound data, and the device vibration data using a linear interpolation method;

[0157] A normalization processing submodule is used to perform normalization processing on the padded thermal infrared imaging image, the sound data of the device, and the vibration data of the device, respectively, to obtain normalized thermal infrared imaging image, the sound data of the device, and the vibration data of the device;

[0158] The feature extraction module includes:

[0159] The first feature extraction submodule is used to extract features from the normalized thermal infrared imaging image using a convolutional neural network (CNN) algorithm to obtain temperature distribution features. The specific mathematical expression for feature extraction is as follows:

[0160] F image =CNN(X image );

[0161] Among them, F image Represents the temperature distribution characteristics, X image represents the normalized thermal infrared imaging image;

[0162] The second feature extraction submodule is used to extract features from the normalized device sound data using short-time Fourier transform (STFT) to obtain time-frequency features. The mathematical expression for the specific feature extraction is as follows:

[0163]

[0164] Among them, F sound (τ, ω) represents the time-frequency characteristics, x sound (n) represents the normalized device sound data, w(n) represents the window function, τ represents time, and ω represents frequency;

[0165] The third feature extraction submodule is used to extract features from the normalized vibration data of the device using the fast Fourier transform (FFT) to obtain frequency domain features. The specific mathematical expression of feature extraction is as follows:

[0166]

[0167] Among them, F vibration (k) represents the kth frequency domain feature, x vibration (n) represents the normalized vibration data of the device, and N represents the total number of normalized vibration data of the device;

[0168] The feature fusion module includes:

[0169] The feature fusion submodule is used to use the attention mechanism to integrate the temperature distribution feature F image , time-frequency features F sound (τ, ω) and frequency domain features F vibration (k) Fusion is performed to obtain the fused multimodal feature F concat , where F concat The specific mathematical expression is as follows:

[0170] F concat =[F image , F sound (τ,ω),F vibration (k);

[0171] The mathematical expression of the attention mechanism is as follows:

[0172]

[0173] Where Q represents the query matrix; K represents the key matrix; V represents the value matrix; d k represents the dimension of the key; softmax(x) represents the normalization function; Attention(x) represents the function of the attention mechanism;

[0174] The model training module includes:

[0175] The multi-layer network structure processing submodule is used to transform the fused multimodal features F concat Input the BiLSTM model to process the multi-layer network structure and obtain the corresponding hidden state h t , where h t The specific mathematical expression is as follows:

[0176] h t =BiLSTM(F concat );

[0177] The conversion submodule is used to convert h t Convert the result into a fault classification type to complete the training of the BiLSTM model. The specific conversion calculation formula is as follows:

[0178] t=softmax(W·h t +b);

[0179] Where y represents the fault classification type result, W represents the weight matrix, and b represents the bias term.

[0180] In this embodiment, the noise removal submodule in the data preprocessing module helps improve data purity. The missing value filling submodule helps ensure data integrity and consistency. The normalization submodule helps unify data scale, facilitating subsequent BiLSTM model processing. In the feature extraction module, the first feature extraction submodule effectively extracts key thermal features, such as abnormally hot areas in the equipment, and reduces the impact of background noise on detection results. The second feature extraction submodule extracts the time and frequency characteristics of the sound signal. The third feature extraction submodule intuitively extracts the frequency components of the equipment vibration data, revealing the core characteristics of the equipment's operating status. In the feature fusion module, the feature fusion submodule better reflects the inherent connections between temperature distribution features, time-frequency features, and frequency domain features. In the model training module, the multi-layer network structure processing submodule helps capture the contextual dependencies in the fused multimodal features, thereby more accurately determining the operating status and fault development trends of the power equipment. The conversion submodule enables the BiLSTM model to accurately classify power equipment faults based on the input multimodal feature data.

[0181] Preferably, the method further includes: a model parameter optimization module for optimizing the parameters of the BiLSTM model using an Adam optimizer, wherein the mathematical expression for optimizing the parameters of the BiLSTM model is as follows:

[0182]

[0183] Among them, θ t represents the parameters optimized in the tth round of the BiLSTM model, η represents the learning rate, m t represents the first-order momentum estimate, v t Represents the second-order momentum estimate, ∈ represents a positive real number, and its value range is 10 -8 ~10 -10 .

[0184] In this embodiment, by setting a model parameter optimization module, the loss function value can be quickly and effectively reduced.

[0185] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.

[0186] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A method for fault diagnosis of electric power equipment based on multimodal data fusion, characterized by: The following steps are involved: Step S1: collecting multimodal data of the operation of the power equipment, wherein the multimodal data of the operation of the power equipment includes thermal infrared imaging images, sound data of the equipment, and vibration data of the equipment; Step S2: preprocessing the multimodal data of the power equipment operation to obtain the preprocessed multimodal data, and dividing the preprocessed multimodal data into a training data set and a test data set; Step S3: extract features from the training data set to obtain multimodal features; Step S4: Use the attention mechanism to perform weighted fusion on the multimodal features to obtain the fused multimodal features; Step S5: construct a bidirectional long short-term memory network BiLSTM model; Step S6: Use the fused multimodal features to train the BiLSTM model to obtain a trained BiLSTM model; Step S7: Collecting real-time data of the power equipment operation and inputting it into the trained BiLSTM model to perform fault diagnosis of the power equipment and output the fault prediction value of the power equipment; Step S8: Calculate the cross entropy loss function value according to the test data set and the fault prediction value of the power equipment, and evaluate the prediction accuracy of the trained BiLSTM model based on the cross entropy loss function value.

2. The method for diagnosing faults of power equipment based on multimodal data fusion according to claim 1, characterized in that: In step S2, the multimodal data of the operation of the power equipment is preprocessed to obtain the preprocessed multimodal data, which specifically includes the following sub-steps: Step S21: using Gaussian filtering to remove noise from the thermal infrared imaging image, and using wavelet transform to remove noise from the sound data and vibration data of the device respectively; Step S22: using linear interpolation to fill in missing values ​​in the denoised thermal infrared imaging image, the device sound data, and the device vibration data; Step S23: normalizing the padded thermal infrared imaging image, the device sound data, and the device vibration data to obtain normalized thermal infrared imaging image, the device sound data, and the device vibration data.

3. The method for diagnosing faults of electric power equipment based on multimodal data fusion according to claim 2, characterized in that: In step S3, the following sub-steps are specifically included: The convolutional neural network (CNN) algorithm is used to extract features from the normalized thermal infrared imaging image to obtain the temperature distribution features. The mathematical expression of the specific feature extraction is as follows: F image =CNN(X image ); Among them, F image Represents the temperature distribution characteristics, X image represents the normalized thermal infrared imaging image; Short-time Fourier transform (STFT) is used to extract features from the normalized device sound data to obtain time-frequency features. The mathematical expression for the specific feature extraction is as follows: Among them, F sound (τ, ω) represents the time-frequency characteristics, x sound (n) represents the normalized device sound data, w(n) represents the window function, τ represents time, and ω represents frequency; Fast Fourier transform (FFT) is used to extract features from the normalized vibration data of the equipment to obtain frequency domain features. The mathematical expression for the specific feature extraction is as follows: Among them, F vibration (k) represents the kth frequency domain feature, x vibration (n) represents the normalized vibration data of the device, and N represents the total number of normalized vibration data of the device.

4. The method for diagnosing faults of electric power equipment based on multimodal data fusion according to claim 3, characterized in that: In step S4, the following sub-steps are specifically included: Use the attention mechanism to transform the temperature distribution feature F image , time-frequency features F sound (τ, ω) and frequency domain features F vibration (k) Fusion is performed to obtain the fused multimodal feature F concat , where F concat The specific mathematical expression is as follows: F concat =[F image ,F sound (t,w),F vibration (k)]; The mathematical expression of the attention mechanism is as follows: Where Q represents the query matrix; K represents the key matrix; V represents the value matrix; d k represents the dimension of the key; softmax(x) represents the normalization function; Attention(x) represents the function of the attention mechanism.

5. The method for diagnosing faults of electric power equipment based on multimodal data fusion according to claim 4, characterized in that: In step S6, the following sub-steps are specifically included: Step S61: The fused multimodal features F concat Input the BiLSTM model to process the multi-layer network structure and obtain the corresponding hidden state h t , where h t The specific mathematical expression is as follows: h t =BiLSTM(F concat ); Step S62: h t Convert the result into a fault classification type to complete the training of the BiLSTM model. The specific conversion calculation formula is as follows: y =softmax(W·h t +b) Where y represents the fault classification type result, W represents the weight matrix, and b represents the bias term.

6. The method for diagnosing faults of electric power equipment based on multimodal data fusion according to claim 1, characterized in that: In step S8, the cross entropy loss function value is calculated based on the test data set and the fault prediction value of the power equipment. The specific calculation formula is as follows: Among them, L represents the cross entropy loss function value, y i represents the true value of the i-th sample, represents the predicted value of the BiLSTM model for the i-th sample, and N' represents the total number of samples in the test dataset.

7. The method for diagnosing faults of electric power equipment based on multimodal data fusion according to claim 1, characterized in that: The following steps are also included: The Adam optimizer is used to optimize the parameters of the BiLSTM model. The mathematical expression for BiLSTM model parameter optimization is as follows: Among them, θ t represents the parameters optimized in the tth round of the BiLSTM model, η represents the learning rate, m t represents the first-order momentum estimate, v t Represents the second-order momentum estimate, ∈ represents a positive real number, and its value range is 10 -8 ~10 -10 .

8. A power equipment fault diagnosis system based on multimodal data fusion, using the power equipment fault diagnosis method based on multimodal data fusion according to any one of claims 1 to 7, characterized in that: The system comprises: A first acquisition module is used to collect multimodal data of the operation of the power equipment, wherein the multimodal data of the operation of the power equipment includes thermal infrared imaging images, sound data of the equipment, and vibration data of the equipment; A data preprocessing module is used to preprocess the multimodal data of the power equipment operation to obtain preprocessed multimodal data; Data partitioning module, used to divide the preprocessed multimodal data into training data set and test data set; Feature extraction module, used to extract features from the training data set to obtain multimodal features; The feature fusion module is used to perform weighted fusion of multimodal features using the attention mechanism to obtain fused multimodal features; Building module for constructing BiLSTM model; The model training module is used to train the BiLSTM model using the fused multimodal features to obtain a trained BiLSTM model; The second acquisition module is used to collect real-time data of power equipment operation; The fault diagnosis module is used to input the real-time data of power equipment operation into the trained BiLSTM model to perform fault diagnosis on the power equipment and output the fault prediction value of the power equipment; A calculation module, used to calculate a cross entropy loss function value based on a test data set and a fault prediction value of the power equipment; The evaluation module is used to evaluate the prediction accuracy of the trained BiLSTM model based on the cross-entropy loss function value.

9. The power equipment fault diagnosis system based on multimodal data fusion according to claim 8, characterized in that: The data preprocessing module includes: A noise removal submodule is used to remove noise from thermal infrared imaging images using Gaussian filtering, and to remove noise from the device's sound data and vibration data using wavelet transform. A missing value filling submodule is used to fill the missing values ​​in the denoised thermal infrared imaging image, the device sound data, and the device vibration data using a linear interpolation method; A normalization processing submodule is used to perform normalization processing on the padded thermal infrared imaging image, the sound data of the device, and the vibration data of the device, respectively, to obtain normalized thermal infrared imaging image, the sound data of the device, and the vibration data of the device; The feature extraction module includes: The first feature extraction submodule is used to extract features from the normalized thermal infrared imaging image using a convolutional neural network (CNN) algorithm to obtain temperature distribution features. The specific mathematical expression for feature extraction is as follows: F image =CNN(X image ); Among them, F image Represents the temperature distribution characteristics, X image represents the normalized thermal infrared imaging image; The second feature extraction submodule is used to extract features from the normalized device sound data using short-time Fourier transform (STFT) to obtain time-frequency features. The mathematical expression for the specific feature extraction is as follows: Among them, F sound (τ, ω) represents the time-frequency characteristics, x sound (n) represents the normalized device sound data, w(n) represents the window function, τ represents time, and ω represents frequency; The third feature extraction submodule is used to extract features from the normalized vibration data of the device using the fast Fourier transform (FFT) to obtain frequency domain features. The specific mathematical expression of feature extraction is as follows: Among them, F vibration (k) represents the kth frequency domain feature, x vibration (n) represents the normalized vibration data of the device, and N represents the total number of normalized vibration data of the device; The feature fusion module includes: The feature fusion submodule is used to use the attention mechanism to integrate the temperature distribution feature F image , time-frequency features F sound (τ, ω) and frequency domain features F vibration (k) Fusion is performed to obtain the fused multimodal feature F concat , where F concat The specific mathematical expression is as follows: F concat =[F image ,F sound (t,w),F vibration (k)]; The mathematical expression of the attention mechanism is as follows: Where Q represents the query matrix; K represents the key matrix; V represents the value matrix; d k represents the dimension of the key; softmax(x) represents the normalization function; Attention(x) represents the function of the attention mechanism; The model training module includes: The multi-layer network structure processing submodule is used to transform the fused multimodal features F concat Input the BiLSTM model to process the multi-layer network structure and obtain the corresponding hidden state h t , where h t The specific mathematical expression is as follows: h t =BiLSTM(F concat ); The conversion submodule is used to convert h t Convert the result into a fault classification type to complete the training of the BiLSTM model. The specific conversion calculation formula is as follows: y =softmax(W·h t +b) Where y represents the fault classification type result, W represents the weight matrix, and b represents the bias term.

10. The power equipment fault diagnosis system based on multimodal data fusion according to claim 8, characterized in that: Also includes: The model parameter optimization module is used to optimize the parameters of the BiLSTM model using the Adam optimizer. The mathematical expression for BiLSTM model parameter optimization is as follows: Among them, θ t represents the parameters optimized in the tth round of the BiLSTM model, η represents the learning rate, m t represents the first-order momentum estimate, v t Represents the second-order momentum estimate, ∈ represents a positive real number, and its value range is 10 -8 ~10 -10 .

Citation Information

Patent Citations

  • Transformer fault diagnosis method based on Bi-LSTM and analysis of dissolved gas in oil

    CN110501585A

  • Aero-engine fault diagnosis method and system based on multi-modal deep learning

    CN116842423A

  • Transformer fault diagnosis method based on multi-mode self-attention mechanism

    CN117725529A

  • Data multi-scale fusion method for digital twinning of power transmission and transformation equipment

    CN119830198A

  • Transformer anomaly detection method based on multi-modal deep learning

    CN119903451A