Multi-modal data fusion acquisition system and method

The system uses Bi-GRU and GNN with cross-attention and WLS optimization to integrate transformer data modalities, addressing the challenge of complex relationships in multi-modal data fusion, thereby improving fault diagnosis accuracy and robustness.

CN120316699APending Publication Date: 2025-07-15CHINA SOUTHERN POWER GRID NEW POWER SYSTEM (BEIJING) RESEARCH INSTITUTE CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510286073.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing multimodal data fusion method has problems such as insufficient data relationship modeling, incomplete feature extraction, and insufficient correlation information between modals in transformer fault diagnosis, making it difficult to accurately capture complex fault modes.

Method used

A multimodal data fusion acquisition system is adopted, including data acquisition units, feature extraction units, graph neural network units and cross-attention mechanism units. Features are extracted through a bidirectional gating neural network, and the graph neural network models the non-Euclidean relationship between modes, and uses the cross-attention mechanism to establish deep connections, and finally outputs the fault category through the classification model.

Benefits of technology

It improves the accuracy and robustness of transformer fault diagnosis, can better handle the dependencies between modes, enhances the quality and noise robustness of feature fusion, and adapts to different operating states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120316699A_ABST
    Figure CN120316699A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal data fusion acquisition system and method, and the method comprises the steps: carrying out the feature extraction of the text information, frequency domain image and infrared image data of a vibration signal through a Bi-GRU network, and fusing the data of a plurality of modals into a comprehensive feature matrix; in the process, key features can be extracted from each mode, and the model can effectively capture the non-Euclidean relationship between different modes by constructing an adjacent matrix and performing graph convolution operation for a subsequent relationship. According to the method, a traditional matrix-based operation mode is broken, the dependency relationship between modes can be better processed, and the quality of feature fusion is improved. And feature fusion is further optimized by using least square weighted least square (WLS) fusion and a multi-modal factorization technology, so that the robustness of the system to noise and abnormal data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of data fusion, and specifically relates to a multi-modal data fusion acquisition system and method. Background Art

[0002] As a core device in the power system, the power transformer undertakes important power transmission and conversion tasks. During long-term operation, the transformer is affected by various factors, such as load fluctuations, temperature changes, mechanical vibrations, etc., and is prone to failures. Transformer failures not only lead to system outages and economic losses but also pose a threat to the safety and stability of the power system. Therefore, real-time monitoring and early diagnosis of transformer failures are particularly important.

[0003] Currently, transformer fault diagnosis methods mainly rely on various types of sensing data such as vibration signals, infrared images, and load data. These data have different characteristics and information, and single-modal data often cannot accurately reflect the operating state of the transformer. However, the multi-modal data fusion method can effectively integrate the advantages of each modality and improve the diagnosis accuracy.

[0004] Existing multi-modal data fusion methods mostly focus on the extraction and processing of single-modal data. However, in practical applications, the fault states of transformers are often complex combinations of multiple factors and multi-modalities. How to effectively fuse these different modal data and extract the deep features and correlation information therein has become the key to improving the accuracy of transformer fault diagnosis. Traditional fusion methods often have problems such as insufficient data relationship modeling, incomplete feature extraction, and insufficient correlation information between modalities, making it difficult to accurately capture complex fault patterns.

[0005] In view of this, the present invention is specifically proposed. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a multi-modal data fusion acquisition system and method to solve the problems mentioned in the above background art.

[0007] The basic concept of the technical solution adopted by the present invention to solve the above technical problem is as follows:

[0008] A multi-modal data fusion acquisition system includes: a data acquisition unit for collecting multi-modal data of the transformer, including vibration signals, infrared image data, and load data;

[0009] A feature extraction unit that uses a bidirectional gated neural network to extract features from the text information of the vibration signal, the frequency domain diagram of the vibration signal, and the infrared image of the transformer to generate a modal feature matrix X L ;

[0010] The graph neural network unit constructs the adjacency matrix A based on the complex relationships between modalities and uses the modality feature matrix X L as the input, models the non-Euclidean relationships between modalities through the graph neural network, and generates the fused modality feature matrix H L .

[0011] The cross-attention mechanism unit applies the cross-attention mechanism to the fused modality feature matrix H L to establish deep connections between modalities, extract more relevant features, and obtain the final multi-modal fusion feature F T ;

[0012] The fault diagnosis unit uses the generated multi-modal fusion feature for fault status identification and finally outputs the fault category through the classification model.

[0013] Optionally, the steps of using the bidirectional gated neural network to extract features from the text information of the vibration signal, the frequency-domain graph of the vibration signal, and the infrared image of the transformer to generate the modality feature matrix are as follows:

[0014] S1: Use the bidirectional gated neural network to process the text information of the vibration signal to obtain the vibration signal text feature matrix where L is the length of the extracted feature vector, represents the extracted feature at each time step;

[0015] S2: Process the frequency-domain image of the vibration signal, use the bidirectional gated neural network to extract the high-frequency component features therein, and obtain the frequency-domain image feature matrix where LLL is the length of the extracted feature vector, represents the frequency-domain feature at each time step;

[0016] S3: Use the bidirectional gated neural network to process the infrared image of the transformer, extract the key features in the infrared image, and obtain the infrared image feature matrix where L is the length of the extracted feature vector, represents the image feature at each time step;

[0017] S4: Fuse the vibration signal text feature matrix P T , the frequency-domain image feature matrix Q T and the infrared image feature matrix O T extracted in S1-S3 to obtain the modality feature matrix X L , XL = Concat(P T , Q T , O T ).

[0018] Optionally, construct an adjacency matrix A based on the complex relationships between modalities, and use the modality feature matrix X L as input, model the non-Euclidean relationships between modalities through a graph neural network to generate a fused modality feature matrix H L The steps are as follows:

[0019] A1: Perform relationship modeling on the modality feature matrices P T , Q T and O T to construct an adjacency matrix A. Take the modality feature matrix X L as input, and combine it with the adjacency matrix A for processing through a graph neural network

[0020] A2: Utilize the input feature matrix X L and the adjacency matrix A to calculate the aggregation and propagation of modality features through a graph convolutional layer to obtain a fused modality feature matrix H L , and its expression is: H L =σ(W h ·(A·X L ))), where A is the adjacency matrix representing the relationships between modalities; X L is the input modality feature matrix that undergoes information propagation through the graph neural network; W h is the weight matrix of the graph neural network; σ is an activation function such as ReLU or Sigmoid; H L is the feature matrix processed by the graph neural network, which contains deep-level associations and fusion information between modalities.

[0021] Optionally, the steps to apply a cross-attention mechanism on the fused modality feature matrix H L to establish deep-level connections between modalities are as follows:

[0022] Take the fused modality feature matrix H L generated by the graph neural network as the input of the cross-attention mechanism. Then, apply the cross-attention mechanism to calculate the mutual influence and correlation between different modalities. The cross-attention mechanism establishes connections between modalities through the calculation of queries, keys, and values, thereby enhancing the association with other modalities and extracting the most relevant features. Its expression is: where Q V is the query matrix, K L is the key matrix representing the interaction relationships between different modalities, d k is the variance coefficient representing the scale of attention calculation, V L is the modality feature matrix; Y V is the vibration signal modality feature obtained through the cross-attention mechanism

[0023] Through the formula Adding multiple levels of cross-attention calculations to further capture deeper associations between modalities, thereby improving the quality of feature fusion, where l represents the number of attention layers. and represent the query and key of the l-th layer respectively. is the output of the previous layer. represents the output feature of the vibration signal modal feature after passing through the L-th layer of cross-attention mechanism.

[0024] Optionally, the step of obtaining the final multi-modal fusion feature F T is as follows:

[0025] In the case of no labeled data, self-supervised learning is used to further enhance the effectiveness of the features, and its expression is: where and represent the feature representations generated by self-supervised learning.

[0026] Using the cross-attention mechanism and self-supervised learning to optimize multiple modal features to obtain the fused important feature vector, where the calculation formula for the fusion process is where α i is the weight dynamically learned through reinforcement learning or gating mechanism, and V i are the features of different modalities.

[0027] Finally, the weighted multi-modal features are further fused through the fusion layer to obtain the final multi-modal fusion feature F T , and after obtaining the final multi-modal fusion feature F T , the least squares WLS fusion is applied to optimize the multi-modal fusion feature.

[0028] Optionally, the steps of applying the least squares WLS fusion to optimize these features after obtaining the final multi-modal fusion feature F T are as follows:

[0029] Construct a weighted matrix W, which is used to represent the weights that should be assigned to each modality during the fusion process. The weight matrix depends on the quality, reliability, or signal-to-noise ratio of each modality. The weighted matrix W can be dynamically adjusted according to the signal quality or importance of each modality.

[0030] After the establishment of the weighted matrix W is completed, through the least squares WLS optimization, the feature matrix F T of each modality is weighted and optimized. First, the optimization objective function is defined as the goal of minimizing the weighted error, and its optimization objective function is: where: F T x i is the i-th modal feature vector. Predicted modal eigenvector; W i is the weighting coefficient for this mode; |·|2 represents the L2 norm, which is used to measure the error;

[0031] According to the weighting coefficient W of each mode i , the features of different modes are weighted to obtain an optimized feature matrix: F T,opt = W·F T , where F T,opt is the fused feature matrix after weighted optimization. Then, multiple weighted and optimized eigenvectors are concatenated to form the final eigenvector.

[0032] Optionally, use historical data to build a convolutional neural network and train a convolutional neural network. Then, input the final eigenvector into the convolutional neural network for convolutional processing. Next, in the convolutional layer, perform a convolutional operation on the input multi-modal fused features to extract feature maps. Convolutional operations help automatically learn important spatial features from the data. After the convolutional operation, a max-pooling layer is used to reduce the weight parameters of the model and effectively reduce the computational complexity. The stride of the max-pooling layer is set to 2.

[0033] Input the pooled data into the fully connected layer, which is used to further integrate features and learn higher-level abstract representations. The output of the fully connected layer contains the prediction information for each fault state.

[0034] After the fully connected layer, connect a Softmax regression layer to output the probability distribution of different faults of the transformer. Its expression is: Fault = Softmax(V), where V is the fused feature output by the fully connected layer, representing the probability distribution of the transformer fault state.

[0035] Use the cross-entropy loss function to calculate the loss of the network, and optimize the parameters of the network according to this loss function. Then, use the Adam optimization algorithm to iterate the gradient and update the hyperparameters of the network to ensure that the network is gradually optimized during training and the error is reduced.

[0036] Set the batch size to 64, the maximum number of iterations to 150, and the initial learning rate of the learning rate to 0.01. To prevent underfitting and improve the generalization ability of the model, a mechanism for dynamically adjusting the learning rate is introduced during training. At iteration numbers 400 and 800, the learning rate is adjusted to 0.001 and 0.0001 respectively to further optimize the training process and accelerate convergence.

[0037] A multi-modal data fusion acquisition method includes the following steps:

[0038] Collect multi-modal data of the transformer, including vibration signal, infrared image data, and load data;

[0039] Use a bidirectional gated neural network to extract features from the text information of vibration signals, the frequency-domain diagrams of vibration signals, and the infrared images of transformers, and generate a modal feature matrix X L ;

[0040] Construct an adjacency matrix A based on the complex relationships between modalities, and use the modal feature matrix X L as the input, and model the non-Euclidean relationships between modalities through a graph neural network to generate a fused modal feature matrix H L ;

[0041] Apply a cross-attention mechanism to the fused modal feature matrix H L to establish deep connections between modalities, extract more relevant features, and obtain the final multi-modal fusion feature F T ;

[0042] Use the generated multi-modal fusion feature for fault state recognition, and finally output the fault category through a classification model.

[0043] After adopting the above technical solutions, the present invention has the following beneficial effects compared with the prior art. Of course, any product implementing the present invention does not necessarily need to achieve all the advantages described below:

[0044] Extract features from the text information, frequency-domain images, and infrared image data of vibration signals through a Bi-GRU network, and fuse the data of multiple modalities into a comprehensive feature matrix. This process can extract key features from each modality, providing more abundant information for subsequent fault diagnosis. In addition, by using a graph neural network to construct complex relationships between modalities, through constructing an adjacency matrix and performing graph convolution operations, the model can effectively capture the non-Euclidean relationships between different modalities. This method breaks the traditional matrix-based operation mode, can better handle the dependence relationships between modalities, and improves the quality of feature fusion. Moreover, the least squares weighted least squares (WLS) fusion and multi-modal factorization techniques are used to further optimize the feature fusion, improving the robustness of the system to noise and abnormal data.

[0045] The following further describes the specific implementation manners of the present invention in detail with reference to the accompanying drawings. Description of the Drawings

[0046] The following drawings in the description are only some embodiments. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the attached

[0047] figures:

[0048] Figure 1It is a block diagram of a multi-modal data fusion acquisition system;

[0049] Figure 2 It is a schematic flow diagram of a multi-modal data fusion acquisition method.

[0050] It should be noted that these drawings and textual descriptions are not intended to limit the scope of the concept of the present invention in any way, but to illustrate the concept of the present invention to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0051] Now, the present invention will be further described in detail with reference to the accompanying drawings.

[0052] Please refer to Figure 1-2 As shown, in this embodiment, a multi-modal data fusion acquisition system is provided, including a data acquisition unit for collecting multi-modal data of a transformer, including vibration signals, infrared image data, and load data;

[0053] A feature extraction unit that uses a bidirectional gated neural network to extract features from the text information of vibration signals, the frequency-domain graph of vibration signals, and the infrared image of the transformer, generating a modal feature matrix X L ;

[0054] A graph neural network unit that constructs an adjacency matrix A based on the complex relationships between modalities and uses the modal feature matrix X L as input, and models the non-Euclidean relationships between modalities through a graph neural network to generate a fused modal feature matrix H L .

[0055] A cross-attention mechanism unit that applies a cross-attention mechanism to the fused modal feature matrix H L to establish deep connections between modalities, extract more relevant features, and obtain the final multi-modal fusion feature F T ;

[0056] A fault diagnosis unit that uses the generated multi-modal fusion feature to identify the fault state, and finally outputs the fault category through a classification model.

[0057] By using a bidirectional gated neural network (Bi-GRU), a graph neural network (GNN), and a cross-attention mechanism, the present invention realizes efficient fault diagnosis, can make full use of various modal data of the transformer, and thus accurately identify and predict the fault state of the transformer. This method not only improves the accuracy of fault diagnosis, but also can adapt to various different operating states of the transformer, and has strong generalization ability.

[0058] A multimodal data fusion acquisition system and method according to claim 1, characterized in that the steps of using a bidirectional gated neural network to extract features from the text information of the vibration signal, the frequency domain diagram of the vibration signal, and the infrared image of the transformer to generate a modal feature matrix are as follows:

[0059] S1: Use a bidirectional gated neural network to process the text information of the vibration signal to obtain a vibration signal text feature matrix where L is the length of the extracted feature vector, represents the extracted feature at each time step;

[0060] S2: Process the frequency domain image of the vibration signal, use a bidirectional gated neural network to extract the high-frequency component features therein, and obtain a frequency domain image feature matrix where LLL is the length of the extracted feature vector, represents the frequency domain feature at each time step;

[0061] S3: Use a bidirectional gated neural network to process the infrared image of the transformer, extract the key features in the infrared image, and obtain an infrared image feature matrix where L is the length of the extracted feature vector, represents the image feature at each time step;

[0062] S4: Combine the vibration signal text feature matrix P T , the frequency domain image feature matrix Q T , and the infrared image feature matrix O T extracted in S1 - S3, and fuse them to obtain a modal feature matrix X L , XL = Concat(P T , Q T , O T ).

[0063] A multimodal data fusion acquisition system and method according to claim 1, characterized in that an adjacency matrix A is constructed based on the complex relationships between modalities, and the modal feature matrix X L is used as the input, and the non-Euclidean relationships between modalities are modeled through a graph neural network to generate a fused modal feature matrix H L The steps are as follows:

[0064] A1: Perform relationship modeling on the modal feature matrices P T , Q T , and O T to construct an adjacency matrix A. Using the modal feature matrix X L as the input, combined with the adjacency matrix A, it is processed through a graph neural network

[0065] A2: Use the input feature matrix XL Given the adjacency matrix A, the aggregation and propagation of modal features are calculated through the graph convolutional layer to obtain the fused modal feature matrix H L , and its expression is: H L =σ(W h ·(A·X L ))), where A is the adjacency matrix representing the relationship between modalities; X L is the input modal feature matrix that propagates information through the graph neural network; W h is the weight matrix of the graph neural network; σ is the activation function, such as ReLU or Sigmoid; H L is the feature matrix processed by the graph neural network, containing the deep-level association and fusion information between modalities.

[0066] For a multi-modal data fusion acquisition system and method according to claim 1, it is characterized in that, applying the cross-attention mechanism to the fused modal feature matrix H L , the steps to establish deep-level connections between modalities are as follows:

[0067] Taking the fused modal feature matrix H L generated by the graph neural network as the input of the cross-attention mechanism. Then, the cross-attention mechanism is applied to calculate the mutual influence and correlation between different modalities. The cross-attention mechanism establishes connections between modalities through the calculation of queries, keys, and values, thereby enhancing the association with other modalities and extracting the most relevant features. Its expression is: where Q V is the query matrix, K L is the key matrix representing the interaction relationship between different modalities, d k is the variance coefficient representing the scale of attention calculation, V L is the modal feature matrix; Y V is the vibration signal modal feature obtained through the cross-attention mechanism

[0068] By formula adding multiple levels of cross-attention calculations to further capture deeper-level associations between modalities, thereby improving the quality of feature fusion, where l represents the number of attention layers, and respectively represent the query and key of the l-th layer, is the output of the previous layer; represents the output feature of the vibration signal modal feature after passing through the L-th layer of the cross-attention mechanism.

[0069] For a multi-modal data fusion acquisition system and method according to claim 1, it is characterized in that, the steps to obtain the final multi-modal fusion feature F T are as follows:

[0070] In the absence of labeled data, self-supervised learning is used to further enhance the effectiveness of features, and its expression is: where and represent the feature representations generated by self-supervised learning;

[0071] The cross-attention mechanism and self-supervised learning are used to optimize multiple modal features to obtain an important fused feature vector. Among them, the calculation formula for the fusion process is where α i is the weight dynamically learned through reinforcement learning or gating mechanism, and V i are the features of different modalities;

[0072] Finally, the weighted multi-modal features are further fused through a fusion layer to obtain the final multi-modal fusion feature F T , and after obtaining the final multi-modal fusion feature F T the least squares WLS fusion is applied to optimize the multi-modal fusion feature.

[0073] According to the multi-modal data fusion acquisition system and method described in claim 1, it is characterized in that, after obtaining the final multi-modal fusion feature F T the steps of optimizing these features by applying the least squares WLS fusion are:

[0074] Construct a weighted matrix W, which is used to represent the weights that should be assigned to each modality during the fusion process. The weight matrix depends on the quality, reliability, or signal-to-noise ratio of each modality. The weighted matrix W can be dynamically adjusted according to the signal quality or importance of each modality.

[0075] After the establishment of the weighted matrix W is completed, through the least squares WLS optimization, the feature matrix F T of each modality is weighted and optimized. First, the optimization objective function is defined as the goal of minimizing the weighted error, and its optimization objective function is: where: F T,i is the i-th modality feature vector; the predicted modality feature vector; W i is the weighted coefficient of this modality; |·|2 represents the L2 norm, which is used to measure the error;

[0076] According to the weighted coefficient W i of each modality, the features of different modalities are weighted to obtain the optimized feature matrix: F T,opt = W·F T where, F T,optis the fused feature matrix after weighted optimization. Then, multiple weighted and optimized feature vectors are concatenated to form the final feature vector.

[0077] After obtaining the final multi-modal fusion features, the present invention applies the least squares weighted least squares (WLS) fusion to optimize these features, which can effectively improve the accuracy and robustness of feature fusion. The least squares WLS fusion method dynamically adjusts the influence of features by assigning different weighted coefficients to the features of each modality according to the signal-to-noise ratio, reliability or quality of each modality. This weighted optimization can minimize errors, suppress the negative impact of low-quality modality features, and at the same time enhance the weight of high-quality features in the final decision. In this way, the WLS optimization not only improves the performance of multi-modal data fusion, but also enhances the robustness of the model to noise and abnormal data, thereby enhancing the accuracy and stability of the overall fault diagnosis system.

[0078] According to the multi-modal data fusion acquisition system and method described in claim 1, it is characterized in that the steps of using the generated multi-modal fusion features for fault state recognition and finally outputting the fault category through a classification model are as follows:

[0079] The final feature vector is input into a convolutional neural network for convolutional processing. Then, in the convolutional layer, convolutional operations are performed on the input multi-modal fusion features to extract feature maps. The convolutional operation helps to automatically learn important spatial features from the data. After the convolutional operation, a max pooling layer is used to reduce the weight parameters of the model and effectively reduce the computational complexity. The stride of the max pooling layer is set to 2.

[0080] The pooled data is input into a fully connected layer, which is used to further integrate features and learn higher-level abstract representations. The output of the fully connected layer contains the prediction information of each fault state.

[0081] After the fully connected layer, a Softmax regression layer is connected to output the probability distribution of different faults of the transformer. Its expression is: Fault = Soffmax(V), where V is the fused feature output by the fully connected layer, representing the probability distribution of the transformer fault state.

[0082] The cross-entropy loss function is used to calculate the loss of the network, and the parameters of the network are optimized according to this loss function. Then, the Adam optimization algorithm is used to iterate the gradient and update the hyperparameters of the network to ensure that the network is gradually optimized during training and the error is reduced.

[0083] Set the batch size to 64, the maximum number of iterations to 150, and the initial learning rate to 0.01. To prevent underfitting and improve the generalization ability of the model, a mechanism for dynamically adjusting the learning rate is introduced during training. When the number of iterations reaches 400 and 800, the learning rate is adjusted to 0.001 and 0.0001 respectively, further optimizing the training process and accelerating convergence.

[0084] A multi-modal data fusion acquisition method includes the following steps:

[0085] Collect multi-modal data of the transformer, including vibration signals, infrared image data, and load data;

[0086] Use a bidirectional gated neural network to extract features from the text information of the vibration signal, the frequency-domain graph of the vibration signal, and the infrared image of the transformer, generating a modal feature matrix X L ;

[0087] Construct an adjacency matrix A based on the complex relationships between modalities, and use the modal feature matrix X L as input, and model the non-Euclidean relationships between modalities through a graph neural network to generate a fused modal feature matrix H L ;

[0088] Apply a cross-attention mechanism to the fused modal feature matrix H L to establish deep connections between modalities, extract more relevant features, and obtain the final multi-modal fusion feature F T ;

[0089] Use the generated multi-modal fusion feature for fault state identification, and finally output the fault category through a classification model.

[0090] In the experiment, the multi-modal fusion (MIF) network was trained. The input of the multi-modal fusion (MIF) network includes multiple modal data such as the text information of the vibration signal, the frequency-domain image, and the infrared image. The core features of the MIF network:

[0091] Multi-modal data input: The input of the MIF network includes multiple modal data such as the text information of the vibration signal, the frequency-domain image, and the infrared image. Each modality provides a different perspective on the operating state of the transformer. The goal of the network is to combine this information and learn the deep associations between different modalities.

[0092] Cross-attention mechanism: The MIF network improves the interaction effect between modalities through the cross-attention mechanism. The cross-attention mechanism allows the model to allocate weights between different modalities, emphasizing the relationships between features in each modality, thereby enhancing the effect of multimodal data fusion. This mechanism can effectively assign different attention values to each modality, capture key modality features, and strengthen their associations.

[0093] Graph Neural Network (GNN): This network uses graph neural networks to model non-Euclidean relationships between modalities. Graph neural networks perform excellently in processing data with complex relationships and can capture the structural dependencies between modalities. By constructing an adjacency matrix, GNN can efficiently aggregate and propagate modality features, thus enhancing the fused feature representation.

[0094] Training of the deep learning model: The MIF network uses cross-validation and a hybrid loss function during training to optimize the model, ensuring that it can effectively learn the diagnostic rules of transformer faults. Combining the features of multiple modality data helps the model better understand the redundant information and associations in multimodal data, thereby improving the accuracy of fault classification.

[0095] Working principle of the MIF network:

[0096] Feature extraction: First, the MIF network extracts features from each modality data. For example, a bidirectional gated recurrent unit (Bi-GRU) is used to extract features from the text information and frequency-domain images of vibration signals.

[0097] Modality fusion: Through the cross-attention mechanism, the network fuses the features of these modalities. The cross-attention mechanism helps the network learn the correlations between different modalities, thereby enhancing the fused feature representation.

[0098] Fault diagnosis: Finally, the fused features are input into a convolutional neural network (CNN) and a fully connected layer (FC). After learning, the fault status prediction of the transformer is output.

[0099] Existing technical solutions usually use traditional neural network models (such as fully connected networks, convolutional neural networks (CNN)) to process single-modal data (such as vibration signals or image data). These methods face great challenges in multimodal data fusion. Especially when there are complex relationships between modalities, traditional methods are prone to losing the deep association information between modalities. Therefore, existing technical solutions often cannot effectively extract important features from multimodal data and have limited performance in fault diagnosis.

[0100] The present invention introduces a Bidirectional Gated Recurrent Unit (Bi-GRU) and a Cross-Attention mechanism, and combines a Graph Neural Network (GNN) to model the complex relationships between modalities. It innovatively achieves multi-modal data fusion through the following steps:

[0101] Bi-GRU (Bidirectional Gated Recurrent Unit): This method is used to extract features from the text information, frequency-domain images, and infrared images of vibration signals. Bi-GRU can capture temporal dependencies and is particularly suitable for processing time-series data and irregular modal features.

[0102] Cross-Attention mechanism: Through the cross-attention mechanism, the model can establish deep connections between modalities, significantly improving the ability of feature fusion. The attention mechanism helps the network dynamically focus on the most relevant features between different modalities, thereby enhancing the effectiveness of the fused features.

[0103] Graph Neural Network (GNN): In the process of modeling non-Euclidean relationships between modalities, GNN captures the complex relationships between modalities by constructing an adjacency matrix, further optimizing the multi-modal feature fusion process. Through graph convolutional layers to propagate and aggregate modal features, GNN can improve the complementarity of information between modalities. The present invention integrates the feature extraction, fusion, and relationship modeling of multi-modal data into a unified framework. Using Bi-GRU, the cross-attention mechanism, and graph neural networks effectively overcomes the complexity problems between modalities. Secondly, by optimizing the feature extraction and fusion process through Least Squares Weighted Least Squares (WLS) fusion, the accuracy and robustness of the model are improved.

[0104] This study uses TensorFlow 2.0 to build and run the multi-modal data fusion network. The experimental hardware platform includes an NVIDIA GeForce GTX 3060 GPU and an Intel i7-11800H CPU, ensuring efficient computing power. All comparison algorithms are run on the same experimental platform to ensure the fairness of the results.

[0105] To evaluate the performance of different algorithms, a 10-fold cross-validation method is used to train each comparison algorithm, and 80,000 groups of data (including vibration signal text information, frequency-domain graphs, and infrared images) are selected as the training set. The dataset is randomly divided into a 75% training set and a 25% test set, and data under other current and voltage conditions are used as the validation set.

[0106] To collect multimodal data of the transformer under different working conditions, a PXI-4461 device and an IEPE piezoelectric acceleration sensor (model CJ-YD500 / N) were used to collect vibration signals, which were installed on the high-voltage side, phase B. Loose faults of the longitudinal tie rod of the winding and the transverse tie rod of the iron core were simulated, and data were collected under different currents (60%, 80%, 100%, 120% IN) and voltages (60%, 80%, 100%, 120% UN).

[0107] Each data set includes 10,000 groups of samples, with a total of 80,000 groups of data, and the sampling frequency is 20 kHz. The training set is trained using data of rated current and voltage, and the validation set uses data of other current and voltage levels.

[0108] The evaluation metrics include Accuracy (Ac), Precision (Pc), F1score (F1), and Recall (R). These metrics are calculated by the following formulas respectively:

[0109]

[0110]

[0111] Among them, TN is the number of negative classes, FN is the number of false negatives, TP is the number of true positives, and FP is the number of false positives.

[0112] Algorithm comparison: The following algorithms were selected for the comparative experiment:

[0113] BGRU (Bidirectional Gated Recurrent Neural Network): Used for feature extraction, but without using the attention mechanism and graph neural network.

[0114] Deep Residual Network (DRSN): A traditional feature extraction method based on deep neural networks.

[0115] Squeeze-and-Excitation Network (SK-net): A lightweight neural network for image processing.

[0116] Training method: All comparative algorithms were trained using the 10-fold cross-validation method. The training set and the test set account for 75% and 25% respectively, and the current and voltage levels in the data set are evenly distributed.

[0117] Hyperparameter settings: The hyperparameter settings of all algorithms are the same. The batch size is 64, the maximum number of iterations is 150, and the initial learning rate is 0.01. To prevent overfitting, the learning rate is dynamically adjusted every 100 iterations (adjusted to 0.001 and 0.0001).

[0118] To evaluate the effectiveness of the present invention and other comparative algorithms, evaluation metrics such as accuracy, precision, recall, and F1-score were used.

[0119] (1) Experimental result display table:

[0120]

[0121]

[0122] As can be seen from the table, the MIF algorithm outperforms other comparative algorithms in terms of accuracy, precision, recall, and F1-score.

[0123] The experimental results show that the MIF (Multi-modal Information Fusion) network performs better than traditional neural network algorithms such as BGRU, Deep Residual Network (DRSN), and Squeeze-and-Excitation Network (SK-net) in the task of transformer fault diagnosis. The MIF algorithm can effectively process multi-modal data, extract deep-level feature relationships between modalities, and significantly improve the accuracy, generalization ability, and robustness of fault diagnosis through technologies such as cross-attention mechanism, graph neural network, and least squares WLS fusion.

[0124] The present invention is not limited to the above embodiments. Anyone should know that structural changes made under the inspiration of the present invention, as long as they have the same or similar technical solutions as the present invention, fall within the protection scope of the present invention. The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies.

Claims

1. A multi-modal data fusion acquisition system and method, characterized in that, Including: A data acquisition unit for collecting multimodal data of the transformer, including vibration signals, infrared image data, and load data; The feature extraction unit uses a bidirectional gated neural network to extract features from the text information of the vibration signal, the frequency-domain diagram of the vibration signal, and the infrared image of the transformer, and generates a modal feature matrix X L ; Graph neural network unit, constructs an adjacency matrix A based on the complex relationships between modalities, and uses the modality feature matrix X L as inputs, models the non-Euclidean relationships between modalities through a graph neural network, and generates a fused modality feature matrix H L ; Cross-attention mechanism unit, applying the cross-attention mechanism to the fused modal feature matrix H L to establish deep connections between modalities, extract more relevant features, and obtain the final multi-modal fusion feature F T ; A fault diagnosis unit that uses the generated multimodal fusion features to identify the fault status and finally outputs the fault category through a classification model.

2. The multimodal data fusion acquisition system and method according to claim 1, characterized in that, The steps of using a bidirectional gated neural network to extract features from the text information of the vibration signal, the frequency domain diagram of the vibration signal, and the infrared image of the transformer to generate a modal feature matrix are as follows: S1: Process the text information of the vibration signal using a bidirectional gated neural network to obtain a vibration signal text feature matrix where L is the length of the extracted feature vector representing the extracted features at each time step S2: Process the frequency-domain image of the vibration signal, and use a bidirectional gated neural network to extract the high-frequency component features therein to obtain a frequency-domain image feature matrix where LLL is the length of the extracted feature vector, representing the frequency-domain features at each time step; S3: Process the infrared image of the transformer using a bidirectional gated neural network, extract the key features in the infrared image, and obtain an infrared image feature matrix where L is the length of the extracted feature vector, represents the image features at each time step; S4: The vibration signal text feature matrix P extracted in S1 - S3 T , the frequency domain image feature matrix Q T and the infrared image feature matrix O T are fused to obtain the modal feature matrix X L , XL = Concat(P T , Q T , O T ).

3. A multimodal data fusion acquisition system and method according to claim 1, characterized in that Construct an adjacency matrix A based on the complex relationships between modalities, and use the modality feature matrix X L as input, model the non-Euclidean relationships between modalities through a graph neural network, and generate a fused modality feature matrix H L The steps are as follows: A1: For the modal feature matrices P T , Q T and O T perform relationship modeling to construct the adjacency matrix A, take the modal feature matrix X L as the input, combine it with the adjacency matrix A, and process it through a graph neural network A2: Utilize the input feature matrix X L and the adjacency matrix A to calculate the aggregation and propagation of modal features through the graph convolutional layer, obtaining the fused modal feature matrix H L , and its expression is: H L =σ(W h ·(A·X L ))), where A is the adjacency matrix representing the relationship between modalities; X L is the input modal feature matrix that undergoes information propagation through the graph neural network; W h is the weight matrix of the graph neural network; σ is the activation function, such as ReLU or Sigmoid; H L is the feature matrix processed by the graph neural network, containing the deep - level associations and fusion information between modalities.

4. A multimodal data fusion acquisition system and method according to claim 1, characterized in that, On the fused modal feature matrix H L Applying the cross-attention mechanism to establish deep connections between modalities involves the following steps: Take the fused modal feature matrix H generated by the graph neural network L as the input of the cross-attention mechanism. Then, apply the cross-attention mechanism to calculate the mutual influence and correlation between different modalities. The cross-attention mechanism establishes the connection between modalities through the calculation of queries, keys, and values, thereby enhancing the association with other modalities and extracting the most relevant features. Its expression is: where Q V is the query matrix, K L is the key matrix, representing the interaction relationship between different modalities, d k is the variance coefficient, representing the scale of attention calculation, V L is the modal feature matrix; Y V is the modal feature of the vibration signal obtained through the cross-attention mechanism Through the formula Add multiple levels of cross-attention calculation to further capture deeper associations between modalities, thereby improving the quality of feature fusion. Here, l represents the number of layers of attention, and respectively represent the query and key of the l-th layer, is the output of the previous layer; represents the output feature of the vibration signal modal feature after passing through the L-th layer of cross-attention mechanism.

5. A multimodal data fusion acquisition system and method according to claim 1, characterized in that Obtain the final multi-modal fusion feature F T The steps are as follows: In the absence of labeled data, self-supervised learning is adopted to further enhance the effectiveness of features, and its expression is as follows: Among them, and represent the feature representations generated by self-supervised learning; Optimize multiple modal features using cross-attention mechanism and self-supervised learning to obtain an important feature vector after fusion. The calculation formula for the fusion process is where α i is the weight dynamically learned through reinforcement learning or gating mechanism, and V i is the feature of different modalities; Finally, the weighted multi-modal features are further fused through a fusion layer to obtain the final multi-modal fusion feature F T , after obtaining the final multi-modal fusion feature F T the least squares WLS fusion is applied to optimize the multi-modal fusion feature.

6. The multimodal data fusion acquisition system and method according to claim 1, wherein After obtaining the final multi-modal fusion feature F T The steps for optimizing these features by applying least squares WLS fusion are as follows: Construct a weighted matrix W, which is used to represent the weights that should be assigned to each modality during the fusion process. The weight matrix depends on the quality, reliability, or signal-to-noise ratio of each modality. The weighted matrix W can be dynamically adjusted according to the signal quality or importance of each modality; After the weighted matrix W is established, through the least squares WLS optimization, for the feature matrix F of each modality T perform weighted optimization. First, define the optimization objective function. The goal is to minimize the weighted error, and its optimization objective function is: where: F T,i is the feature vector of the i-th modality; the predicted modal feature vector; W i is the weighting coefficient of this modality; |·|2 represents the L2 norm, which is used to measure the error; According to the weighted coefficient W of each modality i , the features of different modalities are weighted to obtain the optimized feature matrix: F T,opt = W · F T , where F T,opt is the fused feature matrix after weighted optimization. Then, multiple weighted and optimized feature vectors are concatenated to form the final feature vector.

7. A multimodal data fusion acquisition system and method according to claim 1, wherein The steps of using the generated multimodal fusion features to identify the fault status and finally output the fault category through a classification model are as follows: Use historical data to establish and train a convolutional neural network. Then, input the final feature vector into the convolutional neural network for convolutional processing. Next, in the convolutional layer, perform a convolutional operation on the input multimodal fusion features to extract feature maps. The convolutional operation helps to automatically learn important spatial features from the data. After the convolutional operation, use a max pooling layer to reduce the weight parameters of the model and effectively reduce the computational complexity. The stride of the max pooling layer is set to 2. Input the pooled data into a fully connected layer, which is used to further integrate features and learn higher-level abstract representations. The output of the fully connected layer contains the prediction information for each fault status; After the fully connected layer, connect a Softmax regression layer to output the probability distribution of different faults of the transformer. Its expression is: Fault = Softmax(V), where V is the fusion feature output by the fully connected layer, representing the probability distribution of the transformer fault status; Use the cross-entropy loss function to calculate the loss of the network and optimize the parameters of the network according to this loss function. Then, use the Adam optimization algorithm to iterate the gradients and update the hyperparameters of the network to ensure that the network is gradually optimized during training and the error is reduced.

8. A multimodal data fusion acquisition method, characterized in that Including the following steps: Collect multimodal data of the transformer, including vibration signals, infrared image data, and load data; Use a bidirectional gated neural network to extract features from the text information of vibration signals, the frequency domain diagrams of vibration signals, and the infrared images of transformers, and generate a modal feature matrix X L ; Construct an adjacency matrix A based on the complex relationships between modalities, and use the modality feature matrix X L as the input, model the non-Euclidean relationships between modalities through a graph neural network, and generate a fused modality feature matrix H L ; Apply the cross-attention mechanism to the fused modal feature matrix H L to establish deep connections between modalities, extract features with stronger correlations, and obtain the final multi-modal fusion feature F T ; Use the generated multimodal fusion features to identify the fault status and finally output the fault category through a classification model.

Citation Information

Cited By

  • Method and system for calculating aerodynamic coefficient of heliostat by fusing geometric modal characteristics

    CN120911371A

  • A heliostat aerodynamic coefficient calculation method and system fusing geometric modal features

    CN120911371B

  • Multi-mode interaction system and method integrating multi-source data acquisition and intelligent triggering

    CN122261386A