Multi-modal data fusion GIS fault prediction method and system

Through multimodal data fusion, including infrared imaging, local discharge signals, vibration signals and gas component data, the self-attention and cross-attention mechanisms are used to extract and fusion characteristics, the problem that a single data source in traditional methods is difficult to capture multimodal associations, and higher GIS fault prediction accuracy and robustness are achieved.

CN120030464APending Publication Date: 2025-05-23山东泰开自动化有限公司
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202411993500.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Traditional GIS fault diagnosis methods rely mostly on a single data source, making it difficult to fully capture the complex relationship between multimodal data, resulting in unsatisfactory fault prediction accuracy.

Method used

Using the multimodal data fusion method, by acquiring and preprocessing the infrared imaging data, local discharge signals, mechanical vibration signals and gas component data of GIS devices, low-level features and high-level semantic features are extracted, and these features are fused through self-attention mechanisms and cross-attention mechanisms, and finally fault status prediction is performed through a fully connected network.

Benefits of technology

By deeply fusion of multimodal data, the model can retain both local details and global correlations, significantly improving the accuracy and robustness of GIS fault prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030464A_ABST
    Figure CN120030464A_ABST
Patent Text Reader

Abstract

The invention relates to a GIS fault prediction method and system based on multi-modal data fusion. The method comprises the following steps: acquiring temperature infrared imaging data, a partial discharge signal, a mechanical vibration signal and gas component data of GIS equipment, and preprocessing the temperature infrared imaging data, the partial discharge signal, the mechanical vibration signal and the gas component data; acquiring a low-level feature and a high-level semantic feature of each mode in the preprocessed data, extracting a spatial feature and a time sequence feature in the low-level feature and the high-level semantic feature, acquiring an association relationship between the spatial feature and the time sequence feature according to a self-attention mechanism, and fusing the spatial feature and the time sequence feature; and predicting the fault state of the GIS according to the obtained fusion features through a full-connection network to obtain a prediction result including fault possibility, fault position and fault type. According to the method, low-level features and high-level semantic features of each mode are extracted through a layered architecture, and local spatial features and global time sequence features are captured to ensure that the model can simultaneously keep local details and global association, so that the fault prediction accuracy for the GIS can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault prediction, and in particular to a GIS fault prediction method and system based on multi-modal data fusion. Background Art

[0002] The statements in this section merely provide background information related to the present invention and do not necessarily constitute prior art.

[0003] Gas insulated switchgear (GIS) is a key equipment in the power system, and its operating status is directly related to the safety and reliability of the power grid. Due to the complex structure inside GIS and the high-voltage operating environment, it is easily affected by faults such as partial discharge and mechanical vibration, which may lead to serious power outages. Therefore, fast and accurate prediction of GIS faults is crucial for the safe operation of the power system.

[0004] Traditional fault diagnosis methods mostly rely on a single data source (such as infrared imaging, vibration signals or partial discharge signals), which makes it difficult to fully capture the complex correlation between multimodal data, resulting in unsatisfactory fault prediction accuracy for GIS. Summary of the invention

[0005] In order to solve the technical problems existing in the above-mentioned background technology, the present invention provides a GIS fault prediction method and system of multimodal data fusion, which utilizes infrared imaging data, partial discharge signals and vibration signals to realize deep fusion of multimodal data and fault prediction.

[0006] In order to achieve the above object, the present invention adopts the following technical solution:

[0007] The first aspect of the present invention provides a GIS fault prediction method based on multimodal data fusion, comprising the following steps:

[0008] Acquire temperature infrared imaging data, partial discharge signals, mechanical vibration signals and gas composition data of GIS equipment and pre-process them;

[0009] Obtain low-level features and high-level semantic features of each modality in the preprocessed data, extract spatial features and temporal features, obtain the correlation between spatial features and temporal features based on the self-attention mechanism, and fuse the spatial features and temporal features;

[0010] The obtained fusion features are used through a fully connected network to predict the fault status of GIS and obtain prediction results including fault possibility, fault location and fault type.

[0011] Furthermore, preprocessing includes denoising, standardization and format conversion.

[0012] Furthermore, low-level features of each mode in the preprocessed data are obtained, including frequency domain features of partial discharge signals, time series features of vibration signals, and statistical features of gas components.

[0013] Furthermore, the spatial features and temporal features are extracted, and the correlation between the spatial features and the temporal features is obtained according to the self-attention mechanism and the spatial features and the temporal features are fused, including time alignment and spatial alignment of the obtained features, and mapping the low-level features of different modalities into a shared latent space, unifying the low-level features into the same dimension based on linear transformation, performing nonlinear mapping based on a multi-layer perceptron, and extracting the relationship between the modalities; the cross-modal attention mechanism is used to capture the interactive relationship between the modalities: the attention mechanism is used to capture the dependency between the modalities.

[0014] Furthermore, the modal representation is constructed as a graph structure, with each modal feature as a node and the similarity or dependency between nodes as the edges of the graph to capture global dependencies.

[0015] Furthermore, inter-modal features are fused through dynamic allocation of modal weights and sharing of multi-layer perceptrons.

[0016] Furthermore, high-level semantic features are specifically: utilizing the collaborative characteristics between modalities to extract global information, highlighting important modal features through attention mechanisms or weight adjustment, and generating high-level abstract features for classification or regression tasks.

[0017] The second aspect of the present invention provides a GIS fault prediction system for multimodal data fusion, including.

[0018] The multimodal data acquisition and preprocessing module is configured to: acquire and preprocess the temperature infrared imaging data, partial discharge signal, mechanical vibration signal and gas composition data of the GIS equipment;

[0019] The feature fusion and modeling module is configured to: obtain low-level features and high-level semantic features of each modality in the preprocessed data, extract spatial features and temporal features therein, obtain the correlation between spatial features and temporal features according to the self-attention mechanism, and fuse the spatial features and temporal features;

[0020] The fault prediction module is configured as follows: the obtained fusion features are used through a fully connected network to predict the fault state of the GIS and obtain a prediction result including the fault possibility, fault location and fault type.

[0021] A third aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above-mentioned multimodal data fusion GIS fault prediction method.

[0022] The fourth aspect of the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps in the above-mentioned multimodal data fusion GIS fault prediction method when executing the program.

[0023] Compared with the prior art, one or more of the above technical solutions have the following beneficial effects:

[0024] 1. Through the hierarchical architecture, the low-level features and high-level semantic features of each modality are extracted respectively. By capturing the local spatial features and the global temporal features, the model can retain both local details and global associations, thereby improving the accuracy of fault prediction for GIS.

[0025] 2. The low-level feature extraction stage can fully exploit modal characteristics. Traditional methods often rely on manually designed features or shallow models of a single modality, resulting in insufficient extraction of modal characteristics. This method can capture local patterns (such as local defect morphology in GIS) from image data and long-range dependencies from time series by extracting spatial features in low-level feature extraction. It can mine modal characteristics at a deeper level and reduce information loss. It can also adapt high-dimensional data (such as local image features) and time series data (such as partial discharge signals).

[0026] 3. The advantage of mid-level feature interaction is that it strengthens the capture of correlations between modalities. Traditional methods usually directly concatenate or simply weight modal features, which makes it difficult to fully capture the interactive information between modalities. This method uses a cross-attention mechanism in mid-level feature interactions to deeply explore the interdependence and complementarity between modalities. Use the self-attention mechanism within and between modalities to ensure local consistency within the modality and global consistency between modalities. By introducing graph neural networks, a global structural modeling of complex relationships between modalities is established. It can have a stronger capture capability for the complementarity of multimodal data, for example, by combining local discharge signals and infrared thermal imaging data to jointly reveal the causes of GIS failures. Avoid the problem of insufficient information fusion caused by independent processing of modal features in traditional methods.

[0027] 4. The advantage of high-level feature aggregation is the efficient integration of multimodal features. Traditional methods lack dynamic optimization for global features and rely only on simple classifiers. This method introduces a multi-head attention mechanism in high-level aggregation to adaptively assign importance weights of modal features to ensure the prominent role of key modal features. Use feature compression and dimensionality reduction techniques to effectively remove redundant information and improve model calculation efficiency and prediction accuracy. Combined with joint optimization objectives (such as mutual information, alignment loss, and classification / regression loss), ensure that the fused features have high consistency and discriminability. The importance of modal features is automatically learned by the model without manual intervention, reducing the problem of unreasonable weight distribution. Balance the diversity and consistency of modal features in the global feature representation to adapt to the multimodal data structure of complex GIS faults.

[0028] 5. The overall performance can improve the accuracy and robustness of fault prediction. Traditional methods are mostly based on single-modal features and are easily affected by fluctuations in the quality of a single data. This method uses the complementarity of multi-modal features (such as partial discharge signals detect electrical anomalies, but infrared thermal imaging captures thermal distribution anomalies, and the two work together to improve accuracy). The model has a stronger adaptability to different modes, and even if a certain mode of data is partially lost or noisy, it can rely on other modes to compensate. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0030] Figure 1 It is a schematic diagram of a GIS fault prediction process provided by one or more embodiments of the present invention. DETAILED DESCRIPTION

[0031] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0032] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0033] Most fault diagnosis methods for GIS rely on a single data source (such as infrared imaging, vibration signals or partial discharge signals), which makes it difficult to fully capture the complex correlation between multimodal data. The rapid development of deep learning technology, especially the widespread application of convolutional neural networks (CNN) and Transformer models, has provided new possibilities for the deep fusion of multimodal data and fault prediction. Therefore, the following embodiments propose a GIS fault prediction method and system for multimodal data fusion, which uses infrared imaging data, partial discharge signals and vibration signals to achieve deep fusion of multimodal data and fault prediction.

[0034] Embodiment 1:

[0035] like Figure 1 As shown, the GIS fault prediction method based on multimodal data fusion includes the following steps:

[0036] Data collection and preprocessing, specifically:

[0037] Obtain temperature distribution data of GIS through infrared thermal imaging equipment;

[0038] Collecting partial discharge signals through a partial discharge monitoring device;

[0039] Collect mechanical vibration signals through vibration sensors;

[0040] The above data are denoised, standardized and format converted.

[0041] Multimodal feature extraction, specifically:

[0042] For infrared imaging data, convolutional neural network (CNN) is used to extract spatial features;

[0043] For partial discharge and vibration signals, one-dimensional CNN is used to extract time series features.

[0044] Feature fusion and modeling, specifically:

[0045] The feature vector of multimodal data is input into the Transformer module, and the correlation between multimodal features is captured through its self-attention mechanism;

[0046] In the Transformer module, the multi-head attention mechanism is used to improve the model's ability to capture complex features.

[0047] Fault prediction, specifically:

[0048] The fused feature vector is input into the fully connected network for classification or regression to predict the fault status and fault type of GIS;

[0049] The output includes the prediction results of fault probability, fault location and fault type.

[0050] Model training and optimization, specifically:

[0051] Use historical failure data to train the model through supervised learning;

[0052] Adaptive learning rate optimization algorithm (AdamW) is used to optimize model parameters;

[0053] Cross-validation method was used to evaluate model performance and prevent overfitting.

[0054] This embodiment adopts the mode feature hierarchical processing method, and extracts low-level features (such as frequency domain features of partial discharge signals, time series features of vibration signals, statistical features of gas components) and high-level semantic features of each mode by designing a special hierarchical architecture. CNN is used to capture local spatial features, and Transformer processes global temporal features to ensure that the model can retain local details and global associations at the same time, thereby realizing fault prediction of multi-modal data fusion and improving the accuracy of fault prediction for GIS.

[0055] Low-level feature extraction is the foundation of multimodal fusion models. It aims to mine basic features unique to the modality from raw data and provide high-quality input for subsequent feature interaction and fusion. For partial discharge signals, vibration signals, gas decomposition features, etc., the underlying feature extraction process is: denoising, transformation, and feature pattern extraction.

[0056] 1.1 Partial discharge signals are non-stationary, periodic pulse signals with a wide frequency range, and their characteristics are easily masked by noise.

[0057] 1.1.2.1 Denoising:

[0058] Wavelet denoising: Use wavelet transform to separate the low-frequency components and high-frequency noise of the signal:

[0059]

[0060] Where a is the scale parameter, b is the time shift, and ψ* is the conjugate of the mother wavelet. The high-frequency wavelet coefficients are processed by thresholding:

[0061]

[0062] 1.1.2.2 Feature Transformation:

[0063] Frequency domain feature extraction: Extract spectrum energy features through fast Fourier transform (FFT):

[0064]

[0065] Frequency domain features include main frequency, spectrum center, spectrum energy, etc.

[0066] Time domain features extract statistical features such as mean, standard deviation, pulse width, and pulse amplitude.

[0067] 1.1.2.3 Convolutional feature extraction:

[0068] Use multi-scale convolution kernels to extract local feature patterns and input the transformed results into CNN:

[0069]

[0070] Among them, x[i,j] is the input feature map, k[m,n] is the convolution kernel, and y[i,j] is the output feature map.

[0071] The short-term changing characteristics in the signal are captured by sliding the convolution kernel (filter).

[0072] 1.2 Vibration signals are response signals of mechanical motion, which usually have timing and spectrum characteristics and are easily interfered by environmental noise.

[0073] 1.2.2.1 Denoising processing:

[0074] Low-pass filtering: remove high-frequency interference signals and retain the main vibration components:

[0075]

[0076] where h[k] is the impulse response of the filter.

[0077] 1.2.2.2 Feature Transformation:

[0078] A joint analysis method of time domain and frequency domain is introduced, and short-time Fourier transform (STFT) is combined with CNN to extract specific mechanical vibration features.

[0079] Short-time Fourier transform (STFT): Extract time-frequency distribution features:

[0080]

[0081] Where x(τ) is the input signal and w(τ-t) is the window function.

[0082] 1.2.2.3 Convolutional feature extraction:

[0083] Introduce recursive convolutional unit (RCU) to extract temporal dependency features:

[0084] h t =σ(W x x t +W h h t-1 +b)

[0085] Among them, h t is the hidden state at time step t, combined with the current input x t and the state h of the previous time step t-1 .

[0086] 1.3 Decomposition Gas Analysis (DGA) detects the concentration of gas components inside GIS (such as H 2 , CH 4 , C 2 H 2 ) reflects potential failures. The data is multi-dimensional and has time series characteristics.

[0087] 1.3.2.1 Data Dimensionality Reduction:

[0088] Use principal component analysis (PCA) to reduce dimensionality and remove redundant information:

[0089] Z=XWZ, where W is the principal component direction matrix and Z is the representation after dimensionality reduction.

[0090] 1.3.2.2 Trend extraction:

[0091] Calculate the trend characteristics (growth rate, fluctuation range) of each gas component.

[0092] growth rate:

[0093] Fluctuation range: R = max(x)-min(x)

[0094] 1.3.2.3 Convolutional feature extraction:

[0095] Use one-dimensional convolution to extract time series features:

[0096] yt=ReLU(W t *x t +b)

[0097] The temporal variation pattern of each gas concentration is extracted and aligned with the other modes.

[0098] The core goal of the mid-level feature interaction is to capture the interrelationships and complementary information between different modalities, effectively integrate the modality-specific information extracted by low-level features, and retain the uniqueness and sharing between modalities. The following are the specific steps:

[0099] 2.1 Space-time alignment

[0100] 2.1.1 Time alignment. Dynamic time warping (DTW) is used for asynchronous data to achieve time series alignment by minimizing the distance of the alignment path:

[0101] D(i,j)=||x i-y j ||+min{D(i-1,j),D(i,j-1),D(i-1,j-1)}

[0102] Where D(i,j) represents the cumulative distance to the current step, ||x i -y j || is the local distance between the two modal data.

[0103] 2.1.2 Spatial alignment. Data of different modalities may have different scales and distribution characteristics. Standardization and normalization are used to achieve consistency in feature space:

[0104] Normalize and map the data of each modality to the range [0,1]:

[0105]

[0106] Standardize and process the modal features with zero mean and unit variance:

[0107]

[0108] Among them, μ is the mean and σ is the standard deviation.

[0109] 2.2 Feature Mapping The basis of mid-level feature interaction is to map low-level features of different modalities into a shared latent space for inter-modal interaction.

[0110] 2.2.1 Linear mapping. Use linear transformation to unify low-level features to the same dimension:

[0111] z i =W i x i +b i

[0112] Among them, W i is the weight matrix, b i is the bias, x i is the low-level feature of modality i.

[0113] 2.2.2 Neural network mapping. Use multi-layer perceptron (MLP) for nonlinear mapping to extract complex inter-modal relationships

[0114] z i =ReLU(W 2 (ReLU(W 1 x i +b 1 ))+b 2 )

[0115] 2.3 Feature Interaction. The core of the middle-level interaction is to capture the complementary information between modalities, which is mainly achieved through the following mechanisms:

[0116] 2.3.1 Cross-Attention

[0117] Use the cross-modal attention mechanism to capture the interaction between modalities:

[0118] Use the attention mechanism to capture the dependencies between modalities:

[0119] Query: Extract representations from modality A.

[0120] Key and Value: Extract the representation from modality B.

[0121]

[0122] Among them, Q, K, V are query, key and value matrices respectively, d k is the dimension of the key.

[0123] For modal A and B:

[0124] F AB =Attention(F A ,F B ,F B )

[0125] For example, the interaction between partial discharge features (Fdischarge) and gas features (Fgas) is achieved through the attention mechanism:

[0126] F fusion =Attention(F discharge , F gas , F gas )

[0127] 2.3.2 Self-Attention and Cross-Attention

[0128] Intra-modality self-attention: extracting long-term dependencies within each modality:

[0129] F self =Attention(F modality ,F modality ,F modality )

[0130] Inter-modal Cross-Attention: Capturing interactive information between modalities:

[0131] F cross=Attention(F modality1 ,F modality2 ,F modality2 )

[0132] 2.3.3 Graph Neural Network Interaction (GNN)

[0133] Construct modal representations as graph structures that capture global dependencies:

[0134] Define graph nodes: each modal feature is a node.

[0135] Define graph edges: similarities or dependencies between nodes.

[0136] Graph convolution update:

[0137]

[0138] Among them, N(i) is the neighbor set of node i, and W(l) is the graph convolution weight.

[0139] 2.4 Feature Fusion

[0140] 2.4.1 Dynamic allocation of modal weights

[0141] The importance of modality is measured through a weight distribution mechanism:

[0142] The fused features:

[0143] 2.4.2 Modal Interaction and Aggregation

[0144] Fusion of inter-modality features using shared MLP

[0145] F shared =ReLU(W shared Concat(z 1 ,z 2 ,...,z M )+b shared )

[0146] 2.5 Joint Optimization Objectives

[0147] In the middle-level interaction stage, in order to ensure the effectiveness of the relationship between modalities, the loss function of the joint optimization is designed as follows:

[0148] 2.5.1 Alignment Loss:

[0149] Optimize by minimizing the difference in alignment between modalities:

[0150]

[0151] 2.5.2 Interaction Optimization

[0152] Increase mutual information between modalities:

[0153]

[0154] Here, MI is the mutual information between modalities.

[0155] High-level feature aggregation is the final stage of the multimodal fusion model, which aims to integrate the features generated by the middle-level interactions to form a global semantic representation for the final GIS fault prediction decision. This stage needs to take into account the sharing and differences of information between modalities, while strengthening key features and suppressing redundant information.

[0156] 3.1 High-Level Feature Aggregation Objectives

[0157] The core goal of high-level aggregation is to generate a unified global feature representation Fglobal to achieve:

[0158] Cross-modal information fusion: Utilize the collaborative characteristics between modalities to extract global information.

[0159] Key feature enhancement: highlight important modal features through attention mechanism or weight adjustment.

[0160] Final decision representation: Generate high-level abstract features for classification or regression tasks.

[0161] 3.2 Main methods of high-level aggregation

[0162] 3.2.1 Weighted Aggregation Based on Attention Mechanism

[0163] 3.2.1.1 Adaptive modal weight allocation

[0164] Generate dynamic weights for each modality, emphasizing the relative importance of the modalities:

[0165]

[0166] Among them, F i is the feature representation of modality i, and w is the weight vector.

[0167] 3.2.1.2 Weighted Fusion

[0168] High-level aggregation of modal features is achieved through weighted fusion:

[0169] 3.2.2 Feature Aggregation Based on Multi-Head Attention

[0170] The multi-head attention mechanism can capture global context information and is suitable for high-level aggregation stages. The specific process is as follows:

[0171] 3.2.2.1 Multi-head attention calculation:

[0172] Input: modal features {F1, F2, …, FM}

[0173] For each modal feature, calculate the query, key, and value:

[0174] Q i =W Q F i ,K i =W K Fi,V i =W V F i

[0175] Calculate the attention weights:

[0176] Weighted sum update feature:

[0177] 3.2.2.2 Multi-head mechanism:

[0178] Use multiple attention heads to capture different modality relationships:

[0179] MultiHead(F)=Concat(head 1 ,head 2 ,...,head h )W O

[0180] 3.2.2.3 Output aggregation features:

[0181] The final output global feature is: F global =MultiHead(F 1 ,F 2 ,...,F M )

[0182] 3.2.3 Global Aggregation Based on Graph Neural Network (GNN)

[0183] By constructing a modal feature graph, we can capture the complex relationship between modalities:

[0184] 3.2.3.1 Graph Construction:

[0185] Define each modal feature as a node v i , the similarity between modes is the edge weight e ij :

[0186]

[0187] 3.2.3.2 Graph convolution operation:

[0188] Update node features through graph convolutional network (GCN):

[0189]

[0190] Among them, N(i) represents the neighbor nodes of node i, and deg(i) is the degree of the node.

[0191] 3.2.3.3 Global Representation Generation:

[0192] Finally, global features are generated through global pooling: F global =Pooling(F′ 1 ,F′ 2 , ..., F′ M )3.2.4 Feature stacking based on deep fusion network

[0193] Further integrate modal information through feature stacking:

[0194] 3.2.4.1 Feature stitching:

[0195] Concatenate all modal features directly:

[0196] F concat =Concat(F 1 ,F 2 ,...,F M )

[0197] 3.2.4.2 Deep Fusion Network:

[0198] Use a fully connected layer to perform nonlinear transformation on the concatenated features:

[0199] F global =ReLU(W 2 (ReLU(W 1 F concat +b 1 ))+b 2 )

[0200] 3.2.4.3 Feature Compression:

[0201] Use a dimensionality reduction layer to reduce redundant features:

[0202] F compressed = PCA(F global )

[0203] 3.3 Optimization strategies for high-level aggregation

[0204] In order to improve the high-level aggregation effect, it is necessary to design a joint optimization goal:

[0205] 3.3.1 Choose classification or regression loss function according to the task:

[0206] For classification tasks, use cross entropy loss:

[0207]

[0208] For regression tasks, use mean squared error loss:

[0209]

[0210] 3.3.2 Alignment and Mutual Information Loss

[0211] Ensure the coordination of modal feature fusion:

[0212] Alignment loss:

[0213]

[0214] Mutual Information Optimization:

[0215]

[0216] 3.3.3 Overall loss

[0217] Weighted combination of multiple losses:

[0218]

[0219] Finally, the global feature Fglobal generated by the model through high-level aggregation is used in the GIS fault prediction module to achieve accurate classification or regression prediction. This method can not only capture the complex high-order relationships between multiple modalities, but also fully explore key features to ensure the optimal performance of the model.

[0220] Compared with traditional methods, the CNN+Transformer multimodal fusion model has significant advantages over traditional methods through the design of three stages: low-level feature extraction, mid-level feature interaction, and high-level feature aggregation. Specifically:

[0221] The low-level feature extraction stage can fully exploit modal characteristics. Traditional methods often rely on manually designed features or shallow models of a single modality, resulting in insufficient extraction of modal characteristics. This method uses CNN to extract spatial features in low-level feature extraction, which can capture local patterns (such as local defect morphology in GIS) from image data. With the help of Transformer's sequence modeling capabilities, long-range dependencies can be captured from time series. It can mine modal characteristics at a deeper level and reduce information loss. It can also adapt high-dimensional data (such as local image features) and time series data (such as local discharge signals).

[0222] The advantage of mid-level feature interaction is that it strengthens the capture of correlations between modalities. Traditional methods usually directly concatenate or simply weight modal features, which makes it difficult to fully capture the interactive information between modalities. This method uses a cross-attention mechanism in mid-level feature interaction to deeply explore the interdependence and complementarity between modalities. The self-attention mechanism within and between modalities is used to ensure local consistency within the modality and global consistency between modalities. By introducing graph neural networks, a global structural modeling of complex relationships between modalities is established. It can have a stronger capture capability for the complementarity of multimodal data, for example, by combining partial discharge signals and infrared thermal imaging data, the causes of GIS failures can be jointly revealed. Avoid the problem of insufficient information fusion caused by independent processing of modal features in traditional methods.

[0223] The advantage of high-level feature aggregation is the efficient integration of multimodal features. Traditional methods lack dynamic optimization for global features and rely only on simple classifiers. This method introduces a multi-head attention mechanism in high-level aggregation to adaptively assign importance weights of modal features to ensure the prominent role of key modal features. Feature compression and dimensionality reduction techniques are used to effectively remove redundant information and improve model calculation efficiency and prediction accuracy. Combined with joint optimization objectives (such as mutual information, alignment loss, and classification / regression loss), it ensures that the fused features have high consistency and discriminability. The importance of modal features is automatically learned by the model without manual intervention, reducing the problem of unreasonable weight distribution. Balance the diversity and consistency of modal features in the global feature representation to adapt to the multimodal data structure of complex GIS faults.

[0224] The overall performance can improve the accuracy and robustness of fault prediction. Traditional methods are mostly based on single-modal features and are easily affected by fluctuations in the quality of a single data. This method uses the complementarity of multi-modal features (such as partial discharge signals detect electrical anomalies, but infrared thermal imaging captures thermal distribution anomalies, and the two work together to improve accuracy). The model has a stronger adaptability to different modes, and even if a certain mode of data is partially lost or noisy, it can rely on other modes to compensate.

[0225] In addition, traditional methods have a shallow relationship modeling of multimodal features and are difficult to cope with complex environments. This method dynamically learns the interactive relationship between modalities, such as capturing the deep interaction between spatiotemporal features and frequency features through the multi-head attention of Transformer. It is highly scalable and can easily integrate other modalities (such as acoustic signals and ultrasonic data) to meet future expansion needs.

[0226] Embodiment 2:

[0227] The GIS fault prediction system based on multi-modal data fusion includes:

[0228] Multimodal data acquisition and preprocessing module, configured to: acquire and preprocess temperature infrared imaging data, partial discharge signals, mechanical vibration signals, and gas composition data of GIS devices;

[0229] Feature fusion and modeling module, configured to: obtain low-level features and high-level semantic features of each modality in the preprocessed data, extract spatial features and temporal features therefrom, and obtain the correlation relationship between the spatial features and the temporal features according to the self-attention mechanism and fuse the spatial features and the temporal features;

[0230] Fault prediction module, configured to: the obtained fused features pass through a fully connected network to predict the fault state of GIS, and obtain a prediction result including the fault possibility, fault location, and fault type.

[0231] 1. Hardware deployment:

[0232] Deploy an infrared thermal imager, a partial discharge monitor, and a vibration sensor; transmit real-time data to the computing platform through the data acquisition module.

[0233] 2. Data preprocessing

[0234] 2.1 Data acquisition and multimodal collation

[0235] Collect multimodal data, including: partial discharge signals (time series data), infrared thermal imaging (image data), ultrasonic signals, or other relevant modal data.

[0236] Ensure the synchronization and alignment of modal data: There may be differences in sampling frequencies between time series data and image data, and timestamp alignment and interpolation processing are required.

[0237] Unify the scales of modalities (normalization, standardization).

[0238] 2.2 Data augmentation

[0239] Design data augmentation methods for different modalities:

[0240] Time series data: Add noise, time offset, or data truncation.

[0241] Image data: Use enhancement methods such as rotation, translation, and cropping to increase the generalization ability of the model.

[0242] 2.3 Data segmentation

[0243] Segment the data into a training set, a validation set, and a test set in a ratio of 7:2:1 to ensure a balanced distribution of the dataset.

[0244] For abnormal data (such as fault categories), use undersampling or oversampling methods (SMOTE) to balance the class distribution.

[0245] 3. Model design and implementation

[0246] 3.1 Low-level feature extraction module

[0247] 3.1.1 CNN part (image modality)

[0248] Design a convolutional neural network to extract the spatial features of infrared thermal imaging:

[0249] Convolutional layers: capture local temperature variation patterns.

[0250] Pooling layer: Reduce feature dimensions and retain important information.

[0251] Feature output: Generate a two-dimensional feature map Fimage.

[0252] 3.1.2 Transformer part (time series mode)

[0253] Construct a Transformer model to extract the timing characteristics of partial discharge signals:

[0254] Encoder: Capturing global temporal dependencies via multi-head self-attention.

[0255] Feature output: Generate one-dimensional time feature Fsignal.

[0256] 3.1.3 Unified modal feature dimensions

[0257] Perform linear transformation or zero-padding on the output features of different modalities to make their dimensions consistent, providing a unified input format for middle-level interaction.

[0258] 3.2 Mid-level feature interaction module

[0259] 3.2.1 Cross-Attention Mechanism

[0260] Constructing a modality cross-attention network (Cross-Attention):

[0261]

[0262] Query (Q): Features from modality A.

[0263] Key(K) and Value(V): Features from modality B.

[0264] 3.2.2 Intra-modal and inter-modal fusion

[0265] Intra-modality: Apply self-attention mechanism to each modality to optimize the characteristics of the modality itself.

[0266] Inter-modality: Establish deep associations between modalities through cross-attention to form fusion features Finteraction

[0267] 3.2.3 Information Compensation Mechanism

[0268] When a certain mode has missing data or too much noise, the mode importance weight αi is introduced:

[0269]

[0270] 3.3 High-Level Feature Aggregation Module

[0271] 3.3.1 Global Feature Integration

[0272] Use multi-head attention mechanism to generate unified features:

[0273] F global =MultiHead(F interaction )

[0274] 3.3.2 Optimization target design

[0275] Classification task: Cross Entropy Loss

[0276] Regression task: mean squared error loss

[0277] Cross-modal consistency: alignment loss

[0278] Modal feature complementarity: mutual information loss

[0279] 3.3.3 Joint loss function

[0280] Based on the above objectives, the total loss function is designed:

[0281]

[0282] Model training and optimization

[0283] 4.1 Training Strategy

[0284] The AdamW optimizer is used to dynamically adjust the learning rate, and the initial learning rate is set to 10-4.

[0285] Adopt a pre-training-fine-tuning strategy: perform modality-specific pre-training on low-level modules (CNN or Transformer) and then jointly optimize the full model.

[0286] 4.2 Regularization and Anti-overfitting

[0287] Dropout: Dropout is added to the mid-level interaction and high-level aggregation modules to prevent overfitting.

[0288] Weight regularization: Add L2 regularization term.

[0289] 4.3 Verification and Early Stopping

[0290] Monitor performance metrics (accuracy, mean square error, etc.) on the validation set and stop training when the validation set performance no longer improves.

[0291] 5. Model Deployment

[0292] 5.1 Model Deployment Optimization

[0293] Model compression: Use quantization technology (INT8 quantization) to reduce computational complexity.

[0294] Model pruning: remove redundant neurons and optimize inference speed.

[0295] 5.2 Real-time online monitoring

[0296] Integrate the model into the online monitoring platform, analyze multimodal data input in real time, output fault prediction results, and assist operation and maintenance personnel in diagnosis and decision-making.

[0297] Embodiment three:

[0298] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps in the GIS fault prediction method of multimodal data fusion as described in the above-mentioned embodiment 2 are implemented.

[0299] Embodiment 4:

[0300] This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps in the GIS fault prediction method of multimodal data fusion as described in the above-mentioned embodiment 2 are implemented.

[0301] The steps involved in the above embodiments 2 to 4 correspond to those in embodiment 1. For the specific implementation, please refer to the relevant description of embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood to include any medium that can store, encode or carry an instruction set for execution by a processor and enable the processor to execute any method in the present invention.

[0302] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. The GIS fault prediction method based on multimodal data fusion is characterized by: The following steps are involved: Acquire temperature infrared imaging data, partial discharge signals, mechanical vibration signals and gas composition data of GIS equipment and pre-process them; Obtain low-level features and high-level semantic features of each modality in the preprocessed data, extract spatial features and temporal features, obtain the correlation between spatial features and temporal features based on the self-attention mechanism, and fuse the spatial features and temporal features; The obtained fusion features are used through a fully connected network to predict the fault status of GIS and obtain prediction results including fault possibility, fault location and fault type.

2. The GIS fault prediction method based on multimodal data fusion as claimed in claim 1, characterized in that: Preprocessing includes denoising, normalization and format conversion.

3. The GIS fault prediction method based on multimodal data fusion according to claim 1, characterized in that: The low-level features of each mode in the preprocessed data are obtained, including the frequency domain features of the partial discharge signal, the time series features of the vibration signal, and the statistical features of the gas composition.

4. The GIS fault prediction method based on multimodal data fusion as claimed in claim 1, characterized in that: Extract the spatial and temporal features, obtain the correlation between the spatial and temporal features according to the self-attention mechanism, and fuse the spatial and temporal features, including temporal and spatial alignment of the obtained features, and map the low-level features of different modalities into a shared latent space, unify the low-level features into the same dimension based on linear transformation, perform nonlinear mapping based on multi-layer perceptron, and extract the relationship between modalities; use the cross-modal attention mechanism to capture the interactive relationship between modalities: use the attention mechanism to capture the dependency between modalities.

5. The GIS fault prediction method based on multimodal data fusion as claimed in claim 4, characterized in that: The modal representation is constructed as a graph structure, with each modal feature as a node and the similarity or dependency between nodes as the edges of the graph to capture global dependencies.

6. The GIS fault prediction method based on multimodal data fusion as claimed in claim 5, characterized in that: Inter-modal features are fused through dynamic allocation of modal weights and shared multi-layer perceptrons.

7. The GIS fault prediction method based on multimodal data fusion according to claim 1, characterized in that: High-level semantic features, specifically: utilizing the collaborative characteristics between modalities, extracting global information, highlighting important modal features through attention mechanisms or weight adjustments, and generating high-level abstract features for classification or regression tasks.

8. GIS fault prediction system based on multimodal data fusion, characterized by: include: The multimodal data acquisition and preprocessing module is configured to: acquire and preprocess the temperature infrared imaging data, partial discharge signal, mechanical vibration signal and gas composition data of the GIS equipment; The feature fusion and modeling module is configured to: obtain low-level features and high-level semantic features of each modality in the preprocessed data, extract spatial features and temporal features therein, obtain the correlation between spatial features and temporal features according to the self-attention mechanism, and fuse the spatial features and temporal features; The fault prediction module is configured as follows: the obtained fusion features are used through a fully connected network to predict the fault state of the GIS and obtain a prediction result including the fault possibility, fault location and fault type.

9. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the program is executed by a processor, the steps in the GIS fault prediction method of multimodal data fusion as described in any one of claims 1-7 are implemented.

10. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the method implements the steps in the GIS fault prediction method of multimodal data fusion as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Machine tool thermal error prediction method based on multi-modal deep learning and domain adaptation

    CN120234592A

  • Unmanned aerial vehicle motor fault diagnosis method based on multi-modal data fusion

    CN120354371A

  • Multi-mode fault prediction method and device for energy storage wireless BMS system and storage medium

    CN120446769A

  • Fault detection method and device, equipment and storage medium

    CN120449060A

  • Clean air switch equipment fault diagnosis method and system

    CN120595097A