Transformer fault diagnosis method, system and equipment based on multi-modal deep learning, and storage medium

Through multimodal deep learning methods, combined with integrated gradient wavelet transform, stacked autoencoder and convolutional autoencoder networks, and Dempster-Shafer evidence theory, the multimodal data fusion problem in transformer fault diagnosis is solved, high-precision and high-reliability fault diagnosis is achieved, the misdiagnosis rate is reduced and the ability to detect early faults is improved.

CN120747690APending Publication Date: 2025-10-03GUIZHOU POWER GRID CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510840757.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing transformer fault diagnosis technologies have problems such as insufficient multimodal data fusion, heterogeneous fusion of numerical data and image data, lack of specialized feature extraction methods, poor performance of traditional stacked autoencoders, and lack of theoretical support for decision-level fusion of multi-source information, resulting in the loss of important fault information and a high misdiagnosis rate.

Method used

A multimodal deep learning method is adopted to process vibration signals through continuous wavelet transform with integrated gradient. Stacked autoencoder and convolutional autoencoder networks are constructed to extract deep features of numerical and image data. Dempster-Shafer evidence theory is used for decision-level fusion to achieve efficient fusion of multi-source information and accurate diagnosis.

Benefits of technology

The accuracy and reliability of transformer fault diagnosis have been significantly improved, especially the diagnostic confidence of early and weak faults has been increased by more than 30%, the misdiagnosis rate has been reduced to less than 5%, and the accuracy rate has been maintained at more than 94% under complex working conditions, providing accurate early warning of faults.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120747690A_ABST
    Figure CN120747690A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of power equipment monitoring, in particular to a transformer fault diagnosis method, system and device based on multi-modal deep learning and a storage medium. Obtaining the volume fraction of gas dissolved in oil of the transformer, the local discharge capacity, the sleeve dielectric loss factor, the vibration data and the infrared image data; continuous wavelet transform processing based on integrated gradient is carried out on the vibration data, and key fault frequency bands are dynamically screened to generate a time-frequency map; inputting the numeric data into a stack-type denoising auto-encoder network to extract depth features; respectively inputting the infrared image and the radio frequency map into a double-branch convolution encoder for early feature fusion, and extracting map depth features through a stack type convolution auto-encoder network; and based on the Dempster-Shafer evidence theory, carrying out conflict resolution and evidence synthesis on the fault probability distribution output by the two types of modes, and outputting a diagnosis result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of power equipment monitoring, and in particular to a transformer fault diagnosis method, system, equipment and storage medium based on multimodal deep learning. Background Art

[0002] Power transformers, as key equipment in the power system, undertake the crucial tasks of voltage conversion, power distribution, and transmission. With the continuous growth of electricity demand and the continued expansion of the power grid, the safe and stable operation of transformers has a crucial impact on the reliability of the entire power system. During long-term operation, transformers are subjected to a variety of stresses, including electrical, thermal, mechanical, and chemical stresses. These faults can cause a variety of problems, including winding short circuits, insulation aging, oil degradation, and core failures. If these faults are not detected and addressed promptly, they can not only cause equipment damage and economic losses, but can also trigger widespread power outages, severely impacting social production and daily life.

[0003] Traditional transformer fault diagnosis methods primarily include dissolved gas analysis (DGA), partial discharge detection, and insulation resistance testing. While these methods have achieved some success in practical applications, they all have limitations. DGA primarily analyzes characteristic gases dissolved in transformer oil to determine the fault type, but its sensitivity to early-stage faults is low. Partial discharge detection can detect insulation defects, but is susceptible to on-site electromagnetic interference. Insulation resistance testing requires a power outage, making online monitoring impossible. More importantly, these traditional methods rely on a single source of information for diagnosis and fail to fully reflect the transformer's actual operating status.

[0004] However, existing multi-source information fusion diagnostic technologies still have many shortcomings. First, most methods only process numerical data, such as parameters such as gas concentration in oil, voltage, and current, and lack effective processing methods for unstructured data such as infrared images and vibration signals. Second, even if some methods attempt to fuse image data, they usually use traditional image processing techniques to extract features and cannot fully tap into the deep information in the image. For time series data such as vibration signals, existing methods mostly use simple time domain or frequency domain features, ignoring the importance of joint time-frequency analysis. In addition, the heterogeneity problem between different modal data has not been well solved, and simple feature splicing makes it difficult to achieve true information complementarity. Summary of the Invention

[0005] In view of the problems existing in the prior art, the present invention is proposed.

[0006] Therefore, the problem to be solved by this invention is how to address the following issues existing in existing transformer fault diagnosis technology: first, insufficient multimodal data fusion, especially the heterogeneous fusion of numerical and image data; second, the lack of specialized feature extraction methods for vibration signals and infrared images, which leads to the loss of important fault information; third, the poor performance of traditional stacked autoencoders when processing image data, and the inability to effectively extract spatial features; and fourth, the lack of theoretical support for the decision-level fusion of multi-source information, making simple voting or weighted averaging methods difficult to handle evidence conflicts. By addressing these issues, high-precision and high-reliability diagnosis of transformer faults can be achieved.

[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:

[0008] In a first aspect, an embodiment of the present invention provides a transformer fault diagnosis method based on multimodal deep learning, which includes obtaining the volume fraction of dissolved gas in transformer oil, partial discharge amount, bushing dielectric loss factor, vibration data and infrared image data;

[0009] The vibration data is processed by continuous wavelet transform based on integrated gradient to obtain time-frequency spectrum data;

[0010] The volume fraction of dissolved gas in oil, the amount of partial discharge, and the dielectric loss factor of the casing are input into multiple denoising autoencoders to extract low-dimensional feature representations, and the low-dimensional feature representations are input into a stacked autoencoder network formed by stacking multiple denoising autoencoders to extract numerical modal depth features;

[0011] The infrared image data and time-frequency spectrum data are input into two convolutional encoders respectively, and their feature vectors are extracted and concatenated. The fused features are input into a stacked convolutional autoencoder network formed by stacking multiple convolutional encoders to extract the spectrum modality deep features.

[0012] Based on the Dempster-Shafer evidence theory, the classification outputs of numerical modal depth features and spectral modal depth features are fused to obtain the transformer fault diagnosis results.

[0013] As a preferred solution of the transformer fault diagnosis method based on multimodal deep learning described in the present invention, it includes:

[0014] Collect numerical data and spectral data during transformer operation. The numerical data includes the volume fraction of H2, C2H2, CH4, C2H6, CO, and C2H4 dissolved gases in the oil, partial discharge, and bushing dielectric loss factor. The spectral data includes vibration data and infrared images.

[0015] Perform continuous wavelet transform based on integrated gradient on vibration data to generate vibration time-frequency spectrum;

[0016] Build a stacked autoencoder network to process numerical data and extract deep feature representations through layer-by-layer encoding;

[0017] A stacked convolutional autoencoder network is constructed to process atlas data. The first layer uses two independent convolutional encoders to process infrared images and vibration time-frequency spectra respectively. The second layer and subsequent layers perform feature fusion and deep extraction.

[0018] The fault probability distributions output by the two networks are used as basic probability distributions, and the fusion probability is calculated using the evidence synthesis formula;

[0019] The transformer fault types are determined based on the fused probability distribution, including low-temperature overheating, medium-low-temperature overheating, medium-temperature overheating, high-temperature overheating, low-energy discharge, high-energy discharge and normal state.

[0020] As a preferred solution of the transformer fault diagnosis method based on multimodal deep learning of the present invention, the continuous wavelet transform based on integrated gradient includes:

[0021] The attribution relationship between output and input features is predicted through the integral gradient calculation model; the smooth gradient method is used to reduce the noise in the attribution interpretation; the importance score of each frequency is calculated based on the attribution results; the important frequency range is determined and a continuous wavelet transform is performed within this range.

[0022] As a preferred solution of the transformer fault diagnosis method based on multimodal deep learning described in the present invention, the construction of the stacked autoencoder network includes:

[0023] Each layer of denoising autoencoder adds noise to the input data and then encodes it to learn the hidden features of the data; the decoding part of each layer of denoising autoencoder is removed, and the hidden layer of the encoding part is retained; the hidden layers of each layer are stacked and connected to form a deep network structure; and a Softmax classifier is used in the last layer to output the fault probability.

[0024] As a preferred solution of the transformer fault diagnosis method based on multimodal deep learning described in the present invention, the construction of the stacked convolutional autoencoder network includes:

[0025] The first layer uses two convolutional encoders to extract the features of the infrared image and time-frequency spectrum respectively; the two feature vectors output by the first layer are spliced ​​and fused; the fused features are input into the subsequent convolutional encoder layer to extract deep features layer by layer; the last layer uses the Softmax function to output the fault probability distribution.

[0026] As a preferred solution of the transformer fault diagnosis method based on multimodal deep learning described in the present invention, the fusion based on Dempster-Shafer evidence theory includes:

[0027] The outputs of the stacked autoencoder network and the stacked convolutional autoencoder network are used as two independent evidence sources; the basic probability distribution function corresponding to each evidence source is calculated; the degree of conflict between the evidence is calculated through evidence synthesis rules; and the multi-source evidence is synthesized based on the conflict degree to obtain the final fault diagnosis result.

[0028] As a preferred solution of the transformer fault diagnosis method based on multimodal deep learning of the present invention, the method further includes:

[0029] The stacked autoencoder network is trained using a composite loss function containing a marginal Fisher regularization term. The stacked convolutional autoencoder network is optimized using a layer-by-layer greedy training method. The loss function is minimized using the gradient descent algorithm to obtain the optimal model parameters.

[0030] In a second aspect, an embodiment of the present invention provides a transformer fault diagnosis system based on multimodal deep learning, which includes a data acquisition module for obtaining the volume fraction of dissolved gas in transformer oil, partial discharge, bushing dielectric loss factor, vibration data and infrared image data;

[0031] The data processing module is used to perform continuous wavelet transform based on integrated gradient on the vibration data to generate time-frequency spectrum data and extract deep features of numerical modal data through a stacked autoencoder network;

[0032] The graph feature extraction module is used to extract deep features of graph modality data through a stacked convolutional autoencoder network;

[0033] The decision fusion module is used to fuse the classification results of the two modes based on the Dempster-Shafer evidence theory and output the transformer fault diagnosis results.

[0034] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program instructions are executed by the processor, the steps of the transformer fault diagnosis method based on multimodal deep learning as described in the first aspect of the present invention are implemented.

[0035] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program instructions are executed by a processor, the steps of the transformer fault diagnosis method based on multimodal deep learning as described in the first aspect of the present invention are implemented.

[0036] The beneficial effects of the present invention are as follows: by processing vibration signals through integrated gradient continuous wavelet transform (IG-CWT), fault-sensitive frequency bands are dynamically screened to generate time-frequency spectra, thus solving the problems of large noise interference and blind feature selection in traditional time-frequency analysis; at the same time, a stacked denoising autoencoder is designed for numerical data (gas concentration, discharge amount, etc.), and robust low-dimensional features are extracted through noise injection and layered coding compression; a dual-branch convolutional encoder is used for spectral data (infrared images, time-frequency graphs) to achieve early feature splicing and fusion of heterogeneous spectra at the first layer. This sub-modal dedicated feature extraction architecture has for the first time achieved the collaborative in-depth mining of the dynamic characteristics of vibration signals, gas chemical characteristics, and thermal image spatial characteristics, overcoming the limitations of single modal information and the arbitrariness of traditional feature fusion.

[0037] Based on the Dempster-Shafer (DS) evidence theory, the fault probability distributions of numerical and spectral modal outputs are used as independent sources of evidence. Uncertainty is quantified through a basic probability distribution function, and conflict factors are used to dynamically modify fusion weights. Compared to conventional weighted averaging or voting fusion, this design achieves adaptive arbitration of conflicting evidence in transformer diagnosis for the first time, significantly reducing the misdiagnosis rate. Especially for early, weak faults (such as low-temperature overheating and low-energy discharge), the complementary fusion of multi-source evidence increases diagnostic confidence by over 30%, resolving the pain point of traditional methods that have limited ambiguity in determining complex faults and boundary states.

[0038] By using marginal Fisher regularization to constrain the feature space distribution of stacked autoencoders, similar fault features are clustered and heterogeneous features are separated, significantly improving the generalization ability of small sample faults. Layer-by-layer greedy training and residual connections are used for the convolutional autoencoders to mitigate vanishing gradients while preserving multi-scale spatial features. This combination of "enhanced feature discriminability and optimized structural stability" enables the model to maintain an accuracy of over 94% under complex operating conditions such as strong electromagnetic interference and missing data, far exceeding the fluctuation range of traditional models.

[0039] Through the technical closed loop of "vibration IG-CWT dynamic focusing → numerical / spectral sub-modal deep encoding → DS evidence conflict resolution", this solution has achieved for the first time the accurate distinction between six types of overheating / discharge faults and normal conditions. The diagnostic confidence interval is narrowed by more than 40% compared with existing technologies, and the false alarm rate is reduced to less than 5%, providing irreplaceable technical support for early warning of latent transformer faults. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0041] Figure 1 This is a flow chart of a transformer fault diagnosis method based on multimodal deep learning;

[0042] Figure 2 Computer equipment diagram for transformer fault diagnosis method based on multimodal deep learning;

[0043] Figure 3 Schematic diagram of the IG-CWT framework for transformer fault diagnosis based on multimodal deep learning;

[0044] Figure 4 This is the schematic diagram of the denoising autoencoder for transformer fault diagnosis based on multimodal deep learning;

[0045] Figure 5 This is the schematic diagram of the convolutional autoencoder for the transformer fault diagnosis method based on multimodal deep learning;

[0046] Figure 6 Another flowchart of the transformer fault diagnosis method based on multimodal deep learning. DETAILED DESCRIPTION

[0047] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.

[0048] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0049] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it constitute a separate or selective embodiment that is mutually exclusive with other embodiments.

[0050] Example 1

[0051] Reference Figure 1-Figure 2 , which is the first embodiment of the present invention, provides a transformer fault diagnosis method based on multimodal deep learning, including:

[0052] S100: Obtain the volume fraction of dissolved gas in transformer oil, partial discharge, bushing dielectric loss factor, vibration data, and infrared image data, and process the vibration data through continuous wavelet transform based on integrated gradient to obtain time-frequency spectrum data;

[0053] S200: Input the volume fraction of dissolved gas in oil, the amount of partial discharge, and the dielectric loss factor of the casing into multiple denoising autoencoders to extract low-dimensional feature representations. The low-dimensional feature representations are then input into a stacked autoencoder network formed by stacking multiple denoising autoencoders to extract numerical modal depth features.

[0054] S300: Input the infrared image data and the time-frequency spectrum data into two convolutional encoders respectively, extract their respective feature vectors and perform splicing and fusion, and input the fused features into a stacked convolutional autoencoder network formed by stacking multiple convolutional encoders to extract the spectrum modality deep features;

[0055] S400: Based on the Dempster-Shafer evidence theory, the classification outputs of numerical modal depth features and spectral modal depth features are fused to obtain the transformer fault diagnosis results.

[0056] It should be noted that during actual transformer operation, due to the complex internal structure and harsh operating environment, faults are often diverse and hidden. Traditional single monitoring methods fail to fully reflect the transformer's health. For example, relying solely on dissolved gas analysis in oil can miss early signs of partial discharge, while infrared monitoring alone cannot detect internal insulation degradation. Furthermore, complex correlations exist between different types of monitoring data. For example, abnormal vibration is often accompanied by temperature increases, while the development of partial discharge can lead to changes in the concentration of characteristic gases in the oil. The interweaving of these multi-dimensional and multi-modal data features makes accurate transformer fault diagnosis a significant challenge. Traditional methods are particularly sensitive for early-stage faults and latent defects, as their signals are weak and easily masked by noise. Furthermore, existing technologies often use simple time-domain or frequency-domain analysis when processing vibration signals, ignoring the dynamic evolution of fault characteristics in the time-frequency domain. When processing infrared images, they lack deep feature extraction capabilities and are unable to fully explore the inherent connection between temperature distribution patterns and fault types.

[0057] Therefore, through steps S100-S400, the present invention constructs a complete multimodal deep learning diagnostic framework. First, S100 acquires monitoring data covering multiple physical fields such as chemistry, electricity, acoustics, and thermal science. The integrated gradient-based continuous wavelet transform (IG-CWT) is innovatively introduced to process vibration data. This method can adaptively identify key frequency components related to faults and effectively extract dynamic features in the time-frequency domain. S200 uses a stacked autoencoder network to perform deep feature extraction on numerical data, utilizing a denoising mechanism to enhance the robustness of the model. It learns abstract representations of the data layer by layer, ultimately obtaining highly concentrated numerical modal features. S300, targeting the spatial structure characteristics of image data, employs a stacked convolutional autoencoder network. This convolution operation preserves local correlations and implements feature fusion of infrared images and vibration time-frequency maps in the first layer, fully exploiting the complementary information of different spectral data. Finally, S400 performs decision-level fusion based on the Dempster-Shafer (DS) evidence theory. This theory can effectively handle conflicts and uncertainties between different evidence sources, and obtains more reliable diagnostic results by calculating basic probability distribution and evidence synthesis. The entire method realizes a complete diagnostic process from multi-source heterogeneous data collection, targeted feature extraction, deep feature learning to intelligent decision fusion. It can accurately identify six fault types including low-temperature overheating, medium-low-temperature overheating, medium-temperature overheating, high-temperature overheating, low-energy discharge, and high-energy discharge, as well as normal conditions, significantly improving the accuracy and reliability of transformer fault diagnosis.

[0058] Example 2

[0059] Reference Figures 1-6 , which is the second embodiment of the present invention.

[0060] In the embodiment of the present application, obtaining the volume fraction of dissolved gas in transformer oil, partial discharge amount, bushing dielectric loss factor, vibration data, and infrared image data in step S100 includes the following steps A1-A4:

[0061] A1: Includes:

[0062] Collect numerical data and spectral data during transformer operation. The numerical data includes the volume fraction of H2, C2H2, CH4, C2H6, CO, and C2H4 dissolved gases in the oil, partial discharge, and bushing dielectric loss factor. The spectral data includes vibration data and infrared images.

[0063] Specifically, during the dissolved gas collection process, an online chromatographic analyzer is used to extract the dissolved gas in the transformer oil in real time through an oil-gas separation device. The sampling period is set to 4 hours to ensure that abnormal changes in gas content can be detected in a timely manner. For H2 collection, since it is a sensitive indicator of internal discharge and overheating faults in the transformer, a thermal conductivity detector is used for precise measurement, with a detection accuracy of 0.5μL / L.

[0064] As a characteristic gas of arc discharge, H2 has important diagnostic significance even at extremely low levels. Therefore, a flame ionization detector (FID) is used to ensure a detection limit of 0.1 μL / L. Hydrocarbon gases such as CH4, C2H6, and C2H4 are also detected using a FID. The relative concentrations and growth rates of these gases can reflect overheating faults at different temperature levels. CO, an indicator gas for aging of cellulose insulation materials, is detected using infrared absorption, enabling assessment of the degree of degradation of solid insulation.

[0065] Partial discharge measurements are collected using a combination of ultra-high frequency (UHF) detection and pulse current methods. The UHF sensor is installed in the transformer's observation window or oil drain valve, with a detection frequency range of 300MHz-1500MHz, effectively avoiding on-site electromagnetic interference. The pulse current method uses a high-frequency current sensor connected to the transformer's ground wire, with a detection frequency range of 10kHz-30MHz. Combining these two methods improves the reliability and location accuracy of partial discharge detection. The acquisition system's sampling rate is set to 1GS / s to ensure accurate capture of nanosecond-level discharge pulses.

[0066] Bushing dielectric loss factor is measured online using the grounding current and voltage signals at the end screen. The measurement system utilizes high-precision current and voltage transformers, achieving current measurement accuracy of 0.1% and phase measurement accuracy of 0.01°. Changes in dielectric loss factor can indicate defects in bushing insulation, such as moisture and aging, and are a key parameter for assessing bushing health.

[0067] Vibration data is collected using triaxial accelerometers installed at various locations on the transformer tank, including the high-voltage side, low-voltage side, and near the neutral point. The sensors have a frequency response range of 0.5 Hz to 10 kHz and a sensitivity of 100 mV / g, enabling comprehensive capture of the transformer's vibration characteristics. The sampling frequency is set to 25.6 kHz, meeting the Nyquist sampling theorem. Each measurement point continuously collects vibration signals for 10 seconds, generating a time series consisting of 256,000 data points.

[0068] Infrared images are captured using a cooled infrared thermal imager with a wavelength range of 8-14μm, a temperature resolution of 0.05°C, and a spatial resolution of 640×480 pixels. Images are taken from a distance of 3-5 meters from the transformer surface to ensure coverage of the entire device's temperature distribution. Panoramic infrared images are automatically captured every two hours, with key areas such as the high-voltage bushing, cooler, and tap changer monitored.

[0069] A2: Perform integrated gradient-based continuous wavelet transform on the vibration data to generate a vibration time-frequency spectrum;

[0070] During the vibration signal processing phase, the raw vibration data is first preprocessed, including DC component removal, trend elimination, and outlier removal. The continuous wavelet transform (CWT) method based on integrated gradients is then applied. The core of this method is to automatically identify the frequency components most important for fault diagnosis through the backpropagation mechanism of a machine learning model.

[0071] During the integrated gradient calculation process, the vibration signal is used as input. A pretrained fault classification model is used to calculate the contribution of each frequency component to the final diagnosis result. The integration path is from the baseline (all zero signal) to the actual signal, with a step size of 50 to ensure the accuracy of the gradient calculation. For each frequency f, its importance score is calculated by averaging the gradient values ​​at all time points.

[0072] Smoothed gradients are introduced to reduce the impact of noise in gradient calculations. By adding Gaussian noise (with a standard deviation of 5% of the signal amplitude) to the original signal, generating 100 perturbed samples, calculating the gradient of each sample and averaging it, a more stable frequency importance assessment is obtained.

[0073] Based on the calculated frequency importance scores, a threshold of 1.5 times the average score was set to screen out important frequencies. These frequencies typically correspond to the transformer's natural vibration frequency (such as 100 Hz and its harmonics), the core magnetostrictive frequency, and characteristic frequencies caused by winding looseness. The identified important frequency range serves as the basis for selecting the scale parameter for the continuous wavelet transform.

[0074] In the continuous wavelet transform, the Morlet wavelet was chosen as the mother wavelet due to its excellent localization properties in both the time and frequency domains. The scale parameter was adaptively adjusted based on the critical frequency range, and the time step was set to 0.1ms. The resulting time-frequency spectrum clearly displays the time-frequency distribution characteristics of the vibration signal, providing rich information for subsequent feature extraction.

[0075] A3: Build a stacked autoencoder network to process numerical data and extract deep feature representations through layer-by-layer encoding;

[0076] The stacked autoencoder network is constructed using a layer-by-layer pre-training strategy. The first-layer denoising autoencoder has an input dimension of 8 (corresponding to the six gas contents, partial discharge, and dielectric loss factor), and the number of hidden layer neurons is set to 16. During training, Gaussian noise with a mean of 0 and a standard deviation of 0.1 is added to the input data to force the encoder to learn a robust feature representation of the data.

[0077] The encoder uses fully connected layers and a ReLU activation function to introduce nonlinearity and avoid the vanishing gradient problem. Weight initialization uses the Xavier method to ensure that the activation values ​​of each layer maintain an appropriate variance during forward propagation. The bias term is initialized to 0.01 to avoid neuron death.

[0078] After training the first layer, the decoder was removed, and the encoder output was used as the input for the second layer. The number of neurons in the second hidden layer was set to 32, and training continued using the denoising strategy. In this way, a stacked structure consisting of three hidden layers was constructed, with 16, 32, and 64 neurons, respectively.

[0079] During the training of each layer, the Adam optimizer was used, with an initial learning rate of 0.001 and an exponential decay strategy, which decayed by a factor of 0.95 every 1000 epochs. The batch size was set to 64, and the number of training epochs was 500. To prevent overfitting, a dropout layer was added after each hidden layer, with a dropout rate of 0.2.

[0080] The final layer is connected to a Softmax classifier, with an output dimension of 7, corresponding to the six fault types and the normal state. The loss function for the entire network uses a combination of cross-entropy loss and a marginal Fisher regularization term, with the regularization coefficient determined by tuning on the validation set.

[0081] A4: Build a stacked convolutional autoencoder network to process the atlas data. The first layer uses two independent convolutional encoders to process the infrared image and the vibration time-frequency spectrum respectively. The second layer and subsequent layers perform feature fusion and deep extraction.

[0082] The fault probability distributions output by the two networks are used as basic probability distributions, and the fusion probability is calculated using the evidence synthesis formula;

[0083] The transformer fault types are determined based on the fused probability distribution, including low-temperature overheating, medium-low-temperature overheating, medium-temperature overheating, high-temperature overheating, low-energy discharge, high-energy discharge and normal state.

[0084] The stacked convolutional autoencoder network is specifically designed to process spatially structured atlas data. The first layer contains two parallel convolutional encoder branches, one for processing infrared images and the other for processing vibration time-frequency spectra.

[0085] For the infrared image branch, the input size is 256×256×1 (grayscale image). The first convolutional layer uses 32 3×3 convolution kernels with a stride of 1 and the same padding to maintain the spatial size of the feature map. The activation function uses LeakyReLU with a negative slope of 0.1 to avoid neuron death. This is followed by a 2×2 max pooling layer to halve the feature map size. The second convolutional layer uses 64 3×3 convolution kernels with the same settings, followed by another pooling layer. Through two convolution and pooling operations, the original image is encoded into a 64×64×64 feature representation.

[0086] The vibration time-frequency spectrum branch has a similar structure, but considering the characteristics of time-frequency maps, the first convolutional layer uses 64 5×5 convolution kernels to capture a wider range of time-frequency patterns. The subsequent layers are configured the same as the infrared image branch, ultimately resulting in a 64×64×64 feature representation.

[0087] In the second layer, the feature maps of the two branches are concatenated along the channel dimension to form a 64×64×128 fused feature. This early fusion strategy enables the network to learn the correlation between the two modal data. The fused features are then processed through convolutional layers. The third convolutional layer uses 128 3×3 convolution kernels, and the fourth convolutional layer uses 256 3×3 convolution kernels. Each convolutional layer is followed by a pooling operation.

[0088] To enhance the network's feature extraction capabilities, a batch normalization layer is added after each convolutional layer to accelerate training and improve the model's generalization capabilities. At the same time, residual connections are introduced in deep networks to alleviate the vanishing gradient problem.

[0089] The final encoder output passes through a global average pooling layer to obtain a 256-dimensional feature vector, which is input to the Softmax classifier. The entire network is trained end-to-end using categorical crossentropy as the loss function.

[0090] In an optional embodiment, the data collection in step S100 may also include acquiring sound signals. An acoustic sensor array is deployed around the transformer to collect sound signals during operation. Sound signals can reflect abnormalities such as discharge and vibration within the transformer and are particularly sensitive to early-stage faults. After noise reduction processing, the collected sound signals are extracted using Mel-Frequency Cepstral Coefficient (MFCC) features, which are then input into the diagnostic system as supplementary numerical modal data.

[0091] In another optional embodiment, vibration data processing can also incorporate empirical mode decomposition (EMD). EMD can adaptively decompose nonlinear, nonstationary vibration signals into several intrinsic mode functions (IMFs), each representing a characteristic scale in the signal. By analyzing the energy distribution and frequency characteristics of each IMF component, fault signatures can be more accurately identified. Combining the EMD decomposition results with the time-frequency spectrum of IG-CWT can provide a more comprehensive representation of vibration characteristics.

[0092] In the embodiment of the present application, in step S200, the volume fraction of dissolved gas in oil, the amount of partial discharge, and the dielectric loss factor of the casing are respectively input into multiple denoising autoencoders to extract low-dimensional feature representations, including the following steps B1-B4:

[0093] B1: Continuous wavelet transform based on integrated gradient includes:

[0094] The attribution relationship between output and input features is predicted through the integral gradient calculation model; the smooth gradient method is used to reduce the noise in the attribution interpretation; the importance score of each frequency is calculated based on the attribution results; the important frequency range is determined and a continuous wavelet transform is performed within this range.

[0095] Before inputting numerical data into the denoising autoencoder, detailed preprocessing is required. For dissolved gas data in oil, since the concentration ranges of different gases vary widely (for example, H2 is typically between 10 and 1000 μL / L, while C2H2 may only be between 0.1 and 10 μL / L), directly using the raw data will cause the model to favor features with larger values. Therefore, a logarithmic transformation combined with Z-score normalization is employed: First, the gas concentrations are logarithmically transformed (x' = log(x + 1)) to avoid zero values; then, the mean and standard deviation of the transformed data are calculated and Z-score normalization is performed.

[0096] Partial discharge processing takes into account its pulse characteristics and employs a segmented statistical approach. The collected discharge pulses are divided into 12 intervals (one interval every 30°) based on their phase. The number of discharges, average amplitude, and maximum amplitude in each interval are then counted to form a 36-dimensional feature vector. This representation preserves the phase distribution of the discharge, which is crucial for identifying different types of insulation defects.

[0097] The casing dielectric loss factor is typically a relatively stable parameter, with a normal value between 0.3% and 0.5%. To enhance the model's sensitivity to outliers, the relative rate of change is used as an input feature, calculating the percentage change of the current value relative to the historical average.

[0098] Noise addition is the core mechanism of the denoising autoencoder. Adaptive noise addition strategies are designed based on the characteristics of different data types. For gas concentration data, the noise standard deviation is set to 0.15 times the standard deviation of the feature. For partial discharge features, given their discrete nature, a random zeroing strategy is used, with certain features set to zero with a probability of 0.1. For the dielectric loss factor, uniformly distributed noise is added within the range of [-0.05% to +0.05%].

[0099] B2: Parallel training strategy for multiple denoising autoencoders.

[0100] Considering that different types of numerical data have different physical meanings and statistical characteristics, three independent denoising autoencoders are designed to process gas data, discharge data and dielectric loss data respectively.

[0101] The gas data denoising autoencoder has an input dimension of 6, corresponding to the concentrations of the six gases. The encoder consists of two hidden layers, with 12 and 8 neurons, respectively. This architecture was chosen based on the information bottleneck theory, which states that moderate dimensionality reduction forces the model to learn the most important features. The decoder adopts a symmetrical structure, with 8 and 12 neurons in the hidden layers. The mean squared error (MSE) is used as the reconstruction loss for the entire network.

[0102] The input dimension of the discharge data denoising autoencoder is 36, and the number of neurons in the encoder hidden layer is 24 and 16. Due to the sparsity of discharge data, an L1 regularization term is added to the loss function to encourage the model to learn sparse representations. The regularization coefficient is determined to be 0.001 through grid search.

[0103] Since the dielectric loss data has a low dimension (usually monitoring three-phase bushings, the input dimension is 3), a single hidden layer structure is used with 5 neurons. This simple structure is sufficient to capture the variation pattern of dielectric loss.

[0104] The three denoising autoencoders were trained in parallel, each using its own optimizer and learning rate schedule. The reconstruction error on the validation set was monitored during training, and early stopping was triggered if there was no improvement for 20 consecutive epochs.

[0105] B3: Feature representation extraction and dimension alignment.

[0106] After training is complete, the decoder portion of each denoising autoencoder is removed, leaving only the encoder for feature extraction. For new input data, the three encoders output 8-dimensional, 16-dimensional, and 5-dimensional feature representations, respectively.

[0107] For subsequent feature fusion, features of different dimensions need to be aligned. Using feature mapping, the three sets of features are mapped to a unified 16-dimensional space via a fully connected layer. Batch normalization is used during the mapping process to ensure that features from different sources have similar numerical ranges.

[0108] The weights of the feature mapping layer are learned through a separate training process. Specifically, an auxiliary classification task is constructed, where the mapped features are fed into a simple classifier and the mapping weights are optimized by minimizing the classification loss. This approach ensures that the mapped features retain the discriminative information of the original features.

[0109] B4: Quality evaluation of low-dimensional feature representations.

[0110] To ensure that the extracted low-dimensional features truly contain the key information of the original data, we designed several evaluation metrics. Reconstruction error is the most direct metric, calculating the mean squared error between the original and reconstructed data. For a trained model, the average reconstruction error on the test set should be less than 0.1.

[0111] The discriminability of features is evaluated by linear separability test. Using linear SVM to classify low-dimensional features, if the classification accuracy can reach more than 85%, it means that the feature has good discriminability.

[0112] The distribution characteristics of features are qualitatively evaluated using t-SNE visualization. High-dimensional features are projected into a two-dimensional space to observe whether samples of different fault types form distinct clusters. Ideally, samples of the same type should cluster together, with clear boundaries between samples of different types.

[0113] In an optional implementation, the denoising autoencoder can also be trained using the variational autoencoder (VAE) framework. By introducing probability distribution assumptions, VAEs can learn smoother and more continuous latent space representations. The encoder outputs the mean and variance of the features, which are then sampled using a reparameterization technique to obtain the latent representation. This approach is particularly well-suited for data sparseness, such as when there are few samples of certain rare fault types.

[0114] In another alternative implementation, adversarial training can be introduced to improve feature robustness. Based on the denoising autoencoder, a discriminator network is added to distinguish between the original data and the reconstructed data. Through adversarial training, the autoencoder is forced to produce more realistic reconstructions, while also learning more robust feature representations.

[0115] It should be noted that the design of a denoising autoencoder must consider computational efficiency in practical applications. While deeper network structures may lead to better feature extraction, they also increase inference time. In the case of online transformer monitoring, a balance must be struck between feature quality and computational efficiency. Experiments have shown that a 2-3 hidden layer structure can maintain feature quality while keeping single-sample inference time below 10ms.

[0116] In the embodiment of the present application, in step S300, the infrared image data and the time-frequency spectrum data are respectively input into two convolutional encoders, and their respective feature vectors are extracted and spliced ​​and fused, which includes the following steps C1-C4:

[0117] C1: The construction of the stacked autoencoder network includes:

[0118] Each layer of denoising autoencoder adds noise to the input data and then encodes it to learn the hidden features of the data; the decoding part of each layer of denoising autoencoder is removed, and the hidden layer of the encoding part is retained; the hidden layers of each layer are stacked and connected to form a deep network structure; and a Softmax classifier is used in the last layer to output the fault probability.

[0119] Raw infrared images often contain significant background noise and environmental interference, requiring meticulous preprocessing. First, non-uniformity correction is performed to eliminate fixed pattern noise caused by inconsistent detector pixel responses. A two-point correction method is employed, using blackbody calibration data at two different temperatures to calculate the gain and bias coefficients for each pixel.

[0120] Temperature data is extracted using a lookup table. Based on the infrared detector's response curve and environmental parameters (temperature, humidity, and measurement distance), the raw grayscale values ​​are converted to temperature values. To improve temperature measurement accuracy, the emissivity of the object being measured is input. The emissivity of the painted surface of a transformer is typically set to 0.95.

[0121] Image enhancement uses the Color-Adaptive Histogram Equalization (CLAHE) method, which divides the image into 8×8 sub-blocks. Histogram equalization is performed independently on each sub-block, followed by bilinear interpolation to remove blocking artifacts. The cropping limit parameter is set to 2.0 to prevent excessive noise enhancement. This method enhances local contrast in the image, making areas of temperature anomalies more distinct.

[0122] To highlight areas of abnormal temperature, a statistical anomaly detection method is used. The mean and standard deviation of the temperature across the entire image are calculated, and regions exceeding the mean plus twice the standard deviation are marked as potential hotspots. Morphological processing, including dilation and erosion operations, is then performed on these regions to remove isolated noise points and retain connected high-temperature areas.

[0123] Image normalization is a necessary step for convolutional network input. The preprocessed image is resized to 256×256 pixels and bicubic interpolation is used to maintain image smoothness. The aspect ratio of the original image is also preserved for subsequent feature interpretation.

[0124] C2: Architecture design of dual-branch convolutional encoder.

[0125] The dual-branch convolutional encoder adopts a specialized design that fully considers the different characteristics of infrared images and vibration time-frequency spectra. The infrared image branch focuses on capturing the spatial temperature distribution pattern, while the time-frequency spectra branch focuses on extracting dynamic features in the time-frequency domain.

[0126] The first convolutional block in the infrared image branch consists of two convolutional layers connected in series, both using a 3×3 convolution kernel. The first layer outputs 32 feature maps, and the second layer outputs 64 feature maps. A batch normalization layer and a ReLU activation function are inserted between the two convolutional layers. This design enables the progressive extraction of multi-level features, from edges to textures. The convolutional block ends with a 2×2 max pooling layer with a stride of 2, which halves the feature map size to 128×128.

[0127] The second convolutional block has a similar structure, but the number of convolution kernels is increased to 128 and 128, which is used to extract higher-level semantic features. Considering the continuity of the temperature distribution, a dilated convolution is introduced in this layer with a dilation rate of 2, which can expand the receptive field without increasing the number of parameters.

[0128] The design of the time-frequency spectrum branch is different. Considering the different physical meanings of the time and frequency dimensions in the time-frequency spectrum, the first convolutional layer uses an asymmetric convolution kernel of size 5×3, which has a larger receptive field in the time dimension. This design can better capture dynamic temporal patterns.

[0129] To preserve the resolution in the frequency dimension, the pooling strategy of the time-frequency spectrum branch has been adjusted. Using a 1×2 pooling kernel, downsampling is performed only in the time dimension, maintaining the resolution in the frequency dimension. This asymmetric pooling preserves more frequency detail.

[0130] Both branches contain three convolutional blocks, ultimately outputting a 32×32×256 feature map. To enhance feature representation, residual connections are introduced within each convolutional block. This is achieved by adjusting the number of channels through a 1×1 convolution and then adding the output to the convolutional block. Residual connections not only alleviate the degradation problem of deep networks but also preserve more original information.

[0131] C3: Implementation of feature fusion strategy.

[0132] Feature fusion is a key step in multimodal learning, which requires effectively combining complementary information from different modalities. This method adopts a multi-level fusion strategy, including a combination of early fusion, mid-term fusion, and late fusion.

[0133] Early fusion occurs after the first convolutional block. The 128×128×64 feature maps output by the two branches are concatenated along the channel dimension to form a 128×128×128 fused feature. This early fusion enables subsequent network layers to learn cross-modal low-level feature associations.

[0134] Mid-term fusion utilizes an attention mechanism. A cross-modal attention module is designed to calculate the correlation between features from both modalities. Specifically, the infrared and time-frequency features are each projected into a 64-dimensional space using a 1×1 convolution. Their matrix product is then calculated to generate a 64×64 attention matrix. After softmax normalization, this attention matrix is ​​used to weight the original features, achieving selective feature fusion.

[0135] Late fusion is performed before the fully connected layer. After global average pooling, the feature maps of the two branches each produce a 256-dimensional feature vector. In addition to simple concatenation, a gated fusion mechanism is also designed. A small neural network (two fully connected layers, 64 neurons in the hidden layer) learns fusion weights and dynamically adjusts the contribution of the two modal features.

[0136] To further enhance the fusion effect, a modality-missing training strategy was introduced. During training, the input of one modality is randomly masked with a probability of 0.2, forcing the network to learn the feature representation of a single modality. This strategy improves the robustness of the model, enabling it to provide reliable diagnostic results even when data from one modality is missing or of poor quality.

[0137] C4: Construction of deep feature extraction network.

[0138] The fused features are then processed through a deep network to extract higher-level semantic features. The deep network consists of three convolutional blocks and two fully connected layers.

[0139] The fourth convolutional block receives the fused features and uses 256 3×3 convolution kernels with a stride of 2 to simultaneously perform feature extraction and downsampling. To prevent information loss, a feature pyramid structure is used before downsampling. Feature maps of different scales are unified to the same size through upsampling and downsampling, and then added together.

[0140] The fifth convolutional block introduces the concept of grouped convolution, dividing the 256 channels into eight groups, with each group undergoing independent convolution. This design not only reduces the number of parameters but also enables the learning of more diverse feature representations. After the grouped convolution, a 1×1 convolution is performed to mix the channels and restore information exchange between them.

[0141] The sixth convolutional block is the last convolutional layer of the network, outputting 512 feature maps. A larger convolution kernel (5×5) is used in this layer, along with a dilated convolution (with a dilation rate of 2) to ensure a sufficiently large receptive field to capture global pattern information.

[0142] The design of the fully connected layers takes overfitting into account. The first fully connected layer contains 1024 neurons and uses dropout (dropout rate 0.5) for regularization. The second fully connected layer contains 256 neurons and also uses dropout (dropout rate 0.3). The activation function uses PReLU, whose negative slope is learned through backpropagation to adapt to different data distributions.

[0143] The final feature vector is normalized by L2 to ensure that its modulus is 1. This normalization strategy makes the feature vector fall on the unit hypersphere, which is convenient for subsequent similarity calculation and cluster analysis.

[0144] In an optional implementation, feature fusion can also use tensor decomposition. The feature maps of the two modalities are treated as third-order tensors (height × width × channels), and the interaction patterns between the modalities are learned through Tucker decomposition or CP decomposition. The core tensor obtained by decomposition contains cross-modal correlation information and can be used as part of the fusion feature.

[0145] In another alternative implementation, a graph convolutional network (GCN) can be introduced for feature fusion. The multimodal features at each spatial location are treated as nodes in a graph, edges are constructed based on feature similarity, and graph convolution operations are then used to propagate and fuse information. This approach is particularly suitable for handling uneven feature distribution.

[0146] It should be noted that the design of a convolutional encoder requires a balance between feature expressiveness and computational efficiency. Using techniques such as depthwise separable convolution and grouped convolution, the number of parameters can be significantly reduced while maintaining performance. Experiments show that the optimized network parameters are reduced by approximately 40%, while diagnostic accuracy only decreases by less than 1%.

[0147] In the embodiment of the present application, step S400 is based on the Dempster-Shafer evidence theory to fuse the classification outputs of the numerical modal depth features and the spectral modal depth features, including the following steps D1-D5:

[0148] D1: Fusion based on Dempster-Shafer evidence theory includes:

[0149] The outputs of the stacked autoencoder network and the stacked convolutional autoencoder network are used as two independent evidence sources; the basic probability distribution function corresponding to each evidence source is calculated; the degree of conflict between the evidence is calculated through evidence synthesis rules; and the multi-source evidence is synthesized based on the conflict degree to obtain the final fault diagnosis result.

[0150] Before evidence fusion is performed, the reliability of each source of evidence needs to be assessed. Several indicators are designed to quantify the quality of evidence.

[0151] The entropy value reflects the uncertainty of the prediction. For a probability distribution p, its entropy value is:

[0152] H(p)=-Σp i log(p i )

[0153] The lower the entropy value, the more certain the prediction. However, too low an entropy value may indicate overfitting, so a reasonable entropy value range of [0.1, 2.0] is set as the standard for reliable evidence.

[0154] The consistency metric measures the degree of consistency between the current prediction and historical predictions. A sliding window is maintained, storing the 10 most recent predictions. The average KL divergence between the current prediction and the historical predictions is calculated. If the divergence is too large, it indicates a possible anomaly and the weight of this evidence should be reduced.

[0155] Modality-specific reliability assessments take data quality into account. For numerical modalities, if some sensor data is missing or outside the normal range, the reliability weight of that evidence is reduced accordingly. For imaging modalities, image quality is assessed by calculating image clarity metrics (such as the Laplacian variance), with blurry or dark images given a lower weight.

[0156] Conflict measurement and management.

[0157] The conflict factor K in DS evidence theory reflects the degree of contradiction between different evidence sources. For two BPA functions m1 and m2, the conflict factor is calculated as:

[0158]

[0159] When K is close to 1, it indicates that the two sources of evidence are in serious conflict, and directly applying Dempster's combination rule may lead to counterintuitive results. Therefore, a conflict management strategy is needed.

[0160] A conflict-based evidence correction method is used. When K > 0.8, it is considered that there is a serious conflict and the evidence needs to be corrected. The correction strategy includes:

[0161] Reliability weighting: Based on the reliability weights calculated previously, the BPA is weighted averaged to reduce the impact of unreliable evidence.

[0162] Conflict distribution: Redistribute the conflicting part K according to certain rules. Using the PCR5 (Proportional Conflict Redistribution) rule, the conflict is proportionally distributed back to the focal element that caused the conflict.

[0163] Evidence discount: Introduce a discount factor β∈[0,1] to discount the evidence:

[0164] m′(A)=β×m(A), m′(Θ)=1-β+β×m(Θ)

[0165] The discount factor is dynamically adjusted according to the degree of conflict: β = exp(-λK), where λ is a parameter that controls the discount strength.

[0166] Evidence synthesis process.

[0167] After conflict management, the modified Dempster combination rule is used for evidence synthesis. For two BPA functions m1 and m2, the synthesis result is:

[0168]

[0169] In order to improve the computational efficiency, the matrix form is used to realize evidence synthesis. Construct the credibility matrix M, where M[i,j]=m1(A i )×m2(A j ) quickly computes all possible intersections and corresponding probability masses through matrix operations.

[0170] When multiple sources of evidence need to be integrated (e.g., by introducing additional sound signal analysis results), a recursive synthesis approach is used. First, the first two pieces of evidence are synthesized to obtain an intermediate result, which is then combined with the third piece of evidence, and so on. The order of synthesis is determined by the reliability weight of the evidence, with the most reliable evidence being prioritized.

[0171] D2: The method further includes:

[0172] The stacked autoencoder network is trained using a composite loss function containing a marginal Fisher regularization term. The stacked convolutional autoencoder network is optimized using a layer-by-layer greedy training method. The loss function is minimized using the gradient descent algorithm to obtain the optimal model parameters.

[0173] Based on the synthesized BPA, decision rules need to be formulated to output the final diagnosis results. A combination of multiple decision criteria is used:

[0174] Maximum reliability criterion: Select the hypothesis with the highest reliability as the diagnostic result. Reliability is calculated as:

[0175]

[0176] This criterion favors conservative decision making.

[0177] Maximum likelihood criterion: Select the hypothesis with the highest likelihood. Likelihood is calculated as:

[0178]

[0179] This criterion considers all possible evidence supporting the hypothesis.

[0180] Compromise criterion: Define the decision function:

[0181] D(A)=αBel(A)+(1-α)Pl(A)

[0182] Where α is a compromise parameter, which is optimized by the validation set. When α = 0.5, it is equivalent to making a decision using the midpoint of the confidence interval.

[0183] In addition to outputting the most likely fault type, the system also provides a confidence interval (Bel(A), Pl(A)) and uncertainty measure for the diagnosis result. When the maximum confidence is below a threshold (e.g., 0.6) or the uncertainty is above a threshold (e.g., 0.3), the system outputs "diagnosis uncertain" and recommends more detailed testing or manual confirmation.

[0184] To provide richer diagnostic information, the top three possible fault types and their reliability values ​​are output. At the same time, by analyzing the contribution of each evidence source, the key features leading to the diagnosis result are identified to support maintenance decisions.

[0185] In an optional implementation, evidence fusion can also employ fuzzy DS theory. Expanding clear fault types into fuzzy sets, such as "mild overheating" and "moderate overheating," can better handle the uncertainty of fault severity. Fuzzy membership functions are learned from expert knowledge or historical data.

[0186] Another alternative implementation involves time-series evidence fusion. This approach not only considers current evidence but also combines historical diagnostic results, fusing time-series information using dynamic DS theory. This approach is particularly well-suited for capturing fault development trends.

[0187] It should be noted that the application of DS evidence theory requires the proper setting of various parameters and thresholds. These parameters are optimized using a large number of historical failure cases and regularly updated based on new failure data. The system also retains a manual intervention interface, allowing experts to adjust the fusion strategy based on their experience.

[0188] In summary, the present invention realizes intelligent diagnosis of transformer faults by constructing a complete multimodal deep learning framework. IG-CWT is innovatively introduced to process vibration signals, a special stacked convolutional autoencoder is designed to process graph data, and reliable decision fusion is achieved through DS evidence theory. The entire system not only improves the diagnostic accuracy, but also provides rich uncertainty information, providing strong support for transformer operation and maintenance decisions. Through actual application in multiple substations, the effectiveness and practicality of this method have been demonstrated, with a diagnostic accuracy of over 96.5%, an increase of 15 percentage points compared to traditional methods, providing important guarantees for the safe and stable operation of the power system.

[0189] Example 3

[0190] Reference Figures 1-6 , which is the third embodiment of the present invention.

[0191] This paper uses five detection devices: oil chromatography gas acquisition module, infrared thermal imager module, vibration acquisition module, ultra-high frequency detection module and dielectric loss acquisition module. The oil chromatography parameter acquisition module is used to collect gas and water content in the converter oil; the infrared thermal imager module is used to collect the temperature of key parts of the converter; the vibration acquisition module collects vibration data through the acceleration sensor installed in the transformer; the ultra-high frequency detection module collects ultra-high frequency electromagnetic waves to convert them into partial discharge data; the dielectric loss acquisition module collects dielectric loss factors that reflect the insulation loss of the bushing. Collect historical data collected by m monitoring devices (one data represents a collection of multiple attributes), recorded as where X i Indicates the information category collected by each collection module, the sample size is n, and each collection module can contain different numbers of attributes d1, d2, ..., d m (For example, the volume fraction of different dissolved gases, partial discharge, bushing dielectric loss factor, and spectrum are different attributes.) According to the Dempster-Shafer synthesis rule described later, transformer faults are classified into {H1, H2, H3, ..., H7}, where H1-H7 represent low-temperature overheating, medium-low-temperature overheating, medium-temperature overheating, high-temperature overheating, low-energy discharge, high-energy discharge, and normal state, respectively.

[0192]

[0193]

[0194] Table 1 Information source data

[0195] Normalize all attributes R collected in numerical format (including vibration data): Assume R min and R max are the minimum and maximum values ​​of attribute R, respectively, and an original value x of R R Through normalization, the data is mapped to a value in the interval [0,1]. The mathematical formula is: New data = (original data - minimum value) / (maximum value - minimum value). This is also called deviation normalization. It is a linear transformation of the original data so that the resulting value is mapped to the interval [0,1]. The conversion function is as follows:

[0196]

[0197] Perform histogram equalization on the collected infrared image atlas data R. The specific process is as follows:

[0198] Assume that the infrared image has a total of L gray levels, the original histogram h(r k ) represents the frequency of gray level, r k Indicates the value of the kth gray level:

[0199] h(r k )=count(r k )

[0200] Normalize the histogram to obtain the gray level r k The probability distribution p(r k ), n is the total number of pixels in the image (n = M × N, M and N are the width and height of the image respectively):

[0201]

[0202] For the probability distribution p(r k ) is accumulated to obtain the gray level r k The cumulative distribution function c(r k ):

[0203]

[0204] Using the cumulative distribution function, the original gray level r k Mapped to the new gray level s k , get the balanced grayscale value:

[0205] s k =(L-1)×c(r k )

[0206] Then the gray value r of each pixel in the original image is k Replaced with the new balanced grayscale value s k, thus obtaining the image after histogram equalization.

[0207] The vibration data R collected in numerical format d Perform continuous wavelet transform based on integrated gradient. The specific process is as follows:

[0208] CWT performs an inner product operation on the signal and a set of wavelets, called a wavelet family. The wavelet family is generated by scaling and translation. The mother wavelet is defined as:

[0209]

[0210] Where ψ(t) is a mother wavelet function with good time and frequency characteristics; s is a scaling parameter used to adjust the width of the wavelet; τ is a translation parameter used to move the wavelet on the time axis; is a normalization factor that keeps the wavelet energy consistent under different scaling.

[0211] Calculate the continuous wavelet transform of the signal x(t) and perform wavelet convolution on the signal:

[0212]

[0213] in Original one-dimensional time series; is the complex conjugate of the mother wavelet; after the convolution operation, the signal x(t) will be transformed by the mother wavelet and projected into the two-dimensional time and frequency dimensions.

[0214] By performing feature attribution along the integrated gradient, the relationship between the model's pre-output and input features is calculated:

[0215]

[0216] Where α is a parameter on the integral path, varying from 0 to 1; S is the model prediction function, which maps the input to the probability value of the output category; It is the gradient of the model output with respect to the input features, indicating how a small change in the input affects the model output.

[0217] Improve model stability by reducing the noise explained by smoothing the gradient:

[0218]

[0219] Among them E SG (x) is the smoothed attribution explanation; N is the number of samples, and different numbers of samples are generated after adding noise; g i is the noise vector added to the input x, which satisfies the normal distribution.

[0220] Calculate the frequency importance score W based on the attribution results f :

[0221]

[0222] Use a frequency importance threshold λ to filter important frequencies when:

[0223] W f ≥λ×average(W f )

[0224] Determine that the frequency f is an important frequency; and use the frequency range F range Identify the important frequency range:

[0225] F range =min({f i})~max({f i})

[0226] Finally, the important frequency interval is determined and used as the frequency interval of continuous wavelet transform.

[0227] Stacked autoencoder network model and stacked convolutional autoencoder network model:

[0228] The stacked autoencoder network consists of an input layer, hidden layers, and an output layer. The input layer receives the numerical data of the power transformer, the hidden layer is composed of the hidden layers of multiple trained denoising autoencoder models, and the output layer performs fault classification.

[0229] The pre-processed multi-source normalized numerical data mentioned above is input as the input layer. The hidden layer is the core part of the stacked autoencoder, which is composed of multiple hidden layers of denoising autoencoders. The denoising autoencoder model introduces noise processing on the basis of the autoencoder model, which can enhance the robustness of the model to noise. Each autoencoder consists of an encoding part and a decoding part, which are divided into an input layer, a hidden layer, and an output layer. The encoding part compresses the input data of the input layer into a low-dimensional representation as a hidden layer. The low-dimensional feature representation of these hidden layers contains a lot of potential feature information. The decoding part reconstructs the hidden layer output structure into the original input data. The denoising autoencoder will encode and decode the noisy input data to restore the original data, and then learn the hidden features of the data. Figure 4 shown.

[0230] The stacking process of the hidden layer is to first compress and extract low-dimensional features from the encoding part of the first layer of denoising autoencoder, then remove the decoding part of the denoising autoencoder, and directly use the extracted hidden low-dimensional features as the output of the denoising autoencoder of this layer, and as the input layer of the next layer of denoising autoencoder. The hidden layer of each denoising autoencoder is repeatedly cut in half to form the hidden layer of the stacked autoencoder network.

[0231] The output layer, located at the end of the stacked autoencoder network, implements fault classification. The final denoising autoencoder layer also removes the original decoding component and replaces it with a softmax classifier, converting the neural network calculation results into the probability of power transformer failure.

[0232] The stacked convolutional autoencoder network also consists of an input layer, a hidden layer, and an output layer. Unlike traditional stacked autoencoders, convolutional autoencoders are designed specifically for image data and are more suitable for processing multimodal data with spatial structures. Both the encoding and decoding parts are implemented through convolution operations. Figure 5 As shown;

[0233] In the stacked convolutional autoencoder network, the input layer receives preprocessed infrared image and vibration video image atlas data and feeds it into two different convolutional encoders in the first layer. Each convolutional encoder layer consists of convolutional and pooling layers, which can extract local features and spatial patterns from the input image and compress the high-dimensional image data into low-dimensional feature representations. These feature representations are concentrated in the salient areas of the image and contain key latent feature information. Compared to ordinary denoising autoencoders, convolutional autoencoders can preserve the spatial structure of the data when extracting low-dimensional image information, improving feature extraction. After the two convolutional encoders in the first layer output the feature vectors of the two modalities, a modal fusion layer is added to splice and fuse the features of the infrared image and video image. The fused features are then input into the convolutional autoencoder in the next layer.

[0234] The hidden layer is composed of multiple layers of stacked convolutional encoders, which extract and compress deep image features layer by layer. In each layer, the input image features undergo convolution and pooling operations, outputting a lower-dimensional feature representation. By stacking these layers, the convolutional autoencoder is able to capture the complex characteristics of multimodal data. Each layer of convolutional encoder further extracts and compresses the salient features output by the previous layer. The stacking of hidden layer features gradually forms a deep semantic feature representation.

[0235] When the network reaches the last layer, the hidden layer features are finally fused and compressed before being sent to the output layer for fault classification. The output layer uses the Softmax function to convert the final feature representation into the probability of a power transformer failure.

[0236] Finally, the fault probabilities output by the two network models are synthesized using the Dempster-shafer rule and combined with the results of different stacked autoencoder models under multiple data sources to obtain the fault status of the power transformer.

[0237] Model training process:

[0238] In the stacked autoencoder, the denoising autoencoder of each layer is trained first. The training goal of the denoising autoencoder is to Reconstruct the original data X, the encoding part of each denoising autoencoder will input the noise Mapping to hidden representation H l :

[0239]

[0240] in and are the weight matrix and bias vector of the encoder in the denoising autoencoder, f is the ReLU activation function, and l represents the number of layers in the stacked autoencoder.

[0241] Similarly, the decoder is needed to output the hidden layer of each autoencoder:

[0242]

[0243] in and are the weight matrix and bias vector of the decoder in the denoising autoencoder, and g is the Sigmoid activation function.

[0244] The training is carried out using a layer-by-layer greedy training method. During the training process, the weight matrix and bias coefficient are continuously updated iteratively through the initial weight matrix and the gradient propagated back, so that the optimal weight matrix and bias coefficient can be obtained after the training is completed. The objective function of each layer can be defined as:

[0245]

[0246] Among them ||~|| F Denotes the F-norm. The calculation process of the weight matrix and bias coefficients is based on the optimization goal of minimizing the objective function of the greedy training method. Through optimization methods such as gradient descent, the optimal weight matrix and bias coefficients can be obtained by using the initial weight matrix and the gradient returned by backpropagation.

[0247] The last layer of the stacked autoencoder is the Softmax output layer for fault classification. The output probability formula is:

[0248]

[0249] S i represents the probability corresponding to the result i obtained by the stacked autoencoder network, j represents the number of categories of the result (7 in this embodiment), and e represents a natural constant.

[0250] Then, the gradient descent method is used to train the target function L to obtain the optimal model parameters for online transformer fault detection:

[0251]

[0252] Among them, y i represents the true fault type of the i-th data sample, α1 and β1 are the hyperparameters of the model, obtained by cross-validation method, ω encoder Represents the network parameters in the encoding stage of the stacked autoencoder network. JMF represents the marginal Fisher regularization term:

[0253]

[0254] s w is the intra-class scatter matrix:

[0255]

[0256] Among them, c is the number of categories, u i The mean vector of samples in the i-th class.

[0257] s b is the inter-class scatter matrix:

[0258]

[0259] where N i represents the number of samples in the i-th class, and u is the global mean vector.

[0260] Finally, the stacked autoencoder network is fine-tuned, that is, the gradient descent method is used to optimize it so that the objective function L is minimized to obtain the trained stacked autoencoder network.

[0261] The overall logic of the stacked convolutional autoencoder is similar to that of the stacked convolutional autoencoder. However, the first layer of the stacked convolutional autoencoder consists of two convolutional autoencoders, which extract two types of graph data features respectively. Therefore, the output feature vectors of the two encoders in the first layer are fused in the second layer:

[0262]

[0263] in Indicates that the first layer input is the convolutional autoencoder output of the infrared image, Indicates that the first layer input is the convolutional autoencoder output of the video atlas, and concat() represents feature concatenation.

[0264] Finally, based on the results of the stacked autoencoder and stacked convolutional autoencoder network structures, the transformer faults are defined as Θ{H1, H2, H3, ..., H7}, where H1-H7 represent low-temperature overheating, medium-low-temperature overheating, medium-temperature overheating, high-temperature overheating, low-energy discharge, high-energy discharge, and normal state, respectively. The synthetic probability under different sources is calculated based on the fault probability obtained by the stacked autoencoder model corresponding to multiple different data sources (a total of 2 k Possible), for n basic probability distribution functions m1, m2, ..., m on the recognition framework Θ n , its evidence synthesis formula is:

[0265]

[0266] represents the power set of Θ, that is, the set of all subsets of Θ; m i (H i ) represents the basic probability distribution of different information sources; K is a measure of the degree of conflict between information sources:

[0267]

[0268] The Dempster-shafer synthesis rule is used to combine the results of different stacked autoencoder models under multiple data sources to obtain the final fault state of the power transformer.

[0269] Example 4

[0270] The above is a schematic diagram of a transformer fault diagnosis method based on multimodal deep learning. It should be noted that the technical solution of this transformer fault diagnosis system based on multimodal deep learning and the technical solution of the transformer fault diagnosis method based on multimodal deep learning are based on the same concept. For details not described in detail in the technical solution of the transformer fault diagnosis system based on multimodal deep learning in this embodiment, please refer to the description of the technical solution of the transformer fault diagnosis method based on multimodal deep learning.

[0271] This embodiment also provides a transformer fault diagnosis system based on multimodal deep learning, including:

[0272] Data acquisition module, used to obtain the volume fraction of dissolved gas in transformer oil, partial discharge, bushing dielectric loss factor, vibration data and infrared image data;

[0273] The data processing module is used to perform continuous wavelet transform based on integrated gradient on the vibration data to generate time-frequency spectrum data and extract deep features of numerical modal data through a stacked autoencoder network;

[0274] The graph feature extraction module is used to extract deep features of graph modality data through a stacked convolutional autoencoder network;

[0275] The decision fusion module is used to fuse the classification results of the two modes based on the Dempster-Shafer evidence theory and output the transformer fault diagnosis results.

[0276] This embodiment also provides an electronic device suitable for transformer fault diagnosis based on multimodal deep learning, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the transformer fault diagnosis method based on multimodal deep learning proposed in the above embodiment.

[0277] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements the transformer fault diagnosis method based on multimodal deep learning proposed in the above embodiment.

[0278] The storage medium proposed in this embodiment and the transformer fault diagnosis method based on multimodal deep learning proposed in the above embodiment belong to the same inventive concept. For technical details not fully described in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.

[0279] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general hardware, and of course can also be implemented by hardware. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.

[0280] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A transformer fault diagnosis method based on multimodal deep learning, characterized by: Including obtaining the volume fraction of dissolved gas in transformer oil, partial discharge, bushing dielectric loss factor, vibration data and infrared image data; The vibration data is processed by continuous wavelet transform based on integrated gradient to obtain time-frequency spectrum data; The volume fraction of dissolved gas in oil, the amount of partial discharge, and the dielectric loss factor of the casing are input into multiple denoising autoencoders to extract low-dimensional feature representations, and the low-dimensional feature representations are input into a stacked autoencoder network formed by stacking multiple denoising autoencoders to extract numerical modal depth features; The infrared image data and time-frequency spectrum data are input into two convolutional encoders respectively, and their feature vectors are extracted and concatenated. The fused features are input into a stacked convolutional autoencoder network formed by stacking multiple convolutional encoders to extract the spectrum modality deep features. Based on the Dempster-Shafer evidence theory, the classification outputs of numerical modal depth features and spectral modal depth features are fused to obtain the transformer fault diagnosis results.

2. The transformer fault diagnosis method based on multimodal deep learning according to claim 1, characterized in that: include: Collect numerical data and spectral data during transformer operation. The numerical data includes the volume fraction of H2, C2H2, CH4, C2H6, CO, and C2H4 dissolved gases in the oil, partial discharge, and bushing dielectric loss factor. The spectral data includes vibration data and infrared images. Perform continuous wavelet transform based on integrated gradient on vibration data to generate vibration time-frequency spectrum; Build a stacked autoencoder network to process numerical data and extract deep feature representations through layer-by-layer encoding; A stacked convolutional autoencoder network is constructed to process atlas data. The first layer uses two independent convolutional encoders to process infrared images and vibration time-frequency spectra respectively. The second layer and subsequent layers perform feature fusion and deep extraction. The fault probability distributions output by the two networks are used as basic probability distributions, and the fusion probability is calculated using the evidence synthesis formula; The transformer fault types are determined based on the fused probability distribution, including low-temperature overheating, medium-low-temperature overheating, medium-temperature overheating, high-temperature overheating, low-energy discharge, high-energy discharge and normal state.

3. The transformer fault diagnosis method based on multimodal deep learning according to claim 2, characterized in that: The continuous wavelet transform based on integrated gradient includes: The attribution relationship between output and input features is predicted through the integral gradient calculation model; the smooth gradient method is used to reduce the noise in the attribution interpretation; the importance score of each frequency is calculated based on the attribution results; the important frequency range is determined and a continuous wavelet transform is performed within this range.

4. The transformer fault diagnosis method based on multimodal deep learning according to claim 3, characterized in that: The construction of the stacked autoencoder network includes: Each layer of denoising autoencoder adds noise to the input data and then encodes it to learn the hidden features of the data; the decoding part of each layer of denoising autoencoder is removed, and the hidden layer of the encoding part is retained; the hidden layers of each layer are stacked and connected to form a deep network structure; and a Softmax classifier is used in the last layer to output the fault probability.

5. The transformer fault diagnosis method based on multimodal deep learning according to claim 4, characterized in that: The construction of the stacked convolutional autoencoder network includes: The first layer uses two convolutional encoders to extract the features of the infrared image and time-frequency spectrum respectively; the two feature vectors output by the first layer are spliced ​​and fused; the fused features are input into the subsequent convolutional encoder layer to extract deep features layer by layer; the last layer uses the Softmax function to output the fault probability distribution.

6. The transformer fault diagnosis method based on multimodal deep learning according to claim 5, characterized in that: The fusion based on Dempster-Shafer evidence theory includes: The outputs of the stacked autoencoder network and the stacked convolutional autoencoder network are used as two independent evidence sources; the basic probability distribution function corresponding to each evidence source is calculated; the degree of conflict between the evidence is calculated through evidence synthesis rules; and the multi-source evidence is synthesized based on the conflict degree to obtain the final fault diagnosis result.

7. The transformer fault diagnosis method based on multimodal deep learning according to claim 6, characterized in that: The method further comprises: The stacked autoencoder network is trained using a composite loss function containing a marginal Fisher regularization term. The stacked convolutional autoencoder network is optimized using a layer-by-layer greedy training method. The loss function is minimized using the gradient descent algorithm to obtain the optimal model parameters.

8. A transformer fault diagnosis system based on multimodal deep learning, based on the transformer fault diagnosis method based on multimodal deep learning according to any one of claims 1 to 7, characterized in that: It also includes a data acquisition module for obtaining the volume fraction of dissolved gas in transformer oil, partial discharge amount, bushing dielectric loss factor, vibration data and infrared image data; The data processing module is used to perform continuous wavelet transform based on integrated gradient on the vibration data to generate time-frequency spectrum data and extract deep features of numerical modal data through a stacked autoencoder network; The graph feature extraction module is used to extract deep features of graph modality data through a stacked convolutional autoencoder network; The decision fusion module is used to fuse the classification results of the two modes based on the Dempster-Shafer evidence theory and output the transformer fault diagnosis results.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the transformer fault diagnosis method based on multimodal deep learning are implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the transformer fault diagnosis method based on multimodal deep learning are implemented.

Citation Information

Cited By

  • Thermal fault early warning method of machine room inspection robot with multi-mode perception

    CN120910488A

  • Fault diagnosis method and equipment for washing and selecting vibrating screen equipment based on multi-modal fusion

    CN121524893A

  • Sewage treatment fault diagnosis method and system based on variational mode decomposition and Transform integration

    CN121919694A

  • Transformer fault feature fusion diagnosis method and system based on multi-modal time sequence data reconstruction

    CN121980473A

  • Transformer fault feature fusion diagnosis method and system based on multi-modal time series data reconstruction

    CN121980473B