Transformer fault diagnosis method based on differential decomposition and secondary fusion

By employing differentiated decomposition and two-level fusion methods, KLA-EMD, VMD-EOBL, and SSA-LMD are used to decompose transformer signals. Combined with LFFL and DBN-CP networks, the problems of signal characteristic matching and deep correlation in existing technologies are solved, achieving efficient and interpretable transformer fault diagnosis.

CN121935704APending Publication Date: 2026-04-28HUAIYIN INSTITUTE OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-20
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing transformer fault diagnosis technologies suffer from limitations such as single decomposition algorithms failing to match different signal characteristics, insufficient feature effectiveness, difficulty in uncovering deep correlations in first-level fusion modes, poor interpretability of diagnostic results, low robustness and accuracy, and inability to meet the real-time requirements of integrated "source-grid-load-storage" systems.

Method used

A differentiated decomposition strategy is adopted to decompose multi-source monitoring data through KLA-EMD, VMD-EOBL, and SSA-LMD. The learnable feature fusion layer LFFL and the deep belief network DBN-CP are combined for secondary fusion to achieve feature selection and deep correlation mining, and output fault diagnosis results and multi-source feature contribution analysis.

Benefits of technology

It improves the accuracy, robustness, and efficiency of transformer fault diagnosis, enhances the interpretability of diagnostic results, meets real-time requirements, and is suitable for transformer condition monitoring and fault early warning under complex operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121935704A_ABST
    Figure CN121935704A_ABST
Patent Text Reader

Abstract

The invention provides a transformer fault diagnosis method based on differential decomposition and secondary fusion. The method comprises the following steps: step 1, acquiring multi-source monitoring data of electricity, temperature and working conditions of a transformer; 2, a customized decomposition strategy is adopted for different signal characteristics; step 3, obtaining an intrinsic mode component and carrying out standardization preprocessing; step 4, performing time sequence alignment and matrix splicing on the multi-source intrinsic mode component features; step 5, constructing a secondary fusion mechanism; and step 6, outputting a transformer fault diagnosis result and a multi-source characteristic contribution degree analysis report. According to the method, through differential decomposition of the physical characteristics of the matched signal, secondary fusion reinforcement of feature screening and decision interpretability, the accuracy, robustness and efficiency of fault diagnosis are remarkably improved, and reliable technical support is provided for state operation and maintenance of the transformer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power equipment fault diagnosis technology, and in particular relates to a transformer fault diagnosis method based on differential decomposition and two-level fusion. Background Technology

[0002] The power system is a core infrastructure of the national economy, and transformers, as key equipment for energy conversion and transmission, are crucial to public safety and economic stability. With the development of smart grids, transformer operating conditions are becoming increasingly complex, susceptible to factors such as harmonics, extreme weather, and insulation aging, leading to a rise in fault risks. Transformer faults can easily cause grid paralysis and significant losses; therefore, early and accurate diagnosis and condition warning have become core needs and research hotspots in the field of power equipment fault diagnosis. Current diagnostic technologies have shifted towards online intelligent monitoring, capable of collecting multi-source data such as electrical, temperature, and operating conditions. However, multi-source signals are heterogeneous and complex, easily affected by noise and asynchronous timing, posing challenges to feature extraction and fault identification. Existing technologies have significant shortcomings: single decomposition algorithms cannot match different signal characteristics, resulting in insufficient feature effectiveness; first-level fusion modes struggle to uncover deep correlations, leading to poor interpretability, low robustness, and low accuracy of diagnostic results. Furthermore, the integrated "source-grid-load-storage" system increases the real-time requirements for diagnosis, rendering existing methods inefficient. Summary of the Invention

[0003] Purpose of the invention: The technical problem to be solved by the present invention is to address the shortcomings of the existing technology by providing a transformer fault diagnosis method based on differentiated decomposition and two-level fusion. By matching different signal physical characteristics through customized decomposition strategies and combining a two-level fusion mechanism to enhance feature screening and deep correlation mining, the method achieves a synergistic improvement in fault diagnosis accuracy, robustness and efficiency. At the same time, it quantifies and analyzes the contribution of multi-source features to improve the interpretability of diagnostic results and provides reliable technical support for transformer condition-based operation and maintenance.

[0004] The method includes the following steps:

[0005] Step 1: Synchronously collect multi-source monitoring data of the transformer, including dielectric loss factor and winding DC resistance data, top oil temperature and winding temperature data, load rate and running time data; and bind it with a Coordinated Universal Time (UTC) millisecond-level timestamp.

[0006] Step 2: The multi-source monitoring data collected in Step 1 is decomposed using a differentiated decomposition strategy to obtain the intrinsic mode components of each signal. The specific decomposition methods are as follows: the dielectric loss factor and winding DC resistance signals are decomposed by Kernel Local Alignment-Empirical Mode Decomposition (KLA-EMD); the top oil temperature and winding temperature signals are decomposed by Variational Mode Decomposition-Enhanced Observer-Based Learning (VMD-EOBL); and the load factor and running time signals are decomposed by Singular Spectrum Analysis-Local Mean Decomposition (SSA-LMD).

[0007] Step 3: Perform normalization preprocessing on each intrinsic modal component obtained in Step 2 in sequence;

[0008] Step 4, matrix fusion: The multi-source intrinsic modal component features after standardization and preprocessing are time-series aligned and matrix-concatenated to form a unified multi-source fusion feature matrix;

[0009] Step 5: Construct a two-level fusion mechanism, which completes deep decision fusion and contribution parsing through a learnable feature fusion layer (LFFL) and a deep belief network-contribution parsing (DBN-CP) network;

[0010] Step 6, Output fault diagnosis results: Output transformer fault diagnosis results and multi-source feature contribution analysis report.

[0011] Step 2 includes:

[0012] Step 2.1, Kernel Local Alignment-Empirical Mode Decomposition (KLA-EMD) of Electrical Signals, includes the following steps:

[0013] Step 2.1.1: Fusion of time-varying noise injection and random subspace regression;

[0014] By injecting time-varying noise into the original electrical signal, a balance is achieved between weak feature enhancement and noise suppression. An adaptive ensemble noise injection and random subspace regression fusion framework is employed.

[0015] ,

[0016] in, After noise injection, the first indivual Class of electrical signals, Original electrical signal, It is Gaussian white noise with a mean of 0 and a variance of 1. δ is the time-varying noise figure, δ is the dielectric loss factor, and R is the DC resistance of the winding.

[0017] Step 2.1.2: Divide the dataset into training and validation subsets;

[0018] proportionally The set of extreme points is divided into training and validation subsets:

[0019] ,

[0020] ,

[0021] in, For the first indivual The complete set of minimum points of the signal type, For the first indivual The training subset of complete minimum points of the signal class. For the first indivual A complete set of verification subsets of the minimum points of the signal class; For the first indivual The complete set of maxima of the signal type, For the first indivual The training subset of complete maxima of the signal class. For the first indivual A complete subset of the verification points of the signal class;

[0022] The following partitioning constraints must be met:

[0023] ,

[0024] ,

[0025] ,

[0026] in, For the first indivual Signal type in time amplitude, The sampling interval is... , The time coordinates of the extreme point;

[0027] Step 2.1.3: Kernel polynomial regression and smooth envelope generation;

[0028] Training subsets with maxima Using the input, construct a d-order multinomial regression model:

[0029] ,

[0030] in, In time The predicted amplitude of the fitted envelope is given, with d initially set to order 3-5. The coefficients to be determined; the validation subset will be used. Input a d-th order polynomial regression model, adjust d using a grid search method, and determine the optimal order with the goal of minimizing the mean squared error. :

[0031] ,

[0032] Among them, It is the first The time corresponding to each maximum point coordinate; This corresponds to the corresponding time point. The signal amplitude, where q is the number of samples in the validation subset; based on By fitting the sets of maxima and minima respectively, the upper envelope is obtained. Lower envelope Calculate the rate of change of the first derivative of the envelope; if it is greater than the threshold, add L2 regularization and refit.

[0033] Step 2.1.4: Verification of Intrinsic Mode Function (IMF) using dual criteria;

[0034] The effectiveness of the Intrinsic Mode Components (IMFs) extracted from the envelope generated by kernel multinomial regression and adaptive optimization is verified through quantitative constraints based on dual criteria. The mathematical definition and logic of the dual criteria are as follows:

[0035] Criterion 1, Local Extremum Matching Criterion: Requires the intrinsic mode components (IMF components) to match. The quantization expression is as follows: The difference between the number of upper and lower extreme points does not exceed 1, and the distribution of extreme points within a local region is consistent with the original signal.

[0036] ,

[0037] ,

[0038] in For the intrinsic mode components IMF components The total number of maxima. For the intrinsic mode components IMF components The total number of local minimum points; The preset threshold; Let be any local time ν.

[0039] Criterion 2, Kernel Fitting Residual Constraint Criterion: Requires the intrinsic mode components and IMF components to... Residual with the original signal Satisfy the kernel fitting residual threshold constraint, where The original electrical signal, The intrinsic mode components (IMFs) obtained from the decomposition of the original signal in time The numerical value, quantized as:

[0040]

[0041] in, For residual signal In terms of time length The mean square value, where T is the signal duration. The residual threshold;

[0042] Both criteria must be met simultaneously; otherwise, return to step 2.1.3 to readjust the kernel polynomial order or regularization weights.

[0043] Step 2.2, Variational Mode Decomposition of Temperature Signal - Enhanced Observer Learning VMD-EOBL Decomposition, includes the following steps:

[0044] Step 2.2.1: Define a multi-objective optimization problem that is hierarchical in terms of physical meaning;

[0045] We innovatively construct a fitness function that integrates four criteria, transforming Variational Mode Decomposition (VMD) into a physically interpretable optimization problem:

[0046] ,

[0047] in, For multi-objective fitness functions, The parameters to be optimized include the number of modes. Punishment factor Time scale parameter τ; These are the fundamental oscillatory components with physical meaning obtained after variational mode decomposition (VMD) of the temperature signal. This refers to the k-th fundamental oscillation component obtained after variational mode decomposition of the VMD temperature signal. Characterizes signal fidelity. The function is used to evaluate component sparsity. The function is used to measure the independence of components. The function passes through each fundamental oscillation component The similarity to the preset physical process template library L is used to quantify the degree of matching of physical meaning; , , , These are dynamic weighting coefficients;

[0048] Step 2.2.2: Enhanced observer learns EOBL algorithm parameters training;

[0049] We design an enhanced observer learning EOBL algorithm with a prior experience pool. Historical high-quality parameters are introduced during population initialization to accelerate convergence. The individual position update formula is:

[0050] ,

[0051] in, For individuals In the The new position of the era For individuals In the The position of the substitute (i.e., the current parameter combination, such as...) ), This is the globally optimal solution. The function analyzes population distribution Gradient direction prediction with experience pool For adaptive noise, intermediate parameters Dynamically adjusted with iteration; combined with elite retention before Early stop strategy for individual fitness plateaus to improve training efficiency;

[0052] Step 2.2.3: After decomposition, physical labels are automatically bound to the basic oscillation components IMF, generating feature vectors with clear physical semantics;

[0053] Step 2.3, Singular Spectrum Analysis - Local Mean Decomposition (SSA-LMD) of Operating Condition Signals; includes the following steps:

[0054] Step 2.3.1: Singular Spectrum Analysis of Physical Constraints - SSA Trend-Residual Separation;

[0055] Introduce a constraint function based on operational physics knowledge. Used to evaluate each specific and complete oscillation pattern, trend, or noise component. Physical rationality; definition of physical fit index :

[0056] ,

[0057] in, For similarity function: calculate the reconstructed components The degree of matching with the preset physical process template, Physical constraint function: Domain knowledge-based evaluation Physical rationality; maximizing trend components through constraint optimization. Minimize residual components Output trend terms with clear physical meaning. With residuals ;

[0058] Step 2.3.2: Parameter-adaptive Local Mean Decomposition (LMD) fine decomposition;

[0059] The residual term R(t) from the first-level singular spectral analysis (SSA) trend-residual separation output is input into the local mean decomposition (LMD) for fine decomposition.

[0060] Local Mean Decomposition (LMD) iteratively decomposes the signal into a product function by calculating local extrema, the local mean function, and the envelope estimation function, thus constructing a bi-objective optimization function.

[0061] ;

[0062] in, The overall objective function (loss function) for optimizing Local Mean Decomposition (LMD). The set of parameters to be optimized for Local Mean Decomposition (LMD) For dynamic weights, To decompose the error term:

[0063] ;

[0064] in The first is obtained from the Local Mean Decomposition (LMD) decomposition. Product functions, These are weighting coefficients used to balance reconstruction error and smoothness requirements. To obtain the maximum value of the coefficient of variation among all PF components of the product function, i.e., to focus on the least smooth envelope, To penalize the envelope function in the PF component of the product function, For the physical matching degree item:

[0065] ;

[0066] in, Indicates taking the first The product function PF components and all templates The maximum similarity among them. A template library for physical microprocesses. The time-frequency coherence coefficient, It is the entropy value;

[0067] Step 2.3.3: Implement adaptive collaborative optimization of two-layer parameters;

[0068] Define optimization parameters , This is a subset of parameters to be optimized for the SSA decomposition layer in singular spectral analysis. For the subset of parameters to be optimized in the Local Mean Decomposition (LMD) layer, construct a comprehensive fitness function. :

[0069] ,

[0070] in, To ensure fidelity, For complexity terms, The sum of the physical interpretability scores. , , These are the weighting coefficients;

[0071] The outer ring employs an improved swarm intelligence algorithm to globally search the subset of parameters to be optimized in the Singular Spectrum Analysis (SSA) decomposition layer. The inner loop optimizes the subset of parameters to be optimized in the Local Mean Decomposition (LMD) layer by gradient-guided search. This achieves dual-level parameter coupling optimization; finally, it extracts the statistical and time-frequency features of the trend term and the PF component of the product function to form a hierarchical differentiated feature vector.

[0072] Step 3 includes:

[0073] The intrinsic modal components obtained from step 2 are subjected to normalization preprocessing in sequence, resulting in three types of normalized feature matrices:

[0074] , , ,

[0075] in, For the normalized feature matrix of electrical signals, The normalized feature matrix of the temperature signal, Here, m is the standardized feature matrix of the operating condition signal. As an electrical characteristic dimension, Temperature feature dimension This refers to the dimension of working condition characteristics.

[0076] Step 4 includes:

[0077] Based on a unified UTC millisecond-level timestamp, the three types of standardized feature matrices are first time-series aligned and verified. When the missing rate is less than or equal to the threshold, linear interpolation is used to complete the matrix. Then, the aligned three types of matrices are concatenated column by column to form a multi-source fusion feature matrix. Finally, outliers are removed by the 3σ criterion and the overall missing rate is verified to be less than or equal to the threshold. A qualified multi-source fusion feature matrix is ​​output as the input for the secondary fusion mechanism in step 5.

[0078] Step 5 includes:

[0079] The original high-dimensional feature set extracted from the transformer multi-source monitoring data after preliminary differential decomposition and multi-source fusion constitutes the sample feature matrix. , where n is the total dimension of the original features;

[0080] Step 5.1, First-level fusion: Learnable Feature Fusion Layer (LFFL);

[0081] Step 5.1.1: Multi-source feature integration and parameter initialization;

[0082] Feature selection gating vector ,pass function Parameterization:

[0083] ,

[0084] in, It is a natural constant. For the first One trainable basic parameter component Characterizing the first The retention strength of each original feature It achieves soft selection of features, and the entire computation process is differentiable, supporting gradient backpropagation.

[0085] Nonlinear fusion projection weights: projection matrix With bias vector , where d is the feature dimension after fusion, used to map the filtered features to a low-dimensional semantic space, d=n;

[0086] Step 5.1.2: Gated weighting and nonlinear transformation;

[0087] The Learnable Feature Fusion Layer (LFFL) performs the following differentiable computation sequence on the input features to achieve a fusion of feature importance weighting and dimensionality compression:

[0088] Gated weighted modulation: the gate vector Broadcast to feature matrix Perform element-wise Hadamard product in the same dimension. To achieve feature-level soft selection:

[0089] ,

[0090] in, It is a column vector with all elements being 1; H represents the transpose; Indicates the first The first sample Each feature channel has a global weight. modulation;

[0091] Nonlinear fusion projection: The weighted features are subjected to an affine transformation and nonlinear activation, mapping them to a low-dimensional fusion space.

[0092] ,

[0093] in, The features are after gating and weighting. It is a non-linear activation function. This is the projection weight matrix (learnable parameters). This is the transpose of the bias vector. The dimensionality-reduced and filtered feature representations of the first-level fusion output are used as inputs for the second-level fusion.

[0094] Step 5.2, Second-level fusion: Deep Belief Network - Contribution Analysis DBN-CP;

[0095] Step 5.2.1: Establish the Deep Belief Network-Contribution Analysis (DBN-CP) network structure;

[0096] Suppose that the Deep Belief Network-Contribution Parsing (DBN-CP) network is pre-trained by stacking L Restricted Boltzmann Machines (RBMs) and then fine-tuned using supervised signals. The top-level design of the DBN-CP network is a two-branch structure, including:

[0097] Classification branch: Output the probability distribution of fault categories:

[0098] ,

[0099] in, It is the weight matrix of the classification layer. This is the bias vector of the classification layer, where C is the number of fault categories. The activation value of the last hidden layer, after... After processing, To become a probability vector;

[0100] Contribution analysis branch: This is a fully connected layer that runs parallel to the classification branch, used to analyze high-level semantic representations. The contribution estimate to the original n-dimensional features is reconstructed from the data:

[0101] ,

[0102] in, The weight matrix for the contribution analysis branch. The bias vector for the contribution analysis branch. For the Sigmoid function, the output is... This can be interpreted as a normalized contribution score vector;

[0103] Step 5.3: Construct contribution alignment constraints between the two levels;

[0104] Deep Belief Network - Contribution Analysis (DBN-CP) - Contribution Analysis The gating weights actually used in the LFFL layer, which is a fusion layer of learnable features. The characteristics they imply are of consistent importance:

[0105] ,

[0106] in, To use global weights Based on the benchmark, measure the performance of the first Explanation of the contribution of each sample The amount of additional information it carries; For the explanation of a single sample Based on this, measure the global weight. The amount of additional information it carries; The loss value specifically refers to the feature selection weights used to measure and minimize the "first-level feature fusion layer" during model training. "Contribution analysis output of the second-level decision fusion layer" The loss value of the difference between "".

[0107] ;

[0108] in For the first The degree of contribution of each feature to the sample decision, To give the first The importance weights of each feature;

[0109] Step 5.4, Joint Loss Function and Optimization;

[0110] Total loss function It integrates classification objective, contribution alignment constraint, and feature selection sparsity regularization:

[0111] ,

[0112] in, It is the standard cross-entropy classification loss; The probability distribution of fault categories predicted by the model; Labels for actual faults; These are the weighting coefficients for the alignment loss; It is the L1 regularization coefficient; for Norm, acting on the gated vector This promotes sparsification of feature selection and automatically filters redundant features;

[0113] Step 5.5, backpropagation and parameter update.

[0114] Step 5.5 includes: all parameters Joint updates via stochastic gradient descent or its variants include the following key gradient flows:

[0115] The classification loss gradient is backpropagated from the top of the Deep Belief Network-Contribution Parsing (DBN-CP) to the bottom layer, and further propagated through Z to the parameters of the Learnable Feature Fusion (LFFL) layer. ;

[0116] Alignment loss Two gradient paths are generated:

[0117] A parameter applied to the contribution analysis branch ;

[0118] Another one directly affects the gating parameters , about gradient for:

[0119] .

[0120] in, For alignment loss on the first Gating basic parameters of each feature The partial derivative (gradient) of . For the first The first sample The contribution estimate of each feature.

[0121] In step 6, the output results include the specific fault type of the transformer and a multi-source feature contribution analysis report. The contribution analysis report clarifies the contribution ratio of various electrical, temperature, and operating condition features to the fault diagnosis results.

[0122] The present invention also provides an electronic device, including a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method.

[0123] The present invention also provides a storage medium storing a computer program or instructions that, when the computer program or instructions are run on a computer, execute the steps of the method described.

[0124] Beneficial Effects: Compared with the problems existing in actual transformer fault diagnosis, existing methods still face some challenges. Therefore, this method comprehensively considers the physical characteristics of various signals and adopts a targeted and customized decomposition algorithm. Compared with the above-mentioned single decomposition algorithm, which has a narrow scope of application and is difficult to apply to various types of signals, this method can solve these problems, so that the extracted features can truly reflect the state information of the transformer, thus laying a good foundation for subsequent diagnosis.

[0125] This invention creates multi-source monitoring signals for transformers, each with its own inherent physical properties. Electrical signals (dissipation factor, winding DC resistance) exhibit high-frequency oscillations, temperature signals (top oil temperature, winding temperature) show slow drift, and operating condition signals (load rate, running time, etc.) are constantly changing due to load influence. To address these characteristics, three customized decomposition algorithms—KLA-EMD, VMD-EOBL, and SSA-LMD—were invented. Other decomposition algorithms are general-purpose and cannot meet the physical characteristics of various signals; this is a key focus of this invention. Using targeted decomposition avoids feature redundancy and distortion, resulting in data that accurately captures information on the most critical points of transformer operating status, forming highly pure feature values ​​and ensuring accurate fault diagnosis in subsequent stages.

[0126] This invention proposes a two-level fusion structure model based on a "feature layer + decision layer," which enables efficient integration and deep mining of multi-source features. The feature layer employs a learnable feature fusion layer (LFFL) with a gated weighting mechanism to adaptively select key features useful for fault diagnosis based on different application scenarios. Its nonlinear dimensionality reduction eliminates redundant and interfering information. The decision layer, based on a DBN-CP deep network, mines deep-level connections and semantic information among multi-source features. It improves feature utilization efficiency by combining contribution alignment constraints and sparse regularization, avoiding the drawback of "shallow integration failing to extract internal connections of features." This significantly improves the accuracy and robustness of fault diagnosis and is applicable to various complex operating conditions such as power grid harmonics, load, and extreme weather.

[0127] This invention employs a two-stage collaborative approach, significantly improving diagnostic efficiency. Nonlinear dimensionality reduction at the feature layer compresses feature dimensions, increasing model optimization speed while reducing model data volume. Simultaneously, a joint loss function and gradient backpropagation algorithm are used to collaboratively optimize the parameters of the Learnable Feature Fusion Layer (LFFL) and the DBN-CP-based deep network, overcoming model redundancy and computational inefficiency caused by traditional two-stage training methods. On one hand, it achieves dual optimization of the Learnable Feature Fusion Layer (LFFL) and the DBN-CP-based deep network, thereby greatly reducing model computational complexity; on the other hand, it significantly reduces the diagnostic time per sample, effectively meeting the low-latency requirements of online real-time transformer monitoring. It is better suited for continuous operation and maintenance under complex conditions, providing excellent support for transformer condition monitoring and fault early warning. Attached Figure Description

[0128] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments, and the advantages of the present invention in the above and / or other aspects will become clearer.

[0129] Figure 1 This is a schematic diagram of the KLA-EMD decomposition process for local alignment of electrical signal nuclei provided by the present invention.

[0130] Figure 2 This is a schematic diagram of the temperature signal variational mode decomposition-enhanced observer learning VMD-EOBL decomposition process provided by the present invention.

[0131] Figure 3 A schematic diagram of the Singular Spectrum Analysis-Local Mean Decomposition (SSA-LMD) process for operating signals provided by this invention.

[0132] Figure 4 This is a schematic diagram of the first-level learnable feature fusion layer (LFFL) provided by the present invention.

[0133] Figure 5 This is a schematic diagram of the second-level deep belief network-contribution resolution fusion (DBN-CP) process provided by the present invention.

[0134] Figure 6 A schematic diagram of the collaborative optimization mechanism provided by this invention.

[0135] Figure 7 This is a schematic diagram of the overall system flow provided by the present invention. Detailed Implementation

[0136] This embodiment provides a transformer fault diagnosis method based on differential decomposition and two-level fusion, such as... Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 As shown, it includes:

[0137] Step 1: Synchronously collect multi-source monitoring data of the transformer, including dielectric loss factor and winding DC resistance data, top oil temperature and winding temperature data, load rate and running time data; and bind a unified UTC millisecond-level timestamp.

[0138] Step 2: The multi-source monitoring data collected in Step 1 is decomposed using a differentiated decomposition strategy to obtain the inherent mode components of each signal. The specific decomposition method is as follows: the dielectric loss factor and winding DC resistance signals are decomposed by KLA-EMD, the top oil temperature and winding temperature signals are decomposed by VMD-EOBL, and the load rate and running time signals are decomposed by SSA-LMD.

[0139] Step 3: Perform normalization preprocessing on each intrinsic modal component obtained in Step 2 in sequence;

[0140] Step 4: Matrix fusion. The multi-source intrinsic modal component features after standardization and preprocessing are time-series aligned and matrix-concatenated to form a unified multi-source fusion feature matrix.

[0141] Step 5: Construct a two-level fusion mechanism, which uses a learnable feature fusion layer to achieve gated weighting and nonlinear dimensionality reduction, and completes deep decision fusion and contribution analysis based on the DBN-CP network;

[0142] Step 6: Output fault diagnosis results: Output transformer fault diagnosis results and multi-source feature contribution analysis report.

[0143] Step 1 includes:

[0144] Multi-source monitoring data of the transformer is collected synchronously, specifically covering three types of monitoring data with significant differences in characteristics: electrical data (dielectric loss factor, winding DC resistance), temperature data (top oil temperature, winding temperature), and operating condition data (load rate, operating time). Based on the different physical characteristics and characteristic distribution patterns of the three types of data, data support is provided for the implementation of subsequent differentiated decomposition strategies.

[0145] Step 2 includes the following steps:

[0146] KLA-EMD decomposition of electrical signals:

[0147] The KLA (Kernel+Local Adaptive) core framework is defined as a collaborative system of "kernel method to enhance nonlinear fitting ability - local adaptive adaptation to nonstationary features".

[0148] Step 2.1.1: Fusion of time-varying noise injection and random subspace regression;

[0149] The KLA framework's "adaptive data layer enhancement and kernelization preprocessing" is not a simple dynamic parameter adjustment, but a closed-loop optimization based on the local statistical characteristics of the signal:

[0150] By injecting time-varying noise into the original electrical signal, a balance is achieved between weak feature enhancement and noise suppression. An adaptive ensemble noise injection and random subspace regression fusion framework is employed.

[0151] ,

[0152] in, Original electrical signal, It is Gaussian white noise with a mean of 0 and a variance of 1. The noise figure is a time-varying value that is dynamically adjusted based on the local signal-to-noise ratio of the signal to achieve a balance between weak feature enhancement and noise suppression.

[0153] Step 2.1.2: Partitioning the extreme point set into training / validation subsets;

[0154] The KLA framework emphasizes that "the validity of local data takes precedence over global coverage," and this step provides a dataset that conforms to local characteristics for subsequent kernel fitting.

[0155] proportionally The extreme point set is divided into training / validation subsets, and the specific division formula is as follows:

[0156] ,

[0157] ,

[0158] And it satisfies the partitioning constraints:

[0159] ,

[0160] In the formula:

[0161] ,

[0162] ,

[0163] For a set of local extrema, The sampling interval is... The time coordinates are the extreme points.

[0164] Step 2.1.3: Kernel Multinomial Regression and Smooth Envelope Generation (Kernel Enhancement + Local Adaptation)

[0165] It covers two major dimensions: "kernel function nonlinear enhancement" and "local adaptive parameter tuning".

[0166] Training subsets with maxima (Time coordinates) Amplitude Using as input, construct a d-order multinomial regression model:

[0167] ,

[0168] Where d is initially set to order 3-5. The coefficients to be determined; the validation subset will be used. Input the model, adjust d (range 2-8) using the grid search method, and determine the optimal order with the goal of minimizing the mean square error. :

[0169] ,

[0170] q represents the number of samples in the validation subset; based on By fitting the set of maxima M and the set of minima N respectively, we can obtain the upper and lower envelopes. , Calculate the rate of change of the first derivative of the envelope. If it is greater than 0.5, add L2 regularization (weight: 0.01-0.05) and refit to ensure that the envelope is smooth and free of overshoot.

[0171] Step 2.1.4: IMF dual criterion verification;

[0172] The effectiveness of the IMF extracted from the envelope generated by kernel multinomial regression and adaptive optimization is verified through quantitative constraints based on dual criteria, ensuring that the decomposition results can truly reflect the inherent modal characteristics of electrical signals. The mathematical definition and logic of the dual criteria are as follows:

[0173] Criterion 1 (Local Extremum Matching Criterion): Requires IMF components If the difference between the number of upper and lower extreme points does not exceed 1, and the distribution of extreme points in the local region is consistent with the original signal, the quantization expression is:

[0174] ,

[0175] and ; , These represent the number of maximum and minimum points, respectively. For any local time window;

[0176] Criterion 2 (Kernel Fitting Residual Constraint Criterion): Requires the residuals of the IMF and the original signal to be... Satisfying the kernel fitting residual threshold constraint, the quantization expression is:

[0177] ,

[0178] For signal duration, (The residual threshold is adaptively determined by the amplitude range of the original signal). Both criteria must be met simultaneously; otherwise, return to step 2.1.3 to readjust the kernel polynomial order or regularization weights, forming a closed-loop logic of "fitting-verification-optimization" to ensure the reliability and stability of KLA-EMD decomposition.

[0179] VMD-EOBL decomposition of temperature signal;

[0180] Step 2.2.1: Define a multi-objective optimization problem that is hierarchical in terms of physical meaning;

[0181] We innovatively construct a fitness function that integrates four criteria, transforming VMD decomposition into a physically interpretable optimization problem:

[0182] ,

[0183] in For the parameters to be optimized (number of modes) Punishment factor (Time scale parameter τ) Characterizes signal fidelity. Assess component sparsity. To measure the independence of components, Through each The similarity to the preset physical process template library L is used to quantify the degree of matching of physical meaning; These are dynamic weighting coefficients.

[0184] Step 2.2.2: Dual-observer EOBL parameter training;

[0185] An improved EOBL algorithm with a prior experience pool is designed. Historical high-quality parameters are introduced during population initialization to accelerate convergence. The formula for updating individual positions is:

[0186] ,

[0187] In the formula, This is the globally optimal solution. By analyzing population distribution Gradient direction prediction with experience pool For adaptive noise, Dynamically adjusted with iteration; combined with elite retention (previous) Individuals and early stop strategies during fitness plateaus can improve training efficiency.

[0188] Step 2.2.3: After decomposition, the IMF is automatically bound with physical labels such as "long-term aging" and "load cycle" to generate feature vectors with clear physical semantics.

[0189] SSA-LMD decomposition of operating condition signals;

[0190] Construct a two-layer decomposition model structure guided by physical constraints;

[0191] Step 2.3.1: Improved SSA trend of physical constraints - residual separation;

[0192] Introduce a constraint function based on operational physics knowledge. Used to evaluate each Physical rationality. Define physical fit index:

[0193] ,

[0194] Maximize the trend component through constraint optimization. Minimize residual components Output trend terms with clear physical meaning. With residuals .

[0195] Step 2.3.2: Parameter-adaptive LMD fine decomposition;

[0196] The residual term R(t) from the first layer output is input into the Local Mean Decomposition (LMD) model for fine decomposition. LMD iteratively decomposes the signal into multiple product functions by calculating local extrema, local mean functions, and envelope estimation functions, constructing a bi-objective optimization function that balances decomposition accuracy and physical time-frequency interpretability.

[0197] ;

[0198] Decompose error terms: ;

[0199] Physical matching degree item: ;

[0200] For dynamic weights, A template library for physical microprocesses. The time-frequency coherence coefficient, This is the entropy value.

[0201] Step 2.3.3: Implement adaptive collaborative optimization of two-layer parameters;

[0202] Define optimization parameters Construct the comprehensive fitness function:

[0203]

[0204] The outer ring employs an improved swarm intelligence algorithm for global search. The inner loop optimizes the search through gradient guidance. This achieves dual-level parameter coupling optimization; finally, it extracts the statistical and time-frequency features of the trend term and PF components to form a hierarchical differentiated feature vector.

[0205] The implementation process of step (3) is as follows:

[0206] The intrinsic mode components (IMF / PF) obtained from step (2) are subjected to normalization preprocessing in sequence, resulting in three types of normalized feature matrices. , , Where m is the number of samples, As an electrical characteristic dimension, Temperature feature dimension This refers to the dimension of working condition characteristics.

[0207] Step (4) matrix fusion is implemented as follows:

[0208] Based on the unified UTC millisecond-level timestamp collected synchronously in step (1), the three types of standardized feature matrices are first time-series aligned and verified. When the missing rate is ≤5%, linear interpolation is used to complete the matrix. Then, the aligned three types of matrices are spliced ​​together column by column to form a multi-source fusion feature matrix. Finally, outliers are removed by the 3σ criterion and the overall missing rate is verified to be ≤1%. A qualified multi-source fusion feature matrix is ​​output as the input of the secondary fusion mechanism.

[0209] The implementation process of step (5) is as follows:

[0210] The original high-dimensional feature set extracted from the transformer multi-source monitoring data after preliminary differential decomposition and multi-source fusion constitutes the sample feature matrix. , where m is the number of samples and n is the total dimension of the original features.

[0211] First-level fusion: Feature layer fusion based on differentiable feature selection and nonlinear projection (LFFL);

[0212] Step 5.1: Multi-source feature integration and parameter initialization;

[0213] Feature selection gating vector: ,pass function Parameterization:

[0214] ,

[0215] in These are the basic parameters that can be trained. Characterizes the retention strength of the j-th original feature. Nonlinear fusion projection weights: projection matrix With bias vector Where d is the feature dimension after fusion ( ).

[0216] Step 5.2: Gated weighting and nonlinear transformation

[0217] The LFFL layer performs the following differentiable computation sequence on the input features to achieve a fusion of feature importance weighting and dimensionality compression:

[0218] Gated weighted modulation: Broadcasting the gated vector g to the feature matrix For the same dimension, perform element-wise (Hadamard) multiplication to achieve feature-level soft selection:

[0219] ,

[0220] in, This is a column vector with all elements being 1. This operation allows the gradient to propagate back through g, thereby learning the importance of each feature.

[0221] Nonlinear fusion projection: The weighted features are subjected to affine transformation and nonlinear activation, and mapped to a low-dimensional fusion space.

[0222] ,

[0223] in, It is a non-linear activation function (ReLU). This refers to the dimensionality-reduced and filtered feature representation of the first-level fusion output.

[0224] Second-level fusion: Deep decision fusion (DBN-CP) with decision contribution analysis capabilities.

[0225] Step 5.3: DBN-CP network structure;

[0226] Suppose that DBN-CP is pre-trained by stacking L Restricted Boltzmann Machines (RBMs) and then fine-tuned using supervised signals. Its top-level design is a dual-branch structure:

[0227] Classification branch: Outputs the probability distribution of fault categories.

[0228] ,

[0229] in, This is the activation value of the last hidden layer.

[0230] Contribution analysis branch: This is a fully connected layer that runs parallel to the classification branch, designed to analyze high-level semantic representations. The contribution estimate to the original n-dimensional features is reconstructed.

[0231] ,

[0232] in, , For the Sigmoid function, the output is... It can be interpreted as a normalized contribution score vector.

[0233] Step 5.4: Construct contribution alignment constraints between the two levels;

[0234] Contribution as resolved by DBN-CP The gating weights should match those actually used in the LFFL layer. The importance of the features they contain is highly consistent.

[0235] ,

[0236] in,

[0237] ,

[0238] This is a symmetric divergence measure, forced (Analytical contribution of the i-th sample) and global gating vector Align the distribution. By minimizing The network is constrained to learn in such a way that it is gated. The emphasized features must also be identified as high-contribution features in the DBN-CP decision-making process, and vice versa, thus forming a closed-loop feedback path from decision-making to feature selection.

[0239] Two-level collaborative end-to-end joint optimization training;

[0240] Step 5.5: Joint Loss Function and Optimization;

[0241] The system's total loss function integrates the classification objective, contribution alignment constraints, and feature selection sparsity regularization:

[0242] ,

[0243] in It is the standard cross-entropy classification loss; It is the weighting coefficient of the alignment loss, which controls the strength of the feedback constraint between the two levels; It is the L1 regularization coefficient, which is applied to the gated vector. This encourages sparsification and enables automated feature selection.

[0244] Step 5.6: Backpropagation and parameter update;

[0245] All parameters Joint updates are performed using stochastic gradient descent or its variants. Key gradient flows include:

[0246] The classification loss gradient is backpropagated from the top of DBN-CP to its bottom layers, and further propagated through Z to the parameters of the LFFL layer. .

[0247] Alignment loss Two gradient paths are generated: one acts on the contribution analytical branch parameter. Another one directly affects the gating parameters This is precisely the core manifestation of the feedback mechanism. about The gradient is:

[0248] ,

[0249] This gradient drives Average resolution contribution to the current batch of samples The adjustment enabled the backpropagation of decision knowledge to the feature selector.

[0250] Through iterative optimization, the system eventually converges to a stable state: the LFFL layer learns to select the most critical feature subset for making accurate, high-confidence decisions for DBN-CP, while DBN-CP learns to make interpretable decisions based on these selected features. The two mutually promote and optimize each other, ultimately forming a high-performance, highly reliable fusion diagnostic model.

[0251] Step 6 includes the following steps:

[0252] The output results include the specific fault type of the transformer (such as partial discharge, winding aging, insulation dampness, etc.) and a multi-source feature contribution analysis report. The contribution analysis report clarifies the contribution ratio of various electrical, temperature, and operating conditions features to the fault diagnosis results, providing a basis for fault tracing.

[0253] Previous transformer fault diagnosis methods often employed a "single decomposition + shallow fusion" approach, such as pure EMD+SVM, traditional DBN, and ordinary LSTM. This paper proposes a "differentiated customized decomposition + LFFL-DBN-CP two-level fusion" method, exploring transformer fault diagnosis from a new technological perspective and opening up a superior diagnostic paradigm. Comparative data from these four methods shows significant improvements in accuracy, stability, and generalization ability. Simulation data reveals a significant difference in performance among the different methods at optimal accuracy: pure EMD+SVM achieves 85.2%, traditional DBN reaches 90.1%, ordinary LSTM improves to 92.3%, while the method of this invention achieves an optimal accuracy of 98.5%, a 6.2 percentage point improvement over ordinary LSTM, far exceeding conventional methods in peak fault diagnosis accuracy. In terms of average accuracy, pure EMD+SVM achieves only 83.0%, traditional DBN 88.3%, and ordinary LSTM 90.0%, while the method of this invention reaches 97.0%. This means that the overall performance of this method is more stable across multiple diagnoses, without significant fluctuations in results. Test set accuracy is a key indicator of a model's generalization ability. The test set accuracy of pure EMD+SVM is 82.1%, traditional DBN 87.5%, and ordinary LSTM 89.2%, while the test set accuracy of the method of this invention reaches 96.5%, indicating that it maintains high diagnostic accuracy even on new data not used in training, demonstrating significantly better generalization ability than traditional methods. Standard deviation reflects the volatility of diagnostic results. The standard deviation of pure EMD+SVM is 3.5, traditional DBN is 2.8, and ordinary LSTM is 2.2, while the standard deviation of the method of this invention is only 1.0, far lower than other methods. This indicates that the diagnostic results of this method are more stable, more robust, and can maintain consistently high performance across different scenarios.

[0254] Table 1 Comparison of Optimization, Improvement and Enhancement

[0255]

[0256] This embodiment applies to a partial discharge fault scenario in a 110kV transformer, including:

[0257] I. Site Layout Design

[0258] For 110kV oil-immersed transformers (S11-6300) operating under high-frequency load fluctuation conditions, this invention comprehensively considers the multi-source data acquisition requirements and develops a detailed point-of-sale method to achieve comprehensive acquisition of fault characteristics.

[0259] 1. Electrical data acquisition points: One high-precision dielectric loss factor sensor (measurement range 0~0.1, accuracy ±0.0001) is installed at the end of the high-voltage winding and the output end of the low-voltage winding of the transformer. A DC resistance monitoring module (measurement accuracy ±0.5%) is installed on the neutral point side of the winding. The sampling frequency is set to 10Hz to synchronously acquire dielectric loss factor and winding DC resistance signals, and to capture minute changes in electrical parameters caused by partial discharge.

[0260] 2. Temperature data acquisition points: One fiber optic temperature sensor (measurement range -20~120℃, accuracy ±0.3℃) is placed on the top of the transformer tank and near the high-voltage winding core. One NTC temperature sensor (measurement range -20~100℃, accuracy ±0.5℃) is installed on the side wall of the tank. The sampling frequency is 5Hz to monitor the local temperature rise that may be caused by partial discharge.

[0261] 3. Install current transformers on the secondary side of the transformer to obtain load rate data, and use current transformers with ±1% accuracy to collect data; in addition, based on the substation SCADA system, acquire operating condition data (cumulative running time, voltage level, etc.); the sampling rate is 1Hz, that is, one frame of data per second is collected to reflect the high-frequency load change trend.

[0262] 4. Synchronization Mechanism Deployment: All sensors are equipped with modules featuring UTC millisecond-level timestamps, and industrial Ethernet is used as the data transmission channel, enabling simultaneous data transmission. Based on this, a unified time series standard can be achieved for multi-source data, facilitating better subsequent differential decomposition and fusion.

[0263] II. System Deployment and Operation Process

[0264] (a) Hardware deployment

[0265] 1. Edge computing node setup: Install the Huawei Atlas200 IDKA2 edge computing module in the control room of the substation. This module has a computing power of 16 TOPS and uses a combination of ARM architecture CPU and AI accelerator to undertake real-time data computing tasks, run business-related algorithms, and perform fault diagnosis and analysis. At the same time, the module must also be able to adapt to the low latency requirements of the site.

[0266] 2. Data transmission module setup: 5G industrial modules are used for data transmission. The transmission rate of 5G industrial modules reaches over 100Mbps. They are used to transmit all data between sensors and edge computing nodes, ensuring normal transmission of high-frequency collected data and preventing data loss during transmission. The transmission latency must be kept within 10ms throughout the entire transmission process.

[0267] 3. Early warning output module construction: This module is connected to the substation operation and maintenance monitoring platform and the mobile APP of operation and maintenance personnel. It has the function of real-time push of fault early warning information, and can push the type, confidence level and fault type contribution analysis report of fault early warning information.

[0268] (ii) Software Deployment

[0269] 1. To achieve the deployment of differentiated decomposition algorithms, KLA-EMD, VMD-EBO-L, and SSA-LMD algorithms are embedded into edge computing nodes. For the electrical signal (dielectric loss factor) mentioned in this paper, the system calls the KLA-EMD decomposition module, and sets the corresponding parameters in the module: dynamic adjustment threshold of time-varying noise figure; extreme point set partitioning ratio γ is set to 0.7; the optimization range of polynomial regression order is 2 to 8, etc.

[0270] 2. First, the deployment and implementation of the two-level fusion model are carried out. After loading the LFFL feature layer fusion model and the DBN-CP decision layer fusion model, all the relevant parameters mentioned above are initialized. The gate vector v is initialized as a vector in Rn space, and the projection matrix Wf is initialized as a matrix in R(n×d) space (d is set to 32). Then, for the DBN-CP network, the fully connected network formed by stacking L=3 RBMs during pre-training is used for initialization, and the classification branch and the contribution analysis branch are run together at the same time.

[0271] 3. To implement the deployment and configuration of this joint optimization module, the joint loss function is selected as the optimization criterion, and the parameters α=0.8 and β=0.05 are set. The stochastic gradient descent optimizer is used with a learning rate of 0.001. At the same time, online adaptive updates are enabled every 8 hours. The advantage of this is that the model can be updated to the latest dataset at any time without shutting down the model for retraining, so that the recommendation results for users are always online.

[0272] (III) Operation and Diagnosis Process

[0273] 1. Data acquisition stage: When the transformer is under high-frequency load fluctuation conditions (i.e., load rate of 80-110%), sensors at the same location simultaneously sample data. The obtained data such as dielectric loss factor, winding DC resistance, temperature and operating conditions are transmitted together through the 5G module to the edge computing node for further data processing.

[0274] 2. Differential decomposition process: After the system identifies electrical signals based on the dielectric loss factor, the KLA-EMD decomposition algorithm is called. During the algorithm operation, time-varying noise is added to amplify subtle signal features. Then, starting from the corresponding extreme points, the multinomial regression method is used to smooth the signal and draw the envelope. The verification is completed through the IMF dual criterion. Finally, the IMF components with strong partial discharge characteristics (denoted as IMF1-IMF3) are selected from the original data, which can effectively remove the interference noise generated by high-frequency loads.

[0275] 3. Secondary Fusion Diagnosis Stage: For the LFFL layer, firstly, a gated weighted modulation method (0≤gj≤1) is used, followed by nonlinear projection to obtain the final dimension-reduced feature vector Z. After obtaining Z, the LFFL network can provide the probability information of the fault category, such as the probability of partial discharge fault reaching 99.2%, to the DBN-CP network. Furthermore, the contribution ratio of different features to the operating condition is quantitatively analyzed through the contribution analysis branch. Among them, the contribution rate of electrical features is 78.3%, and the contribution of the dielectric loss factor IMF2 component is 0.32, which is the core fault feature of this operating condition.

[0276] 4. Early Warning and Verification Steps: The system's diagnostic response time is only 2.5ms, and it can push partial discharge fault alarm information and contribution analysis reports to the operation and maintenance platform 72 hours in advance. After receiving the report, the operation and maintenance personnel can check the high-voltage winding end by referring to the key prompts in the report, and accurately locate the fault point using an ultrasonic partial discharge detector. If the fault point is located, insulation reinforcement can be carried out immediately at the fault point to prevent further faulting that could cause winding burnout and power grid outage, ensuring normal equipment operation and reducing the occurrence of accidents.

[0277] This invention proposes a transformer fault diagnosis method based on differential decomposition and a two-level fusion strategy, which has significant advantages over traditional transformer fault feature extraction and diagnostic decision algorithms. Furthermore, there are various possible implementation methods for this technical solution; this paper selects only the most suitable preferred implementation method, which provides more room for creativity and imagination for later developers. It should be emphasized that those skilled in the art can make adjustments or modifications to the specific structure and operation process without departing from the essence of this invention, and the technical effects obtained by these modifications should also fall within the scope of protection of this invention. In addition, parts not described in detail in this embodiment can be replaced by existing mature technologies without affecting the effects of this embodiment or the technical effects of this invention.

Claims

1. A transformer fault diagnosis method based on differential decomposition and two-level fusion, characterized in that, Includes the following steps: Step 1: Synchronously collect multi-source monitoring data of the transformer, including dielectric loss factor and winding DC resistance data, top oil temperature and winding temperature data, load rate and running time data; and bind it to a Coordinated Universal Time (UTC) millisecond-level timestamp; Step 2: The multi-source monitoring data collected in Step 1 is decomposed using a differentiated decomposition strategy to obtain the intrinsic mode components of each signal. The specific decomposition methods are as follows: the dielectric loss factor and winding DC resistance signals are decomposed by kernel local alignment-empirical mode decomposition (KLA-EMD); the top oil temperature and winding temperature signals are decomposed by variational mode decomposition-enhanced observer learning (VMD-EOBL); and the load rate and running time signals are decomposed by singular spectrum analysis-local mean decomposition (SSA-LMD). Step 3: Perform normalization preprocessing on each intrinsic modal component obtained in Step 2 in sequence; Step 4, matrix fusion: The multi-source intrinsic modal component features after standardization and preprocessing are time-series aligned and matrix-concatenated to form a unified multi-source fusion feature matrix; Step 5: Construct a two-level fusion mechanism, which completes deep decision fusion and contribution analysis through the learnable feature fusion layer LFFL and the deep belief network-contribution analysis DBN-CP network; Step 6, Output fault diagnosis results: Output transformer fault diagnosis results and multi-source feature contribution analysis report.

2. The method according to claim 1, characterized in that, Step 2 includes: Step 2.1, Kernel Local Alignment-Empirical Mode Decomposition (KLA-EMD) of Electrical Signals, includes the following steps: Step 2.1.1: Fusion of time-varying noise injection and random subspace regression; By injecting time-varying noise into the original electrical signal, a balance is achieved between weak feature enhancement and noise suppression. An adaptive ensemble noise injection and random subspace regression fusion framework is employed. , in, After noise injection, the first indivual Class of electrical signals, Original electrical signal, It is Gaussian white noise with a mean of 0 and a variance of 1. δ is the time-varying noise figure, δ is the dielectric loss factor, and R is the DC resistance of the winding. Step 2.1.2: Divide the training and validation subsets; proportionally The set of extreme points is divided into training and validation subsets: , , in, For the first indivual The complete set of minimum points of the signal type, For the first indivual The training subset of complete minimum points of the signal class. For the first indivual A complete set of verification subsets of the minimum points of the signal class; For the first indivual The complete set of maxima of the signal class For the first indivual The training subset of complete maxima of the signal class. For the first indivual A complete subset of the verification points of the signal class; The following partitioning constraints must be met: , , , in, For the first indivual Signal type in time amplitude, The sampling interval is... , The time coordinates of the extreme point; Step 2.1.3: Kernel polynomial regression and smooth envelope generation; Training subsets with maxima Using the input, construct a d-order multinomial regression model: , in, In time The predicted value of the fitted envelope amplitude. The coefficients to be determined; the validation subset will be used. Input a d-th order polynomial regression model, adjust d using a grid search method, and determine the optimal order with the goal of minimizing the mean squared error. : , Among them, It is the first The time corresponding to each maximum point coordinate; This corresponds to the corresponding time point. The signal amplitude, where q is the number of samples in the validation subset; based on By fitting the sets of maxima and minima respectively, the upper envelope is obtained. Lower envelope Calculate the rate of change of the first derivative of the envelope; if it is greater than the threshold, add L2 regularization and refit. Step 2.1.4: Verification of Intrinsic Modal Components (IMF) using dual criteria; The effectiveness of the Intrinsic Mode Components (IMFs) extracted from the envelope generated by kernel multinomial regression and adaptive optimization is verified through quantitative constraints based on dual criteria. The mathematical definition and logic of the dual criteria are as follows: Criterion 1, Local Extremum Matching Criterion: Requires the intrinsic mode components (IMF components) to match. The quantization expression is as follows: The difference between the number of upper and lower extreme points does not exceed 1, and the distribution of extreme points within a local region is consistent with the original signal. , , in For the intrinsic mode components IMF components The total number of maxima. For the intrinsic mode components IMF components The total number of local minimum points; The preset threshold; Let be any local time ?; Criterion 2, Kernel Fitting Residual Constraint Criterion: Requires the intrinsic mode components and IMF components to... Residual with the original signal Satisfy the kernel fitting residual threshold constraint, where The original electrical signal, The intrinsic mode components (IMFs) obtained from the decomposition of the original signal in time The numerical value, quantized as: , in, For residual signal In terms of time length The mean square value, where T is the signal duration. The residual threshold; Both criteria must be met simultaneously; otherwise, return to step 2.1.3 to readjust the kernel polynomial order or regularization weights. Step 2.2, Variational Mode Decomposition of Temperature Signal - Enhanced Observer Learning VMD-EOBL Decomposition, includes the following steps: Step 2.2.1: Define a multi-objective optimization problem that is hierarchical in terms of physical meaning; We innovatively construct a fitness function that integrates four criteria, transforming Variational Mode Decomposition (VMD) into a physically interpretable optimization problem: , in, For multi-objective fitness functions, The parameters to be optimized include the number of modes. Punishment factor Time scale parameter τ; These are the fundamental oscillatory components with physical meaning obtained after variational mode decomposition (VMD) of the temperature signal. This refers to the k-th fundamental oscillation component obtained after variational mode decomposition of the VMD temperature signal. Characterizes signal fidelity. The function is used to evaluate component sparsity. The function is used to measure the independence of components. The function passes through each fundamental oscillation component The similarity to the preset physical process template library L is used to quantify the degree of matching of physical meaning; , , , These are dynamic weighting coefficients; Step 2.2.2: Enhanced observer learns EOBL algorithm parameters training; We design an enhanced observer learning EOBL algorithm with a prior experience pool. Historical high-quality parameters are introduced during population initialization to accelerate convergence. The individual position update formula is: , in, For individuals In the The new position of the era For individuals In the The position of the generation, This is the globally optimal solution. The function analyzes population distribution Gradient direction prediction with experience pool For adaptive noise, intermediate parameters Dynamically adjusted with iteration; combined with elite retention before Early cessation strategy for individual fitness plateau; Step 2.2.3: After decomposition, physical labels are automatically bound to the basic oscillation components IMF, generating feature vectors with clear physical semantics; Step 2.3, Singular Spectrum Analysis - Local Mean Decomposition (SSA-LMD) of Operating Condition Signals; includes the following steps: Step 2.3.1: Singular Spectrum Analysis of Physical Constraints - SSA Trend-Residual Separation; Introduce a constraint function based on operational physics knowledge. Used to evaluate each specific and complete oscillation pattern, trend, or noise component. Physical rationality; definition of physical fit index : , in, For similarity function: calculate the reconstructed components The degree of matching with the preset physical process template, Physical constraint function: Domain knowledge-based evaluation Physical rationality; maximizing trend components through constraint optimization. Minimize residual components Output trend terms with clear physical meaning. With residuals ; Step 2.3.2: Parameter-adaptive Local Mean Decomposition (LMD) fine decomposition; The residual term R(t) from the first-level singular spectral analysis (SSA) trend-residual separation output is input into the local mean decomposition (LMD) for fine decomposition. Local Mean Decomposition (LMD) iteratively decomposes the signal into a product function by calculating local extrema, the local mean function, and the envelope estimation function, thus constructing a bi-objective optimization function. ; in, The overall objective function for Local Mean Decomposition (LMD) optimization is... The set of parameters to be optimized for Local Mean Decomposition (LMD) For dynamic weights, To decompose the error term: ; in The first local mean decomposition (LMD) obtained by the local mean decomposition Product functions, These are the weighting coefficients. To obtain the maximum value of the coefficient of variation among all PF components of the product function, To penalize the envelope function in the PF component of the product function, For the physical matching degree item: ; in, Indicates taking the first The product function PF components and all templates The maximum similarity among them. A physical microprocess template library, The time-frequency coherence coefficient, It is the entropy value; Step 2.3.3: Implement adaptive collaborative optimization of two-level parameters.

3. The method according to claim 2, characterized in that, Step 2.3.3 includes: defining optimization parameters. , This is a subset of parameters to be optimized for the SSA decomposition layer in singular spectral analysis. For the subset of parameters to be optimized in the Local Mean Decomposition (LMD) layer, construct a comprehensive fitness function. : , in, To ensure fidelity, For complexity terms, The sum of the physical interpretability scores. , , These are the weighting coefficients; The outer ring employs an improved swarm intelligence algorithm to globally search the subset of parameters to be optimized in the Singular Spectrum Analysis (SSA) decomposition layer. The inner loop optimizes the subset of parameters to be optimized in the Local Mean Decomposition (LMD) layer by gradient-guided search. This achieves dual-level parameter coupling optimization; finally, it extracts the statistical and time-frequency features of the trend term and the PF component of the product function to form a hierarchical differentiated feature vector.

4. The method according to claim 3, characterized in that, Step 3 includes: The intrinsic modal components obtained from step 2 are subjected to normalization preprocessing in sequence, resulting in three types of normalized feature matrices: , , , in, For the normalized feature matrix of electrical signals, The normalized feature matrix of the temperature signal, Here, m is the standardized feature matrix of the operating condition signal. As an electrical characteristic dimension, Temperature feature dimension This refers to the dimension of working condition characteristics.

5. The method according to claim 4, characterized in that, Step 4 includes: Based on a unified UTC millisecond-level timestamp, the three types of standardized feature matrices are first time-series aligned and verified. When the missing rate is less than or equal to the threshold, linear interpolation is used to complete the matrix. Then, the aligned three types of matrices are concatenated column by column to form a multi-source fusion feature matrix. Finally, outliers are removed by the 3σ criterion and the overall missing rate is verified to be less than or equal to the threshold. A qualified multi-source fusion feature matrix is ​​output as the input for the secondary fusion mechanism in step 5.

6. The method according to claim 5, characterized in that, Step 5 includes: The original high-dimensional feature set extracted from the transformer multi-source monitoring data after preliminary differential decomposition and multi-source fusion constitutes the sample feature matrix. , where n is the total dimension of the original features; Step 5.1, First-level fusion: Learnable Feature Fusion Layer (LFFL); Step 5.1.1: Multi-source feature integration and parameter initialization; Feature selection gating vector ,pass function Parameterization: , in, It is a natural constant. For the first One trainable basic parameter component Characterizing the first The retention strength of each original feature; Nonlinear fusion projection weights: projection matrix With bias vector , where d is the feature dimension after fusion; Step 5.1.2: Gated weighting and nonlinear transformation; The Learnable Feature Fusion Layer (LFFL) performs the following differentiable computation sequence on the input features to achieve a fusion of feature importance weighting and dimensionality compression: Gated weighted modulation: the gate vector Broadcast to feature matrix In the same dimension, perform element-wise Hadamard product e to achieve feature-level soft selection: , in, It is a column vector with all elements being 1; H represents the transpose; Indicates the first The first sample One eigenvalue; Nonlinear fusion projection: The weighted features are subjected to an affine transformation and nonlinear activation, mapping them to a low-dimensional fusion space. , in, The features are after gating and weighting. It is a non-linear activation function. Let be the projection weight matrix. This is the transpose of the bias vector. The dimensionality-reduced and filtered feature representations of the first-level fusion output are used as inputs for the second-level fusion. Step 5.2, Second-level fusion: Deep Belief Network - Contribution Analysis DBN-CP; Step 5.2.1: Establish the Deep Belief Network-Contribution Analysis (DBN-CP) network structure; Suppose that the Deep Belief Network-Contribution Parsing (DBN-CP) network is pre-trained by stacking L Restricted Boltzmann Machines (RBMs) and then fine-tuned using supervised signals. The top-level design of the DBN-CP network is a dual-branch structure, including: Classification branch: Output the probability distribution of fault categories: , in, It is the weight matrix of the classification layer. This is the bias vector of the classification layer, where C is the number of fault categories. The activation value of the last hidden layer, after... After processing, To become a probability vector; Contribution analysis branch: This is a fully connected layer that runs parallel to the classification branch, used to analyze high-level semantic representations. The contribution estimate to the original n-dimensional features is reconstructed from the data: , in, The weight matrix for the contribution analysis branch. The bias vector for the contribution analysis branch. For the Sigmoid function, the output is... This can be interpreted as a normalized contribution score vector; Step 5.3: Construct contribution alignment constraints between the two levels; Deep Belief Network - Contribution Analysis (DBN-CP) - Contribution Analysis The gating weights actually used in the LFFL layer, which is a fusion layer of learnable features. The characteristics they imply are of consistent importance: , in, To use global weights Based on the benchmark, measure the performance of the first Explanation of the contribution of each sample The amount of additional information it carries; Explanation based on a single sample Based on this, measure the global weight. The amount of additional information it carries; This is the loss value; ; in For the first The degree of contribution of each feature to the sample decision, To give the first The importance weights of each feature; Step 5.4, Joint Loss Function and Optimization; Total loss function It integrates classification objective, contribution alignment constraint, and feature selection sparsity regularization: , in, It is the standard cross-entropy classification loss; The probability distribution of fault categories predicted by the model; Labels for actual faults; These are the weighting coefficients for the alignment loss; It is the L1 regularization coefficient; for Norm, acting on gated vectors ; Step 5.5, backpropagation and parameter update.

7. The method according to claim 6, characterized in that, Step 5.5 includes: all parameters Joint updates via stochastic gradient descent or its variants include the following key gradient flows: The classification loss gradient is backpropagated from the top of the Deep Belief Network-Contribution Parsing (DBN-CP) to the bottom layer, and further propagated through Z to the parameters of the Learnable Feature Fusion (LFFL) layer. ; Alignment loss Two gradient paths are generated: A parameter applied to the contribution analysis branch ; Another one directly affects the gating parameters , about gradient for: , in, For alignment loss on the first Gating basic parameters of each feature The partial derivative (gradient) of . For the first The first sample The contribution estimate of each feature.

8. The method according to claim 7, characterized in that, In step 6, the output results include the specific fault type of the transformer and a multi-source feature contribution analysis report. The contribution analysis report clarifies the contribution ratio of various electrical, temperature, and operating condition features to the fault diagnosis results.

9. An electronic device, characterized in that, It includes a processor and a memory, the memory storing program code that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 8.

10. A storage medium, characterized in that, It stores a computer program or instructions that, when run on a computer, perform the steps of the method as described in any one of claims 1 to 8.