Transfer learning transformer fault diagnosis method and system based on intermediate harmonic distribution
By employing a transfer learning method based on intermediate harmonic distribution in transformer fault diagnosis, intermediate harmonic distribution samples with the smallest geometric distance from the dual-domain distribution are generated. This solves the problems of conditional distribution offset and lack of label semantics in cross-domain diagnosis, and achieves efficient and accurate fault diagnosis.
Patent Information
- Application Number
- CN202510870334.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-31
AI Technical Summary
Existing transformer fault diagnosis methods suffer from conditional distribution offset and lack of label semantic information in cross-domain diagnosis, which leads to model confusion between different fault categories. Especially when the target domain label data is scarce, the propagation of false label errors deteriorates the diagnostic performance, and federated learning has difficulty achieving stable convergence in heterogeneous distribution alignment.
An intermediate harmonic distribution-based transfer learning method is adopted. By configuring residual networks in the source and target domains respectively, fitting Gaussian mixture models, generating intermediate harmonic distribution samples, and training stacked autoencoders with minimum transfer cost, the residual network parameters are dynamically updated to achieve collaborative training and iterative optimization between the client and server.
It significantly improves cross-domain diagnostic accuracy, avoids pattern collapse and distribution shift, achieves efficient and accurate fault diagnosis, and protects data privacy.
Smart Images

Figure CN120873668A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep transfer learning fault diagnosis technology for power system transformer equipment. Specifically, it relates to a transfer learning method and system for transformer fault diagnosis based on intermediate harmonic distribution, and also relates to a corresponding computer terminal and computer-readable storage medium. Background Technology
[0002] With the increasing intelligence of power systems, transformers, as core equipment, have attracted much attention for fault diagnosis technology. In recent years, deep transfer learning technology has been introduced into this field to improve the generalization ability of models through cross-domain feature alignment, but the following key bottlenecks still exist:
[0003] Existing methods often focus on aligning edge distributions, neglecting the conditional distribution shifts of equipment failure modes across different domains. For example, the vibration characteristic distribution of the same failure type may undergo nonlinear distortions in laboratory simulations and real industrial scenarios due to factors such as load fluctuations and noise interference. Simply aligning the overall feature distribution can easily lead to model confusion between different failure categories, especially when labeled data is scarce in the target domain, where pseudo-label error propagation can further degrade diagnostic performance. Some studies have attempted to construct intermediate mediators using generative adversarial networks or random prior distributions to bridge cross-domain differences, but these suffer from the problem of generated samples deviating from the true distribution. Furthermore, existing methods do not inject label semantic information into the intermediate distribution generation process, resulting in a lack of class discriminativeness in the alignment process and an inability to achieve refined failure mode transfer.
[0004] Although federated learning has been proposed for distributed training with privacy protection, it is mainly designed for homogeneous data scenarios and has not been optimized for heterogeneous distribution alignment in cross-domain diagnostics. Directly transmitting model parameters or gradients can easily introduce domain-specific noise, and the global model update direction is limited by local domain biases, making it difficult to achieve stable convergence. Summary of the Invention
[0005] To address the aforementioned shortcomings in the prior art, this invention provides a method and system for diagnosing transformer faults based on intermediate harmonic distribution through transfer learning, along with a corresponding computer terminal and computer-readable storage medium.
[0006] According to one aspect of the present invention, a transfer learning method for transformer fault diagnosis based on intermediate harmonic distribution is provided, comprising:
[0007] The source domain transformer vibration signal and the target domain transformer vibration signal are acquired separately.
[0008] Configure the first residual network on the source domain client, extract the high-dimensional features of the source domain transformer vibration signal, fit the Gaussian mixture model of the source domain features, and upload the parameters of the Gaussian mixture model to the server.
[0009] Configure a second residual network on the target domain client, extract sample features of the transformer vibration signal in the target domain, fit a Gaussian mixture model of the target domain features, and upload the parameters of the Gaussian mixture model to the server.
[0010] The server receives the Gaussian mixture model parameters of the source domain and the target domain, generates simulated samples, and trains a stacked autoencoder with minimum transfer cost to generate intermediate harmonic distribution samples.
[0011] The intermediate harmonic distribution samples are broadcast to the client. The source domain and target domain clients align their local feature distributions to the intermediate harmonic distribution by minimizing the minimum migration cost, dynamically update the corresponding residual network parameters, and perform collaborative training and iterative optimization on the client and server until convergence. Transformer fault classification is then performed to complete fault diagnosis in the target domain.
[0012] According to another aspect of the present invention, a transfer learning transformer fault diagnosis system based on intermediate harmonic distribution is provided, comprising:
[0013] The feature extraction module is used to configure a first residual network in the source domain client to extract high-dimensional features of the source domain transformer vibration signal; and to configure a second residual network in the target domain client to extract sample features of the target domain transformer vibration signal.
[0014] The Gaussian Mixture Model (GMMM) modeling module is used to fit a Gaussian mixture model of source domain features on the source domain client and upload the parameters of the Gaussian mixture model to the server; and to fit a Gaussian mixture model of target domain features on the target domain client and upload the parameters of the Gaussian mixture model to the server.
[0015] An intermediate harmonic distribution generation module is used to receive Gaussian mixture model parameters of the source domain and the target domain on the server side, generate simulated samples, train a stacked autoencoder with minimum transfer cost to generate intermediate harmonic distribution samples, and broadcast the intermediate harmonic distribution samples to the client.
[0016] A local feature alignment module is used to align the local feature distribution to an intermediate harmonic distribution by minimizing the minimum migration cost.
[0017] The client-server collaborative training module is used to dynamically update the corresponding residual network parameters, perform collaborative training and iterative optimization on the client and server until convergence, and perform transformer fault classification to complete fault diagnosis in the target domain.
[0018] According to a third aspect of the present invention, a computer terminal is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can be used to perform the methods described above in the present invention, or to run the system described above in the present invention.
[0019] According to a fourth aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, can be used to perform the methods described above in the present invention, or to run the system described above in the present invention.
[0020] By adopting the above technical solution, the present invention has at least one of the following beneficial effects compared with the prior art:
[0021] The present invention provides a method and system for transformer fault diagnosis based on intermediate harmonic distribution. Based on the federated transfer diagnosis framework of intermediate harmonic distribution, it generates an intermediate harmonic distribution with the smallest geometric distance from the dual-domain distribution as a transfer medium, and achieves conditional distribution alignment by minimizing the transfer cost, thereby significantly improving the cross-domain diagnosis accuracy while protecting data privacy.
[0022] The present invention provides a method and system for transformer fault diagnosis based on intermediate harmonic distribution, which adopts a distributed harmonic mechanism and uses stacked autoencoders to dynamically generate intermediate harmonic samples that have both source domain label semantics and target domain distribution characteristics, thereby avoiding mode collapse and distribution shift.
[0023] The present invention provides a method and system for transformer fault diagnosis based on intermediate harmonic distribution, which adopts the method of minimizing the transfer cost and designs a category-aware transfer cost function to constrain the alignment of similar fault features to the intermediate distribution and suppress negative transfer.
[0024] The present invention provides a method and system for transformer fault diagnosis based on intermediate harmonic distribution through transfer learning. It adopts federated collaborative training, and completes efficient and accurate fault diagnosis by updating the server, updating the local model parameters on the client and updating the intermediate harmonic distribution medium through collaborative training. Attached Figure Description
[0025] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:
[0026] Figure 1 This is a flowchart illustrating the workflow of a transfer learning transformer fault diagnosis method based on intermediate harmonic distribution in a preferred embodiment of the present invention.
[0027] Figure 2This is a schematic diagram of the components of a transfer learning transformer fault diagnosis system based on intermediate harmonic distribution in a preferred embodiment of the present invention.
[0028] Figure 3 This is a curve showing the relationship between a rigid object and the output amplitude of a flexible graphene sensor in a specific application example of the present invention.
[0029] Figure 4 This refers to the training time for each neural network model in a specific application example of the present invention.
[0030] Figure 5 This is a graph showing the average training accuracy of each network in a specific application example of the present invention.
[0031] Figure 6 This is a weight parameter-accuracy graph in a specific application example of the present invention.
[0032] Figure 7 This is an iterative diagram of the loss function of each network in a specific application example of the present invention. Detailed Implementation
[0033] The embodiments of the present invention are described in detail below: These embodiments are implemented based on the technical solution of the present invention, and provide detailed implementation methods and specific operation processes. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention.
[0034] To address the shortcomings of existing technologies, one embodiment of the present invention provides a transfer learning method for transformer fault diagnosis based on intermediate harmonic distribution. This method acquires transformer vibration signals using an optimized flexible graphene vibration sensor; it uses the vibration signals of the transformer undergoing fault diagnosis as the target domain and vibration signals acquired during laboratory fault experiments as the source domain; it deploys local feature extraction modules on both the source domain transformer device and the target domain transformer client, constructs a federated learning framework, extracts high-dimensional features of the transformer vibration signals using a deep residual network, and fits a Gaussian mixture model; it further extracts high-dimensional features of the transformer vibration signals using a deep residual network and... A Gaussian mixture model is fitted. On the server side, based on the intermediate harmonic distribution mechanism of the Gaussian mixture model and stacked autoencoders, parameter information of the feature distributions in the source and target domains is aggregated. The autoencoder generates an intermediate distribution with the minimum distance to the dual-domain distribution as a transfer medium. Local feature distributions are aligned to the medium by minimizing the minimum transfer cost, driving local features to align to the intermediate distribution. Finally, a collaborative optimization and fault diagnosis process between the client and server is implemented. The Gaussian network model parameters are iteratively and dynamically updated and encrypted, then transmitted back to the server. Collaborative training and iterative optimization continue until convergence, and transformer fault classification is performed, thus completing fault diagnosis in the target domain. This method effectively performs transfer diagnosis of transformers, significantly improving the accuracy and reliability of predictions.
[0035] Specifically, such as Figure 1 As shown, the transfer learning transformer fault diagnosis method based on intermediate harmonic distribution provided in this embodiment may include:
[0036] S1, acquire the source domain transformer vibration signal and the target domain transformer vibration signal respectively;
[0037] S2, Configure the first residual network on the source domain client, extract the high-dimensional features of the source domain transformer vibration signal, fit the Gaussian mixture model of the source domain features, and upload the parameters of the Gaussian mixture model to the server;
[0038] S3. Configure the second residual network on the target domain client, extract sample features of the transformer vibration signal in the target domain, fit a Gaussian mixture model of the target domain features, and upload the parameters of the Gaussian mixture model to the server.
[0039] S4, the server receives the Gaussian mixture model parameters of the source domain and the target domain, generates simulated samples, and trains a stacked autoencoder with minimum transfer cost to generate intermediate harmonic distribution samples;
[0040] S5 broadcasts the intermediate harmonic distribution samples to the client. The source domain and target domain clients align their local feature distributions to the intermediate harmonic distribution by minimizing the minimum migration cost, dynamically update the corresponding residual network parameters, and perform collaborative training and iterative optimization on the client and server until convergence. Finally, transformer fault classification is performed to complete fault diagnosis in the target domain.
[0041] In some preferred embodiments, the above-mentioned S1, which acquires the source domain transformer vibration signal and the target domain transformer vibration signal respectively, may further include:
[0042] S11, an optimized flexible graphene vibration sensor is used to collect the vibration signal of the transformer in the fault experiment as the source domain transformer vibration signal, which is a fully labeled transformer vibration signal.
[0043] S12, an optimized flexible graphene vibration sensor is used to collect the vibration signal of the transformer used for fault diagnosis as the target domain transformer vibration signal. The target domain transformer vibration signal is a partially marked or unmarked transformer vibration signal.
[0044] in:
[0045] The optimized flexible graphene vibration sensor includes a flexible graphene vibration sensor body and a rigid weight attached to the flexible graphene vibration sensor body. Furthermore, the rigid weight is added to the flexible graphene vibration sensor to enhance its output amplitude. The rigid weight converts external acceleration into a force on the flexible graphene elastomer through inertial effects. This force causes deformation of the elastomer. When the sensor is vibrated, the inertial force of the rigid weight drives the flexible structure to undergo compressive deformation, leading to the reconstruction of its internal conductive network. This deformation causes a faster change in the internal resistance of the material, ultimately increasing the output amplitude of the output electrical signal.
[0046] In some preferred embodiments, the above-mentioned S2, configuring the first residual network on the source domain client, may further include:
[0047] S21, the input data for the first residual network is: a set of fully labeled transformer vibration signals. in, This is the i-th input sample in the source domain transformer vibration signal; For the corresponding The label indicates the target category of the vibration signal; n s This represents the total number of tagged vibration signal samples in the source domain client.
[0048] S22, the output data of the first residual network is:
[0049] The network architecture for constructing the first residual network includes: four residual blocks arranged sequentially; each residual block includes two convolutional layers connected by skip connections; each convolutional layer is followed by a feature normalization layer; the feature normalization layer maps the high-dimensional features of the input data to a unit hypersphere, represented as:
[0050]
[0051] In the formula, These are the normalized feature vectors; Let represent the feature extraction function of the first residual network, and be the parameters of the convolutional layer; Here is the weight matrix of layer F2; σ is the LeakyReLU activation function;
[0052] S23, a Gaussian mixture model that fits the features of the source domain, may further include:
[0053] The parameters of the Gaussian mixture model are estimated, where each healthy state corresponds to a Gaussian component. The source domain features are fitted using the EM algorithm, and the result is expressed as follows:
[0054]
[0055] In the formula, G s (v s |θ s ) represents a Gaussian mixture model in the source domain; v s Let θ be the input random variable of the source domain Gaussian mixture model, representing the distribution of the source domain eigenvectors; s For the parameters of the source domain Gaussian mixture model, in, For mixed weights, The mean, R represents the covariance; R is the total number of transformer fault state categories; N is the Gaussian distribution symbol.
[0056] In some preferred embodiments, the above-described S3, configuring the second residual network on the target domain client, may further include:
[0057] S31, the input to the second residual network is: a set of partially labeled anchor points. and unlabeled sample set in, For the k-th input sample in the target domain transformer vibration signal, there are a total of R different transformer state types, and each state is labeled with at least one label. composition; The j-th unlabeled vibration signal sample in the target domain; n t This represents the total number of unlabeled samples in the target domain.
[0058] S32, the output of the second residual network is:
[0059] The network architecture for constructing the second residual network includes: four residual blocks arranged sequentially; each residual block includes two convolutional layers connected by skip connections; each convolutional layer is followed by a feature normalization layer; the feature normalization layer maps the high-dimensional features of the input data to a unit hypersphere, represented as:
[0060]
[0061] In the formula, These are the normalized feature vectors; Let represent the feature extraction function of the first residual network, and be the parameters of the convolutional layer; Here is the weight matrix of layer F2; σ is the LeakyReLU activation function;
[0062] S33, a Gaussian mixture model that fits the features of the target domain, may further include:
[0063] The parameters of the Gaussian mixture model are estimated, where each healthy state corresponds to a Gaussian component. The target domain features are then fitted using the EM algorithm, as follows:
[0064]
[0065] In the formula, G t (v t |θ t ) represents a Gaussian mixture model for the target domain; v t Vibration signal in the target domain The feature representation is extracted and normalized using the second residual network; θ t For the parameters of the Gaussian mixture model in the target domain, in, For mixed weights, The mean, R represents the covariance; R is the total number of transformer fault state categories; N is the Gaussian distribution symbol.
[0066] In some preferred embodiments, S4 above, generating a simulated sample, includes:
[0067] S41, sampling m through a source domain Gaussian mixture model. s From a sample, a simulated sample is obtained.
[0068] S42, sampling m through the Gaussian mixture model of the target domain. t From a sample, a simulated sample is obtained.
[0069] In some preferred embodiments, S4 above, training a stacked autoencoder with minimum transfer cost, includes:
[0070] S43, simulated sample and The data is merged as input data to the autoencoder. The encoder maps the variables uniformly to the latent space, generating latent variables z. i :
[0071]
[0072] In the formula, h enc (·) represents the encoder; This is the set of simulated samples sampled from the source domain Gaussian mixture model; The set of simulated samples sampled from the Gaussian mixture model of the target domain;
[0073] S44, the decoder utilizes the latent variable z i Perform sample reconstruction to obtain reconstructed samples.
[0074]
[0075] In the formula, h dec (·) represents the decoder;
[0076] S45, Establish the loss function, including:
[0077] S451, Establish the reconstruction loss L rccon The method of minimizing the mean square error of the input and output is expressed as:
[0078]
[0079] In the formula, m s Number of samples generated for the source domain; m t Generate the number of samples for the target domain; Reconstructing samples using an autoencoder;
[0080] S452, Establish the Wasserstein distance loss L g The source domain, target domain, and intermediate harmonic distribution G are calculated using the minimum migration cost. m The distance is expressed as:
[0081] L g =λ·OT(G s ||G m )+(1-λ)·OT(G t ||G m )
[0082] In the formula, OT(Gs ||G m ) represents the source domain distribution G s (i.e., the source domain Gaussian mixture model) to the intermediate harmonic distribution G m Minimum migration cost; OT(G t ||G m ) represents the target domain distribution G t (i.e., the Gaussian mixture model of the target domain) to the intermediate harmonic distribution G m The minimum migration cost; λ represents the weight coefficient, which is the priority for balancing the alignment of the source and target domains. The intermediate harmonic distribution G... m The set of labeled samples generated for the decoder; the process of generating this set is the intermediate harmonic distribution process.
[0083] In some preferred embodiments, the minimum migration cost in S452 above is:
[0084]
[0085] In the formula, Γ is the transmission plan matrix; i,j To make G s The transfer of the i-th sample to G m The proportion of the j-th sample; C i,j Indicates transmission cost;
[0086] The constraints are as follows:
[0087]
[0088] h dec (z j The sample generated by the decoder belongs to the intermediate harmonic distribution G. m ;z j These are latent variables.
[0089] In some preferred embodiments, the above-described S4, generating intermediate harmonic distribution samples, may further include:
[0090] S46, source domain features Input the trained stacked autoencoder to generate latent variables
[0091] S47, the decoder utilizes latent variables Perform sample reconstruction to generate intermediate harmonic distribution samples.
[0092] S48, the intermediate harmonic distribution sample By directly associating the source domain labels, a labeled intermediate harmonic distribution sample set is obtained.
[0093] In the above steps, when generating intermediate harmonic distribution samples, the input directly uses only the feature data of the source domain (because it has complete fault labels), but the core of the generated samples - the decoder - has been deeply constrained in training: the server forces the autoencoder to be trained to make its output closely resemble the features of both domains by combining simulated distribution data of the source domain and the target domain. The final generated samples are essentially a fusion of source domain semantics (preserving labels) and target domain morphology (distribution characteristics), becoming a transfer bridge connecting the two domains.
[0094] In some preferred embodiments, the above-described S5, where the source domain and target domain clients align their local feature distributions to an intermediate harmonic distribution by minimizing the minimum migration cost, may further include:
[0095] S51, Calculate source domain features With intermediate harmonic distribution samples Transmission cost
[0096]
[0097] In the formula, The transmission cost required to migrate the i-th feature sample from the source domain to the j-th sample in the intermediate harmonic distribution;
[0098] S52, update ResNet parameters using backpropagation gradients:
[0099] S521 will reduce transmission costs As a loss function;
[0100] S522, Calculate the gradient Update ResNet parameters θ s This aligns the source domain features with a harmonic distribution towards the center.
[0101] S53 uses the same calculation method and backpropagation gradient update method to align the target domain features with the intermediate harmonic distribution.
[0102] In some preferred embodiments, the above-mentioned S5, which involves collaborative training and iterative optimization of the client and server until convergence, may further include:
[0103] S54, the server continuously generates new intermediate tonal distribution samples through stacked autoencoders and broadcasts them to the client;
[0104] S55: The client optimizes the local residual network parameters by continuously updating the intermediate harmonic distribution samples, fits the new Gaussian mixture model of local features, and uploads the parameters of the Gaussian mixture model to the server to participate in the generation of new intermediate harmonic distribution samples.
[0105] S56, Repeat the above process for collaborative training and iterative optimization until the convergence condition is met. More preferably, the convergence condition includes: if the accuracy of the target domain validation set does not improve by 0.1% in 10 consecutive iterations, then stop the iteration and use the converged target domain parameters for fault diagnosis.
[0106] Based on the same inventive concept, an embodiment of the present invention also provides a transfer learning transformer fault diagnosis system based on intermediate harmonic distribution.
[0107] Specifically, such as Figure 2 As shown, the transfer learning transformer fault diagnosis system based on intermediate harmonic distribution provided in this embodiment may include:
[0108] The feature extraction module is used to configure a first residual network in the source domain client to extract high-dimensional features of the source domain transformer vibration signal; and to configure a second residual network in the target domain client to extract sample features of the target domain transformer vibration signal.
[0109] The Gaussian Mixture Model (GMMM) modeling module is used to fit a Gaussian mixture model of source domain features on the source domain client and upload the parameters of the Gaussian mixture model to the server; and to fit a Gaussian mixture model of target domain features on the target domain client and upload the parameters of the Gaussian mixture model to the server.
[0110] The intermediate harmonic distribution generation module is used to receive Gaussian mixture model parameters from the source and target domains on the server side, generate simulated samples, train a stacked autoencoder with minimum transfer cost to generate intermediate harmonic distribution samples, and broadcast the intermediate harmonic distribution samples to the client.
[0111] A local feature alignment module is used to align the local feature distribution to an intermediate harmonic distribution by minimizing the minimum migration cost.
[0112] The client-server collaborative training module is used to dynamically update the corresponding residual network parameters, perform collaborative training and iterative optimization on the client and server until convergence, and perform transformer fault classification to complete fault diagnosis in the target domain.
[0113] In some preferred embodiments, the above system may further include:
[0114] The data acquisition module uses an optimized flexible graphene vibration sensor to collect the vibration signal of the transformer in the fault experiment as the source domain transformer vibration signal, which is a fully labeled transformer vibration signal; and uses an optimized flexible graphene vibration sensor to collect the vibration signal of the transformer used for fault diagnosis as the target domain transformer vibration signal, which is a partially labeled or unlabeled transformer vibration signal.
[0115] The following describes in detail, with reference to preferred embodiments, the specific implementation methods of each functional module constituting the system provided in the above embodiments of the present invention.
[0116] The transfer learning transformer fault diagnosis system based on intermediate harmonic distribution provided in the above embodiments of the present invention, through the above functional modules, achieves the following functions as a whole: First, it uses an optimized flexible graphene vibration sensor to detect the vibration signal of the transformer; second, it uses the vibration signal of the transformer undergoing fault diagnosis as the target domain and the vibration signal collected in laboratory fault experiments as the source domain, deploying local feature extraction modules on the source domain transformer equipment and the target domain transformer client respectively, constructing a federated learning framework, extracting high-dimensional features of the transformer vibration signal through a deep residual network and fitting a Gaussian mixture model; third... The system employs a Gaussian mixture model and a stacked autoencoder-based intermediate harmonic distribution mechanism. On the server side, it aggregates parameter information from the source and target domain feature distributions and uses the autoencoder to generate an intermediate distribution that minimizes the distance to the dual-domain distribution as a transfer medium. The fourth part aligns local feature distributions to the intermediate distribution by minimizing the minimum transfer cost, driving local features to align with the intermediate distribution. The fifth part implements a collaborative optimization and fault diagnosis process between the client and server. This involves iteratively updating and encrypting model parameters and transmitting them back to the server, collaboratively training and iteratively optimizing until convergence, and then classifying transformer faults to complete fault diagnosis in the target domain. This system effectively performs transformer transfer diagnosis, significantly improving the accuracy and reliability of predictions.
[0117] More preferably, the feature extraction module configures a first residual network on the source domain client, specifically in the following ways:
[0118] Source domain residual network input: Fully labeled transformer vibration signal set in: The i-th input sample in the source domain is specifically the transformer vibration signal data; correspond The label indicates the target category of the vibration signal.
[0119] The first residual network architecture specifically includes:
[0120] Convolutional layer;
[0121] Four residual blocks, each containing two convolutional layers, with skip connections preserving the original features;
[0122] The feature normalization layer maps high-dimensional features to a unit hypersphere, enhancing distribution alignment stability.
[0123]
[0124] Parameter description:
[0125] Normalized feature vectors;
[0126] The feature extraction function of the first residual network is the convolutional layer parameter.
[0127] F2 layer weight matrix.
[0128] σ: LeakyReLU activation function.
[0129] More preferably, the feature extraction module configures a second residual network on the target domain client, specifically implemented as follows:
[0130] Input to the target domain residual network: a small number of labeled anchor points and unlabeled samples in: For the k-th input sample in the target domain, there are a total of R different transformer states, and each state has at least one label. composition.
[0131] The first residual network architecture specifically includes:
[0132] The network architecture is the same as that of the first residual network.
[0133] More preferably, the Gaussian mixture model modeling module, which fits a Gaussian mixture model to local features, is specifically implemented in the following ways:
[0134] Gaussian mixture model parameter estimation:
[0135] Source domain:
[0136] Each health state corresponds to a Gaussian component, and the features are fitted using the EM algorithm:
[0137]
[0138] Parameter description:
[0139] θ s Represents the parameters of the source domain Gaussian mixture model. Mixed weights, mean Covariance;
[0140] Target domain: Fitting G using the same process as described above. t (v t |θ t ), to obtain the Gaussian mixture model parameters θ in the target domain. t ;
[0141] After obtaining the Gaussian mixture model of the source and target domains, it is transmitted to the server for processing.
[0142] More preferably, the intermediate harmonic distribution generation module generates simulated samples, and the specific implementation includes:
[0143] Sampling is performed using a Gaussian mixture model of the source and target domains:
[0144] From G s () Sample m s One sample:
[0145] From G t () Sample m t One sample:
[0146] The above steps generate a simulated sample.
[0147] More preferably, the intermediate harmonic distribution generation module trains a stacked autoencoder using minimum transfer cost, specifically implemented as follows:
[0148] Training via stacked autoencoders:
[0149] (1) Sample input: and The data is merged into the input data and uniformly mapped to the latent space.
[0150] (2) Generation of latent variables:
[0151]
[0152] (3) Sample reconstruction: The decoder reconstructs samples from latent variables:
[0153]
[0154] Where the encoder is h enc (·) and decoder h dec (·);
[0155] (4) Loss function design:
[0156] Reconstruction loss: Minimizes the mean square error (MSE) of the input and output.
[0157]
[0158] Wasserstein distance loss: Calculates the source domain, target domain, and intermediate harmonic distribution G by minimizing the migration cost. m Distance:
[0159] L g =λ·OT(G s ||G m )+(1-λ)·OT(G t ||G m )
[0160] OT(G s ||G m Source domain distribution G s To the intermediate harmonic distribution G m The minimum migration cost.
[0161] OT(G t ||G m ): Target domain distribution G t To the intermediate harmonic distribution G m The minimum migration cost.
[0162] λ: Weighting coefficient, which balances the priority of alignment between the source and target domains.
[0163] More preferably, the minimum migration cost is specifically:
[0164]
[0165] Constraints:
[0166]
[0167] h dec (z j The samples generated by the decoder belong to the intermediate harmonic distribution G. m
[0168] z j : Variables in the latent space.
[0169] More preferably, the intermediate harmonic distribution generation module generates an intermediate harmonic distribution based on a trained stacked autoencoder, and its specific implementation includes:
[0170] (1) Latent variable mapping: mapping source domain features Input encoder, generate latent variables:
[0171]
[0172] (2) Generation of intermediate harmonic distribution samples: The decoder generates intermediate harmonic distribution samples from latent variables:
[0173]
[0174] (3) Label inheritance: Directly associate the source domain labels to obtain a labeled intermediate harmonic distribution sample set:
[0175]
[0176] More preferably, the local feature alignment module aligns the local features of the source and target domains to an intermediate harmonic distribution, and the specific implementation includes:
[0177] After receiving the intermediate harmonic distribution sample set in the source domain client:
[0178] (1) Calculate the intermediate harmonic distribution sample Transmission cost:
[0179]
[0180] (2) Backpropagation gradient update ResNet parameters:
[0181] Transmission costs As a loss function;
[0182] Directly calculate gradient Update ResNet parameters θ s This aligns the source domain features with a harmonic distribution towards the center.
[0183] After receiving the intermediate harmonic distribution sample set in the target domain client, perform the same operations described above to update the target domain ResNet parameters θ. t This aligns the target domain features with a harmonic distribution towards the center.
[0184] More preferably, the client-server collaborative training module performs client-server collaborative training and iterative optimization until convergence, and the specific implementation methods include:
[0185] The source and target domain clients respectively set GMM parameters;
[0186] The server generates a new intermediate tone distribution Gm by stacking an autoencoder and broadcasts it to the client.
[0187] Repeat the following steps until convergence:
[0188] Server update: Generating better G m ;
[0189] Client update: Optimize local model parameters for both the source and target domains;
[0190] Convergence condition: If the accuracy of the target domain validation set does not improve by 0.1% in 10 consecutive iterations, the iteration is stopped, and the converged target domain parameters are used for fault diagnosis.
[0191] More preferably, the data acquisition module employs an optimized flexible graphene vibration sensor, comprising: a flexible graphene vibration sensor body and a rigid weight attached to the flexible graphene vibration sensor body, used to enhance the output amplitude of the flexible graphene vibration sensor. The rigid weight converts external acceleration into a force on the flexible graphene elastomer through inertial effect. This force causes deformation of the elastomer. When the sensor is vibrated, the inertial force of the rigid weight drives the flexible structure to undergo compressive deformation, resulting in the reconstruction of its internal conductive network. The deformation causes the internal resistance of the material to change more rapidly, ultimately affecting the increase of the output amplitude of the output electrical signal.
[0192] It should be noted that the steps in the method provided by the present invention can be implemented using the corresponding components in the system. Those skilled in the art can refer to the technical solution of the system to implement the steps of the method, and can also refer to the technical solution of the method to implement the composition of the system. That is, the embodiments in the system and the embodiments in the method can be understood as preferred examples of each other, which will not be elaborated here.
[0193] The technical solution provided by the above embodiments of the present invention will be further described in detail below with reference to a specific application example.
[0194] This specific application example uses a transformer in a substation as an example. An optimized flexible graphene vibration sensor was used to collect transformer vibration signal data under four operating conditions (D1-D4) for experiments. All transfer tasks were defined as source domain → target domain. In the experiments, the transformer was subjected to various states, including normal (N), partial discharge (PD), low-energy discharge (F1), high-energy discharge (F2), medium-low temperature overheating (T1), and high-temperature overheating (T2). Transfer tasks were designed as S1-D1, S1-D2, S1-D3, and S1-D4. Transfer task S1-D1 means training the model on a labeled source dataset collected in the laboratory and transferring it to an unlabeled target dataset collected under operating condition D1. For example... Figure 3 As shown, the curves represent the relationship between the mass of the rigid object and the output amplitude of the sensor. The output amplitude reaches its maximum at 5g as the mass of the rigid object increases. Therefore, a mass of 5g is selected for the optimized flexible graphene vibration sensor used to acquire diagnostic signals.
[0195] To validate the model's performance, it is compared with the following neural networks: DRAN is a deep rational attention network with an embedded thresholding strategy, used for mechanical fault diagnosis. This network utilizes an attention mechanism to focus on the most important parts, thereby improving diagnostic accuracy, and is particularly suitable for mechanical fault diagnosis, especially performing well when handling complex time-series data. TCA is a kernel-based domain adaptation technique that achieves feature alignment by maximizing the similarity between the source and target domains. This method is suitable for unsupervised domain adaptation tasks, especially performing well when handling high-dimensional data. ADDA minimizes the feature distribution difference between the source and target domains through adversarial training, suitable for unsupervised domain adaptation tasks. This method is widely used in computer vision tasks such as image classification and object detection. DDC achieves domain adaptation by minimizing the domain confusion loss, making the model perform better in the target domain. This method is suitable for unsupervised domain adaptation tasks, especially performing well when handling image and text data. SAFN reduces the feature distribution difference between the source and target domains through self-adversarial feature normalization, improving the model's generalization ability. This method is suitable for unsupervised and semi-supervised domain adaptation tasks, especially performing well when handling high-dimensional data.
[0196] In this specific application example, the transfer learning-based transformer fault diagnosis method based on intermediate harmonic distribution includes the following steps:
[0197] Step S1: Acquire the vibration signal of the transformer using an optimized flexible graphene vibration sensor;
[0198] Step S2: Configure the first residual network in the source domain client to extract high-dimensional features of the labeled vibration signal, fit the Gaussian mixture model of the source domain features, and upload the Gaussian mixture model parameters to the server;
[0199] Step S3: Configure the second residual network on the target domain client to extract target domain sample features containing a small number of labeled anchor points and unlabeled samples, fit the target domain Gaussian mixture model and upload the parameters to the server;
[0200] Step S4: The server receives the Gaussian mixture model parameters of the source and target domains, generates simulated samples, and trains a stacked autoencoder with minimum transfer cost to generate an intermediate harmonic distribution medium.
[0201] Step S5: Broadcast the intermediate harmonic distribution samples to the client. Align the local feature distributions of the source domain and the target domain to the intermediate harmonic distribution by minimizing the minimum migration cost. Dynamically update the model parameters, encrypt them, and send them back to the server. Co-train, iteratively optimize until convergence, and perform transformer fault classification to complete the fault diagnosis of the target domain.
[0202] The specific implementation process of each of the above steps is the same as that of each step or the implementation method of each functional module given in the preferred embodiment of the above method or system of the present invention, and will not be repeated here.
[0203] like Figure 4 The diagram shows the training time, training accuracy, and error for each neural network model under various tasks. Figure 5 The image shows the average training accuracy of each network across various tasks. Figure 6 The diagram shown is an iterative graph of the loss function for each network in this invention.
[0204] The results show that the model proposed in this invention achieves the highest accuracy, demonstrating a significant advantage over other methods. While DRAN can improve diagnostic accuracy by focusing on important parts through an attention mechanism, it primarily relies on single-scale feature extraction, which may have limited effectiveness when handling multi-scale features and exhibits poor robustness under varying load and noise conditions. TCA achieves feature alignment by maximizing the similarity between the source and target domains, making it suitable for unsupervised domain adaptation tasks, but its performance is limited when handling temporal features in time-series data. ADDA minimizes the feature distribution difference between the source and target domains through adversarial training, but it performs poorly when handling complex operating conditions and noisy data, and is susceptible to negative transfer. DDC achieves domain adaptation by minimizing neighborhood confusion loss, but its effectiveness is limited when handling complex time-series data such as gas concentration signals. SAFN reduces the feature distribution difference between the source and target domains through self-adversarial feature normalization, but it has limitations when handling data under different operating conditions.
[0205] During training, to investigate the impact of the Wasserstein distance loss weight parameters on model accuracy, experiments were conducted on the relationship between the weight parameters and diagnostic accuracy. For example... Figure 7 As shown, the model achieves optimal accuracy on the dataset when the weight parameter is set to 0.55. Therefore, this parameter configuration was chosen.
[0206] An embodiment of the present invention also provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it can be used to execute the method of any of the above embodiments of the present invention, or to run the system of any of the above embodiments of the present invention.
[0207] Optionally, the memory is used to store programs; the memory may include volatile memory, such as random-access memory (RAM), such as static random-access memory (SRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM), etc.; the memory may also include non-volatile memory, such as flash memory. The memory is used to store computer programs (such as application programs, functional modules, etc. that implement the above methods), computer instructions, etc., and the aforementioned computer programs, computer instructions, etc., can be partitioned and stored in one or more memories. Furthermore, the aforementioned computer programs, computer instructions, data, etc., can be accessed by the processor.
[0208] A processor is used to execute computer programs stored in memory to implement the various steps of the methods or various modules of the systems involved in the above embodiments. For details, please refer to the relevant descriptions in the preceding method and system embodiments.
[0209] The processor and memory can be separate structures or integrated structures. When the processor and memory are separate structures, they can be coupled together via a bus.
[0210] An embodiment of the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, can be used to perform the method of any of the above embodiments of the present invention, or to run the system of any of the above embodiments of the present invention.
[0211] Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of computer programs from one place to another. Storage media can be any available medium accessible to a general-purpose or special-purpose computer. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. Alternatively, the ASIC can reside in a user device. Of course, the processor and storage medium can also exist as discrete components in a communication device.
[0212] The above embodiments of the present invention provide a method and system for transformer fault diagnosis based on intermediate harmonic distribution using transfer learning. This method utilizes the vibration signal of a transformer obtained from an optimized flexible graphene vibration sensor. The vibration signal of the transformer undergoing fault diagnosis is used as the target domain, while vibration signals collected during laboratory fault experiments are used as the source domain. Local feature extraction modules are deployed on both the source domain transformer and the target domain transformer client, constructing a federated learning framework. High-dimensional features of the transformer vibration signal are extracted using a deep residual network, and a Gaussian mixture model is fitted. On the server side, based on the intermediate harmonic distribution mechanism of the Gaussian mixture model and stacked autoencoders, the parameter information of the feature distributions in the source and target domains is aggregated. The autoencoder generates an intermediate distribution with the minimum distance to the dual-domain distribution as a transfer medium. Local feature distributions are aligned to the medium by minimizing the minimum transfer cost, driving local features to align to the intermediate distribution. Finally, a collaborative optimization and fault diagnosis process between the client and server is implemented. The model parameters are iteratively updated, encrypted, and sent back to the server. Collaborative training and iterative optimization continue until convergence, and transformer fault classification is performed, thus completing the fault diagnosis in the target domain. Experimental results show that this method can effectively perform transfer diagnosis of transformers, significantly improving the accuracy and reliability of predictions.
[0213] Any matters not covered in the above embodiments of the present invention are well-known in the art.
[0214] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the specific embodiments described above, and those skilled in the art can make various modifications or variations within the scope of the claims, which do not affect the essence of the present invention.
Claims
1. A transfer learning method for transformer fault diagnosis based on intermediate harmonic distribution, characterized in that, include: The source domain transformer vibration signal and the target domain transformer vibration signal are acquired separately. Configure the first residual network on the source domain client, extract the high-dimensional features of the source domain transformer vibration signal, fit the Gaussian mixture model of the source domain features, and upload the parameters of the Gaussian mixture model to the server. Configure a second residual network on the target domain client, extract sample features of the transformer vibration signal in the target domain, fit a Gaussian mixture model of the target domain features, and upload the parameters of the Gaussian mixture model to the server. The server receives the Gaussian mixture model parameters of the source domain and the target domain, generates simulated samples, and trains a stacked autoencoder with minimum transfer cost to generate intermediate harmonic distribution samples. The intermediate harmonic distribution samples are broadcast to the client. The source domain and target domain clients align their local feature distributions to the intermediate harmonic distribution by minimizing the minimum migration cost, dynamically update the corresponding residual network parameters, and perform collaborative training and iterative optimization on the client and server until convergence. Transformer fault classification is then performed to complete fault diagnosis in the target domain.
2. The method for transformer fault diagnosis based on intermediate harmonic distribution according to claim 1, characterized in that, The acquisition of source domain transformer vibration signals and target domain transformer vibration signals respectively includes: An optimized flexible graphene vibration sensor was used to collect the vibration signal of the transformer in the fault experiment as the source domain transformer vibration signal, which is a fully labeled transformer vibration signal. The vibration signal of the transformer used for fault diagnosis is collected by an optimized flexible graphene vibration sensor as the target domain transformer vibration signal, which is a partially marked or unmarked transformer vibration signal. in: The optimized flexible graphene vibration sensor includes a flexible graphene vibration sensor body and a rigid weight attached to the flexible graphene vibration sensor body.
3. The method for transformer fault diagnosis based on intermediate harmonic distribution according to claim 1, characterized in that, The configuration of the first residual network on the source domain client includes: The input data for the first residual network is: a set of fully labeled transformer vibration signals. in, This is the i-th input sample in the source domain transformer vibration signal; For the corresponding The label indicates the target category of the vibration signal; n s This represents the total number of tagged vibration signal samples in the source domain client. The output data of the first residual network is: The network architecture for constructing the first residual network includes: four residual blocks arranged sequentially; each residual block includes two convolutional layers connected by skip connections; each convolutional layer is followed by a feature normalization layer; the feature normalization layer is used to map the high-dimensional features of the input data to a unit hypersphere, as shown below: In the formula, These are the normalized feature vectors; Let represent the feature extraction function of the first residual network, and be the parameters of the convolutional layer; Here is the weight matrix of layer F2; σ is the LeakyReLU activation function; The Gaussian mixture model that fits the source domain features includes: The parameters of the Gaussian mixture model are estimated, where each healthy state corresponds to a Gaussian component. The source domain features are fitted using the EM algorithm, and the result is expressed as follows: In the formula, G s (v s |θ s ) represents a Gaussian mixture model in the source domain; v s Let θ be the input random variable of the source domain Gaussian mixture model, representing the distribution of the source domain eigenvectors; s For the parameters of the source domain Gaussian mixture model, in, For mixed weights, The mean, R represents the covariance; R is the total number of transformer fault state categories; N is the Gaussian distribution symbol.
4. The method for transformer fault diagnosis based on intermediate harmonic distribution according to claim 1, characterized in that, The configuration of the second residual network on the target domain client includes: The input to the second residual network is: a set of partially labeled anchor points. and unlabeled sample set in, For the k-th input sample in the target domain transformer vibration signal, there are a total of R different transformer state types, and each state is labeled with at least one label. composition; The j-th unlabeled vibration signal sample in the target domain; n t This represents the total number of unlabeled samples in the target domain. The output of the second residual network is: The network architecture for constructing the second residual network includes: four residual blocks arranged sequentially; each residual block includes two convolutional layers connected by skip connections; each convolutional layer is followed by a feature normalization layer; the feature normalization layer maps the high-dimensional features of the input data to a unit hypersphere, as shown below: In the formula, These are the normalized feature vectors; Let represent the feature extraction function of the first residual network, and be the parameters of the convolutional layer; Here is the weight matrix of layer F2; σ is the LeakyReLU activation function; The Gaussian mixture model for fitting the features of the target domain includes: The parameters of the Gaussian mixture model are estimated, where each healthy state corresponds to a Gaussian component. The target domain features are then fitted using the EM algorithm, as follows: In the formula, G t (v t |θ t ) represents a Gaussian mixture model for the target domain; v t Vibration signal in the target domain The feature representation is extracted and normalized using the second residual network; θ t For the parameters of the Gaussian mixture model in the target domain, in, For mixed weights, The mean, R represents the covariance; R is the total number of transformer fault state categories; N is the Gaussian distribution symbol.
5. The method for transformer fault diagnosis based on intermediate harmonic distribution according to claim 1, characterized in that, The generation of simulated samples includes: m is sampled using the source domain Gaussian mixture model. s From a sample, a simulated sample is obtained. Sample m using the Gaussian mixture model of the target domain t From a sample, a simulated sample is obtained. The step of training a stacked autoencoder with minimum transfer cost includes: Training via stacked autoencoders includes: simulated samples and The data is merged as input data to the autoencoder. The encoder maps the variables uniformly to the latent space, generating latent variables z. i : In the formula, h enc (·) represents the encoder; This is the set of simulated samples sampled from the source domain Gaussian mixture model; The set of simulated samples sampled from the Gaussian mixture model of the target domain; The decoder utilizes the hidden variable z i Perform sample reconstruction to obtain reconstructed samples. In the formula, h dec (·) represents the decoder; Establish the loss function, including: Establish reconstruction loss L rccon The mean square error of the input and output is minimized, expressed as: In the formula, m s Number of samples generated for the source domain; m t Generate the number of samples for the target domain; Reconstructing samples using an autoencoder; Establish Wasserstein distance loss L g The source domain, target domain, and intermediate harmonic distribution G are calculated using the minimum migration cost. m The distance is expressed as: L g =λ·OT(G s ||G m )+(1-λ)·OT(G t ||G m ) In the formula, OT(G s ||G m ) represents the source domain distribution G s To the intermediate harmonic distribution G m The minimum migration cost; OT(G t ||G m ) represents the target domain distribution G t To the intermediate harmonic distribution G m The minimum migration cost; λ represents the weight coefficient, which is the priority for balancing the alignment of the source and target domains; The minimum migration cost is: In the formula, Γ is the transmission plan matrix; i,j To make G s The transfer of the i-th sample to G m The proportion of the j-th sample; C i,j Indicates transmission cost; The constraints are as follows: h dec (z j The sample generated by the decoder belongs to the intermediate harmonic distribution G. m ;z j These are latent variables; The generation of intermediate harmonic distribution samples includes: Source domain features Input the trained stacked autoencoder to generate latent variables The decoder utilizes the hidden variables Perform sample reconstruction to generate intermediate harmonic distribution samples. The intermediate harmonic distribution sample By directly associating the source domain labels, a labeled intermediate harmonic distribution sample set is obtained.
6. The method for transformer fault diagnosis based on intermediate harmonic distribution according to claim 1, characterized in that, The source and target domain clients align their local feature distributions to an intermediate harmonic distribution by minimizing the minimum migration cost, including: Calculate source domain features With intermediate harmonic distribution samples Transmission cost In the formula, The transmission cost required to migrate the i-th feature sample from the source domain to the j-th sample in the intermediate harmonic distribution; Transmission costs As a loss function; Calculate gradient Update ResNet parameters θ s This aligns the source domain features with a harmonic distribution towards the center. Using the same computational method and backpropagation gradient update method to align the target domain features with the central harmonic distribution.
7. The method for transformer fault diagnosis based on intermediate harmonic distribution according to claim 1, characterized in that, The process of collaboratively training and iteratively optimizing the client and server until convergence includes: The server continuously generates new intermediate tonal distribution samples by stacking autoencoders and broadcasts them to the client; The client optimizes the local residual network parameters by continuously updating intermediate harmonic distribution samples, fits a Gaussian mixture model of new local features, and uploads the parameters of the Gaussian mixture model to the server to participate in the generation of new intermediate harmonic distribution samples. Repeat the above process for collaborative training and iterative optimization until the convergence condition is met.
8. A transformer fault diagnosis system based on intermediate harmonic distribution using transfer learning, characterized in that, include: The feature extraction module is used to configure the first residual network on the source domain client and extract high-dimensional features of the source domain transformer vibration signal. Configure a second residual network on the target domain client to extract sample features of the transformer vibration signal in the target domain; The Gaussian Mixture Model (GMMM) modeling module is used to fit a Gaussian mixture model of source domain features on the source domain client and upload the parameters of the Gaussian mixture model to the server; and to fit a Gaussian mixture model of target domain features on the target domain client and upload the parameters of the Gaussian mixture model to the server. An intermediate harmonic distribution generation module is used to receive Gaussian mixture model parameters of the source domain and the target domain on the server side, generate simulated samples, train a stacked autoencoder with minimum transfer cost to generate intermediate harmonic distribution samples, and broadcast the intermediate harmonic distribution samples to the client. A local feature alignment module is used to align the local feature distribution to an intermediate harmonic distribution by minimizing the minimum migration cost. The client-server collaborative training module is used to dynamically update the corresponding residual network parameters, perform collaborative training and iterative optimization on the client and server until convergence, and perform transformer fault classification to complete fault diagnosis in the target domain.
9. A computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it can be used to perform the method of any one of claims 1-7, or to run the system of claim 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program can be used to perform the method of any one of claims 1-7, or to run the system of claim 8.