A transformer fault diagnosis method based on adaptive fault attention residual network
By using adaptive fault attention residual network and MEMS photoacoustic sensor in transformer fault diagnosis, combined with local maximum mean difference and domain discriminator, the problem of insufficient generalization ability in the existing technology is solved, high accuracy and adaptability are achieved, and the transparency of the model is enhanced.
Patent Information
- Application Number
- CN202411469846.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-21
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-10-21
AI Technical Summary
When existing transformer fault diagnosis technology processes data under different operating conditions, the model has insufficient generalization capabilities and lacks transparency and interpretability.
The method based on the adaptive fault attention residual network is adopted, and the attention module is improved through the adaptive activation function and the fault index function, data is collected in combination with the MEMS photoacoustic sensor, and feature distribution is adjusted using local maximum mean difference and domain discriminator to optimize model performance.
It improves the accuracy of fault diagnosis and generalization capabilities of the model, enhances the adaptability to domain offsets, and realizes the transparency and interpretability of the model.
Smart Images

Figure CN119513649B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of transformer fault diagnosis, and in particular to a transformer fault diagnosis method based on an adaptive fault attention residual network. Background Art
[0002] In power systems and industrial applications, transformers are core components, and their operating status is directly related to the stability and reliability of the overall system. As the equipment ages and wears, transformers may experience faults such as overheating and partial discharge, which may lead to the generation of dissolved gas in the oil. Traditionally, transformer fault diagnosis relies on expert experience and knowledge in specific fields. Although this method is intuitive, it is time-consuming and has limited diagnostic accuracy.
[0003] In recent years, with the development of artificial intelligence technology, especially the advancement of deep learning algorithms, intelligent fault diagnosis technology has developed rapidly. It can automatically extract fault features and identify different fault types through training models, realizing real-time monitoring and predictive maintenance of key equipment such as transformers.
[0004] In response to the needs of transformer fault diagnosis, researchers have explored a variety of solutions. For example, Jin et al. proposed a style normalization and restoration method, which achieved effective domain adaptation by separating and normalizing the style information in the source domain and then restoring it to the target domain. Wen et al. developed a hierarchical domain adaptation method to improve the generalization ability of the model through the use of local feature patterns. Hu et al. proposed a transfer learning method based on deep feature decoupling to predict the remaining service life of transformers under different working conditions, further improving the adaptability and diagnostic ability of the model. These methods improve the performance of the model in different environments from different angles, providing valuable ideas for solving the domain adaptation problem.
[0005] In terms of sensor technology, especially for transformer fault diagnosis, MEMS (micro-electromechanical system) photoacoustic sensors have shown great potential. Due to their small size, low cost and fast response, MEMS photoacoustic sensors have broad application prospects in the condition monitoring and fault diagnosis of power equipment. For example, the acoustic signals obtained by MEMS photoacoustic sensors can be used to detect abnormal conditions such as partial discharge inside the transformer, and then provide fault warning. Studies have shown that MEMS photoacoustic sensors can capture weak acoustic signals, which is crucial for early detection of potential faults. In addition, some studies have proposed an online monitoring system based on MEMS photoacoustic sensors, which can continuously monitor the operating status of the transformer and detect abnormalities in a timely manner, thereby improving the safety of the system.
[0006] Although existing technologies have improved the efficiency and accuracy of fault diagnosis to a certain extent, they still face many challenges in practical applications, especially when processing data from different operating conditions, the generalization ability of the model is particularly important. In addition, how to make the model more transparent and explainable is also a focus of researchers. Summary of the invention
[0007] In order to solve the deficiencies of the above-mentioned prior art, the present invention proposes a transformer fault diagnosis method and system based on an adaptive fault attention residual network, aiming to further improve the accuracy of fault diagnosis and the generalization ability of the model. The method first designs an adaptive fault attention mechanism, and improves the traditional attention module by adding an adaptive activation function and a fault index function, wherein the adaptive activation function reduces the attention weight of irrelevant features, and the fault index function enhances the weight of features related to diagnosis. Then, techniques such as local maximum mean difference (LMMD) and domain discriminator are used to adjust the edge and conditional distribution, so that the features between the source domain and the target domain are more consistent, so that the model has generalization ability and adaptability to domain shift. Finally, the model performance is optimized by weighing the hyperparameters, and the losses in various aspects are balanced to achieve dynamic balance.
[0008] To achieve the above objectives, this application provides the following solutions:
[0009] A transformer fault diagnosis method based on an adaptive fault attention residual network is characterized in that it includes the following steps:
[0010] Step 1: Use a MEMS photoacoustic sensor with a capacitive resonator to analyze the acoustic wave signal through photoacoustic spectroscopy and signal processing technology to collect dissolved gas in the oil and form a parameter sequence.
[0011] Step 2: An adaptive fault attention mechanism is designed to improve the traditional attention module by adding an adaptive activation function and a fault index function. The adaptive activation function reduces the attention to diagnosis-irrelevant features, and the fault index function enhances the attention weight to diagnosis-related features to generate diagnostic features that are highly correlated with fault features.
[0012] Step 3: Use the local maximum mean difference and domain discriminator to adjust the feature distribution of the source domain and the target domain to enable the model to have generalization ability and domain adaptability.
[0013] Step 4: Optimize model performance by weighing hyperparameters and balance the losses in all aspects to achieve dynamic balance.
[0014] Furthermore, in the step 1, the MEMS photoacoustic sensor with a capacitive resonator has a large number of comb-tooth structures for increasing the sensitivity of capacitive detection, and the gap between the movable structure and the substrate is increased to reduce gas damping.
[0015] When the resonator vibrates, the overlapping area between the movable and fixed teeth changes. The direction of motion of the resonator is perpendicular to the direction of the electric field lines, which separates the acoustic wave excitation and capacitive sensing. By increasing the anchor height of the resonator, the distance between the acoustic wave receiving area and the glass substrate is increased, which reduces the gas damping effect on the resonator motion. At the same time, the distance between the two electrodes can still be designed to be small to ensure high sensitivity of detection.
[0016] Further, the adaptive activation function in the adaptive fault attention mechanism designed in step 2 is designed to help reduce the attention weight of diagnosis-irrelevant features during model training. The activation function in the attention mechanism is modified by an adaptive threshold that satisfies the following properties: Shrinkage: The attention weight of diagnosis-irrelevant features should be small enough. Maintain: The attention weight of diagnosis-related features remains unchanged.
[0017] The defined adaptive activation function and its derivative form are as follows:
[0018]
[0019] Where x is the input value, y is the output value, and τ is a learnable adaptive threshold. Since the function is not differentiable at x = τ, its derivative value is artificially set to 1 for ease of training. In addition, when the number of neurons is small, network training may be affected by zero gradients, so the function can take the form of a non-zero derivative.
[0020] Furthermore, the fault index function in the adaptive fault attention mechanism designed in step 2 is designed to help increase the attention weight of diagnosis-related features. In the MEMS-based photoacoustic sensor transformer fault diagnosis, the diagnosis-related features are the changes in various gas concentrations and their ratios to diagnose the fault type inside the transformer; therefore, the fault index function should contain the following properties:
[0021] The fault index function should be sensitive to changes in gas concentration and its ratio in the transformer fault signal. In order to increase the attention weight of the periodic points during model training, a concept of the ratio of fault-related gas components to non-fault-related gases (FRG) is introduced. The ratio of fault-related gas components to non-fault-related gases (FRG) is added as a loss function, which can be the ratio of fault-related gas components to non-fault-related gas components. When the refined feature map contains more periodic components, the value of FRG will increase, which can be calculated by the autocorrelation function. Assuming r x (0) is the autocorrelation function r x The global maximum value of (τ), its normalized form r x′(τ) can be defined as formula (2)
[0022]
[0023] where r x (τ) is the autocorrelation function of sample x(n) as shown in equation (3), n==1,2,…,N, N is the number of sample points, is the sample mean.
[0024]
[0025] The FRG can be determined by identifying the first local maximum corresponding to the fault-related gas component, as shown in Equation (4):
[0026]
[0027] where r x ′(τ max ) is r x The first local maximum of ′(τ).
[0028] In order to capture the abnormal gas components in the fault signal, kurtosis is used to detect the peak in the signal, that is, the peak value of the abnormal gas concentration. During the model training process, a function that is sensitive to changes in gas concentration is added as a loss function.
[0029]
[0030] Where x(n) is the sample, n = 1, 2, ..., N, N is the number of sample points, is the mean value, σ x is the standard deviation.
[0031] Furthermore, the adaptive fault attention mechanism in step 2 is proposed in step 2. The proposed adaptive fault attention mechanism mainly includes a sequence attention block and a channel attention block.
[0032] Sequential Attention Module, the feature map first passes through a 1×1×1 convolution layer and a sigmoid activation function to generate sequential attention weights. Specifically, moments related to faults will receive higher attention weights. These attention weights are then input into an adaptive activation function, which resets the attention weights of features not related to diagnosis to zero. At the same time, a 64×1×1 convolution layer is added to enrich feature information. Finally, the output results are multiplied with each other, and residual connections are added to prevent information loss and gradient disappearance.
[0033] The channel attention module first uses a global average pooling layer to generate descriptors that indicate the channel relationship between feature maps. These descriptors are then fed into a multi-layer perceptron (MLP) and a sigmoid activation function to calculate the channel attention weights. Finally, similar to the sequence attention approach, the channel attention weights are also passed through an adaptive activation function and a residual connection is used to form an improved feature map.
[0034] In addition, a transformer fault index loss L is added during model training. tf To increase the attention weight of diagnosis-related features. Specifically, the weighted feature map is first sent to the global maximum pooling layer to obtain a pooled feature map of size (n, 1) Then, according to formulas (2) to (8), To calculate:
[0035] L tf =|avg(HNR)| (8)
[0036] where avg(·) represents the average value in each batch, and |.| represents the absolute value.
[0037] By calculating the mean and absolute value of these statistics, the model can more effectively identify patterns associated with failures.
[0038] Furthermore, the adaptive fault attention mechanism residual network model in step 3 uses the local maximum mean difference and domain discriminator to adjust the source domain and target domain feature distribution, so that the model has generalization ability and domain adaptability.
[0039] represents the source domain, where is the i-th source sample and its corresponding label These samples are drawn from the distribution P s Extract from the middle, n s Indicates the number of source samples. Source samples and |C s |Category associated, where C s Represents the label space of the source domain.
[0040] let represents the target domain, where is the i-th unlabeled sample in the target domain, which is drawn from the distribution P t Extract from the middle, n t represents the number of target samples. The target samples and |C t |Category associated, where C t Represents the label space of the target domain.
[0041] Assume that the data in the source domain and the target domain are sufficient and obey different marginal probability distributions, that is, P s ≠P t The source domain and the target domain share the same label space, namely, C s =C t ,The number of samples in all categories in the source and target ,domains are balanced.
[0042] Among them, design a function or neural network f * To minimize the difference between marginal distribution and conditional distribution, and learn invariant features across domains, and According to the covariate shift assumption, f * Be consistent across areas.
[0043] LMMD and domain discriminator are used to minimize the conditional and marginal distribution differences of features. LMMD is inspired by the maximum mean difference (MMD), which uses kernel embedding in the reproducing kernel Hilbert space (RKHS) to measure data distribution and is defined as follows:
[0044]
[0045] in Is the source domain X s The i-th sample in is the target domain X t The i-th sample in n s and n t are the number of source samples and target samples, respectively. is the Gaussian kernel function that maps samples to RKHS.
[0046] LMMD embeds the category weights derived from pseudo labels, further aligns the subdomain distributions, and uses the loss function L LMMD The definition is as follows:
[0047]
[0048] Where C=|C s |=|C t | is the number of categories in the dataset, and They are and The weight corresponding to the category, φ(x) is the Gaussian kernel function that maps the sample to RKHS.
[0049] The domain discriminator was originally used by the Domain Adversarial Neural Network (DANN) to generate domain-invariant features, where the feature extractor G makes the domain discriminator D unable to correctly classify the domain label until a Nash equilibrium is reached. The loss function L d The definition is as follows:
[0050]
[0051] where g i is the true domain label of the i-th sample, g i ∈{0,1},D(G(x i ) is the domain label predicted for the i-th sample, D(G(x i ))∈{0,1}, m is the number of samples in the source data and the target data.
[0052] Finally, for the classification loss of the source data It is obtained through the cross entropy function, as shown in formula (12):
[0053]
[0054] in is the i-th sample label in the source domain, is the kth element in the classifier output, n s is the number of source samples, and 1{·} represents the indicator function.
[0055] Furthermore, in step 4, the model performance is optimized by weighing the hyperparameters and balancing the losses in various aspects to achieve dynamic balance.
[0056] In the proposed adaptive fault attention mechanism residual network training process, the following four optimization objectives are included: Objective 1: Minimize the use of source domain data Objective 2: Maximize L using source and target domain data d ; Objective 3: Minimize L using source and target domain data LMMD Objective 4: Maximize L using source and target domain data tf The final optimization goal can be expressed by combining the above four optimization goals, as shown in formula (13):
[0057]
[0058] Among them, λ, μ, and γ are trade-off hyperparameters. These hyperparameters are used in the model to balance the mutual influence between different loss function components, ensuring that the model can achieve an optimal compromise in performing such operations as minimizing the classification error of the source domain, maximizing the difference between domains, minimizing the maximum mean difference between domains, and maximizing the benefits brought by the attention mechanism.
[0059] Compared with the prior art, the present invention has the following beneficial effects:
[0060] 1) Traditional transformer fault diagnosis mostly relies on electrical parameters or oil sample chemical analysis. The present invention combines MEMS photoacoustic sensors with photoacoustic spectroscopy for the first time to collect parameter data of dissolved gases in transformer oil. This sensor has a capacitive resonator design, which increases the sensitivity of capacitive detection through a large number of comb teeth structures, while optimizing the gap between the movable structure and the substrate to reduce gas damping, thereby achieving high-precision and high-sensitivity gas detection.
[0061] 2) An adaptive fault attention mechanism is designed to improve the traditional attention module by adding adaptive activation functions and fault index functions. This mechanism is able to generate diagnostic features that are highly correlated with fault characteristics, improving the accuracy and efficiency of fault diagnosis. Moreover, the adaptive activation function dynamically adjusts the input data characteristics by introducing a learnable adaptive threshold, reducing the attention to features that are not related to diagnosis. The fault index function uses statistics such as the ratio of fault-related gas components to non-fault-related gases (FRG) and kurtosis to enhance the attention weight of diagnosis-related features. This design not only improves the diagnostic ability of the model, but also enhances its adaptability and robustness.
[0062] 3) Fault diagnosis across working conditions is achieved by using transfer learning technology. The feature distribution of the source domain and the target domain is adjusted through the local maximum mean difference (LMMD) and the domain discriminator, so that the model has generalization ability and domain adaptability. The problem of performance degradation of traditional fault diagnosis models under different working conditions is solved. By minimizing the conditional and marginal distribution differences of features, the model can learn invariant features across domains and show good generalization ability on the target domain. It not only improves the practicality of the model, but also reduces the dependence on a large amount of labeled data.
[0063] 4) Optimize model performance by weighing hyperparameters and balance the losses in various aspects to achieve dynamic balance. This strategy covers multiple optimization objectives such as minimizing source domain classification error, maximizing inter-domain differences, minimizing the maximum mean difference between domains, and maximizing attention mechanisms. It ensures that the model can reach an optimal compromise state when performing different tasks. By fine-tuning hyperparameters, the model can improve the accuracy of fault diagnosis while maintaining good generalization ability and adaptability. This strategy provides new ideas and methods for model optimization in the field of transformer fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 The present invention is a flow chart of a transformer fault diagnosis method based on an adaptive fault attention residual network (RNB-AFAM).
[0065] Figure 2 It is the residual network model of the adaptive fault attention mechanism of the present invention.
[0066] Figure 3 The diagnostic accuracy and training time of the model of the present invention are compared with those of various advanced networks under different diagnostic tasks.
[0067] Figure 4 is the confusion matrix of this model.
[0068] Figure 5 It is the t-sne diagram of the present invention.
[0069] Figure 6 This is a structural diagram of a MEMS photoacoustic sensor with a capacitive resonator according to the present invention. DETAILED DESCRIPTION
[0070] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0071] The migration task is defined as source domain → target domain. The source domain and target domain represent the transformer parameter data collected under different working conditions. The goal of transfer learning is to use the model trained on the source domain (i.e., known or labeled dataset) to improve the prediction performance on the target domain (i.e., unknown or unlabeled dataset).
[0072] The transformer parameter data collected by the MEMS photoacoustic sensor with a capacitive resonator under four working conditions (D1-D4) in this embodiment. The various states of the transformer monitored include normal state (N), partial discharge (PD), low energy discharge (F1), high energy discharge (F2), medium and low temperature overheating (T1) and high temperature overheating (T2). Each fault mode contains 409,600 data points, which are divided into 100 samples, each of which contains 4096 data points. In the data set, each sample consists of 4096 data points containing sufficient fault information, and the number of samples obtained under each fault is 1000. The data set is divided into training set and test set in a ratio of 7:3. The transfer tasks are designed as D1-D2, D1-D3, D1-D4, D2-D3, D2-D4, and D3-D4. Each task means that the model needs to be trained on a labeled data set collected under one working condition (source domain), and then migrated to an unlabeled data set collected under another working condition (target domain) for prediction. For example, transferring task D1-D2 means that the model is trained on a labeled source dataset collected under working condition D1 and transferred to an unlabeled target dataset collected under working condition D2.
[0073] See also Figure 1 , Figure 1 This is a flow chart of the transformer fault diagnosis method based on the adaptive fault attention residual network (RNB-AFAM) of the present invention. As shown in the figure, the method uses a MEMS photoacoustic sensor to collect transformer parameter data and realizes fault diagnosis across working conditions through transfer learning technology. The following is a detailed introduction:
[0074] Step S1: Data acquisition: A MEMS photoacoustic sensor with a capacitive resonator is used to analyze the acoustic wave signal through photoacoustic spectroscopy and signal processing technology to collect dissolved gas in the oil and form a parameter sequence.
[0075] Among them, the MEMS photoacoustic sensor with capacitive resonator is used to increase the capacitance detection sensitivity through a large number of comb-tooth structures. The gap between the movable structure and the substrate is increased to reduce gas damping.
[0076] Step S2: Design an adaptive fault attention mechanism, improve the traditional attention module by adding adaptive activation function and fault index function, and generate diagnostic features that are highly correlated with fault features; wherein, the adaptive activation function is used to reduce the attention of features irrelevant to diagnosis, and the activation function is modified by adaptive threshold to achieve shrinkage and maintenance properties. The fault index function is used to enhance the attention weight of features related to diagnosis, and detect abnormal gas components by the ratio (FRG) and kurtosis of fault-related gas components to non-fault-related gases. The attention module, including the sequence attention block and the channel attention block, generates attention weights using convolutional layers, sigmoid activation functions, multi-layer perceptrons (MLPs), etc., and forms improved feature maps through adaptive activation functions and residual connections.
[0077] Specifically: The adaptive activation function in the adaptive fault attention mechanism designed in step 2 is intended to help reduce the attention weight of features not related to diagnosis during model training. An adaptive activation function is a function designed to dynamically adjust its activation behavior based on the characteristics of the input data. The adaptive activation function enhances flexibility by introducing a learnable adaptive threshold, enabling the function to adjust its behavior based on the characteristics of the input data.
[0078] Properties of the adaptive activation function, including:
[0079] Shrinkage: For features that are not relevant to diagnosis (i.e., those that have little impact on the final output), the adaptive activation function should be able to reduce the attention weights of these features, i.e., the attention weights should be small enough.
[0080] Maintain: For features that are directly relevant to diagnosis, the adaptive activation function should keep the attention weights of these features unchanged or minimize their changes.
[0081] The definition of the adaptive activation function and its derivative form are as follows:
[0082]
[0083] Where x is the input value, y is the output value, and τ is the learnable adaptive threshold.
[0084] Since the adaptive activation function is not differentiable at x = τ (i.e., its derivative does not exist or is difficult to calculate), the derivative value at this point is usually set to 1 for ease of training. In addition, when the number of neurons is small, network training may be affected by zero gradients, because zero gradients mean that no gradient signal can be passed to the previous layer at this point, thereby preventing the update of parameters. To avoid this, the adaptive activation function can take the form of a non-zero derivative, even at the threshold. This ensures that throughout the training process, no matter how the input value changes, there is a gradient signal that can be passed to the previous layer, allowing the model to continue learning and optimizing.
[0085] The fault index function in the adaptive fault attention mechanism designed in step 2 is intended to increase the attention weight of features related to diagnosis, especially the gas concentration and its ratio change characteristics for the diagnosis of internal fault types of transformers. Therefore, the fault index function should contain the following attributes:
[0086] The core purpose of the fault index function is to sensitively capture the gas concentration and ratio changes in the transformer fault signal in order to more accurately diagnose the fault type inside the transformer. In transformer fault diagnosis, the concentration and ratio changes of different gases (such as hydrogen, methane, ethane, ethylene, acetylene, etc.) can provide important information about the fault type (such as overheating, discharge, arc, etc.).
[0087] Ratio of fault-related gas components to non-fault-related gases (FRG)
[0088] In order to increase the attention weight of the periodicity, the concept of FRG is introduced. FRG is the ratio of fault-related gas components to non-fault-related gas components and is added to the model as a loss function. When the refined feature map contains more periodic components, the value of FRG increases, which can be calculated by the autocorrelation function.
[0089] The autocorrelation function is used to measure the similarity between a signal and itself at different time delays. The global maximum reflects the periodic component of the signal. Let r x (0) is the autocorrelation function r x The global maximum value of (τ), its normalized form r x ′(τ) can be defined as follows:
[0090]
[0091] Among them, r x (τ) is the autocorrelation function of sample x(n), as shown in formula (3):
[0092]
[0093] in, is the sample mean, x(n) is the sample, n == 1, 2, …, N, N is the number of sample points;
[0094] The value of FRG is determined by identifying the first local maximum corresponding to the fault-related gas component, as shown in equation (4):
[0095]
[0096] Among them, r x ′(τ max ) is r x The first local maximum of ′(τ) reflects the intensity of periodic changes of specific gas components in the fault signal.
[0097] In order to capture the abnormal gas components in the fault signal, kurtosis is used to detect the peak in the signal, that is, the peak value of the abnormal gas concentration. In transformer fault diagnosis, the peak value of the abnormal gas concentration may indicate a specific fault type or fault degree.
[0098] Calculate the kurtosis value of the sample and add the kurtosis function as a loss function to the model training to increase the model's sensitivity to changes in gas concentration, especially abnormal peaks.
[0099]
[0100] Among them, σ x is the standard deviation.
[0101] The fault index function improves the model's sensitivity to changes in gas concentration and its ratio in transformer fault signals by combining statistics such as FRG and kurtosis. FRG captures the periodic components of the signal through the autocorrelation function, while kurtosis is used to detect abnormal peaks in the signal. These characteristics make the fault index function more accurate and reliable in MEMS-based photoacoustic sensor transformer fault diagnosis. During the model training process, by adding these statistics as part of the loss function, the model's ability to identify fault features can be optimized.
[0102] The adaptive fault attention mechanism designed in step 2 includes two modules: sequence attention and channel attention.
[0103] Sequential attention module: First, the feature map passes through a 1×1×1 convolution layer and a Sigmoid activation function to generate sequential attention weights. Specifically, moments related to faults will receive higher attention weights. These attention weights are then input into an adaptive activation function, which resets the attention weights of features not related to diagnosis to zero. At the same time, a 64×1×1 convolution layer is added to enrich feature information. Finally, the adaptively activated attention weights are multiplied with the original feature map (or the feature map after the 64×1×1 convolution layer) to achieve feature weighting. And residual connections are added to prevent information loss and gradient disappearance.
[0104] Channel attention module: First, a global average pooling layer is used to generate descriptors. The global average pooling layer averages the feature maps of each channel to generate a descriptor that represents the overall information of the channel. These descriptors indicate the channel relationship between feature maps, that is, which channels may be associated or complementary. These descriptors are fed into a multi-layer perceptron (MLP) and a Sigmoid activation function to calculate the channel attention weights. The Sigmoid activation function ensures that the channel attention weights are between 0 and 1. Finally, similar to the sequence attention approach, the channel attention weights are also passed through an adaptive activation function and a residual connection is used to form an improved feature map, that is, the original feature map is added to the weighted feature map.
[0105] In addition, during the model training process, the transformer fault index loss L is introduced tf , in order to increase the attention weight of diagnosis-related features. Specifically, the weighted feature map is first fed into the global maximum pooling layer to obtain a pooled feature map of size (n, 1) Then, according to formulas (2) to (8), To calculate the transformer fault index loss L tf :
[0106] L tf =|avg(HNR)| (8)
[0107] Here, avg(·) represents the average value in each batch, and |.| represents the absolute value.
[0108] By calculating the mean and absolute values of these statistics, the model can more effectively identify fault-related patterns, which is crucial to improving the accuracy and robustness of fault diagnosis.
[0109] Step S3: The feature distribution of the source domain and the target domain is adjusted by using the local maximum mean difference (LMMD) and the domain discriminator, so that the model has generalization ability and domain adaptability. Among them, LMMD is used to measure the data distribution and define the category weight to further align the subdomain distribution, and the domain discriminator is used to generate domain-invariant features until the Nash equilibrium is reached.
[0110] The source domain is the labeled dataset used when training the model, and the target domain is the unlabeled dataset encountered when the model is applied. The two may come from different distributions but share the same label space. represents the source domain, where is the i-th source sample and its corresponding label These samples are drawn from the distribution P s Extract from the middle, n s Indicates the number of source samples. Source samples and |C s |Category associated, where C s represents the label space of the source domain. Let represents the target domain, where is the i-th unlabeled sample in the target domain, which is drawn from the distribution P t Extract from the middle, n t represents the number of target samples. The target samples and |C t |Category associated, where C t Represents the label space of the target domain.
[0111] Assume that the data in the source domain and the target domain are sufficient and obey different marginal probability distributions, that is, P s ≠P t The source domain and the target domain share the same label space, namely, C s =C t ,The number of samples in all categories in the source and target ,domains are balanced.
[0112] Among them, design a function or neural network f * To minimize the difference between marginal distribution and conditional distribution, and learn invariant features across domains, and According to the covariate shift assumption, f * Be consistent across areas.
[0113] LMMD and domain discriminator are used to minimize the conditional and marginal distribution differences of features. LMMD is inspired by the maximum mean difference (MMD), which uses kernel embedding in the reproducing kernel Hilbert space (RKHS) to measure data distribution and is defined as follows:
[0114]
[0115] in, Is the source domain X s The i-th sample in is the target domain X t The i-th sample in n s and n t are the number of source samples and target samples, respectively. is the Gaussian kernel function that maps samples to RKHS.
[0116] LMMD embeds the category weights derived from pseudo labels, further aligns the subdomain distributions, and uses the loss function L LMMD The definition is as follows:
[0117]
[0118] Where C=|C s |=|C t | is the number of categories in the dataset, and They are and The weight corresponding to the category, φ(x) is the Gaussian kernel function that maps the sample to RKHS.
[0119] The domain discriminator is used by a domain adversarial neural network (DANN) to generate domain invariant features, where the feature extractor G makes the domain discriminator D unable to correctly classify the domain label until a Nash equilibrium is reached. The loss function L d The definition is as follows:
[0120]
[0121] Among them, g i is the true domain label of the i-th sample, g i ∈{0,1},D(G(x i ) is the domain label predicted for the i-th sample, D(G(x i ))∈{0,1}, m is the number of samples in the source data and the target data.
[0122] Finally, for the classification loss of the source data It is obtained through the cross entropy function, as shown in formula (12):
[0123]
[0124] in, is the i-th sample label in the source domain, is the kth element in the classifier output, n s is the number of source samples, and 1{·} represents the indicator function.
[0125] By jointly optimizing the three loss functions of LMMD loss, domain discriminator loss, and classification loss, the model can learn invariant features across domains and show good generalization ability and domain adaptability in the target domain.
[0126] Step S4: Optimize model performance by weighing hyperparameters and balance the losses in all aspects to achieve dynamic balance.
[0127] In the proposed adaptive fault attention mechanism residual network training process, the following four optimization objectives are included, corresponding to different aspects of model training.
[0128] Objective 1: Minimize classification error using source domain data This is usually achieved by minimizing the cross entropy loss on the source domain;
[0129] Objective 2: Maximize the inter-domain difference L using source and target domain data d ;
[0130] Objective 3: Minimize the maximum mean difference L between domains using source and target domain data LMMD ;
[0131] Objective 4: Maximize the attention mechanism L using source and target domain data tf .
[0132] The final optimization goal can be expressed by combining the above four optimization goals, as shown in formula (13):
[0133]
[0134] Among them, λ, μ, and γ are trade-off hyperparameters. These hyperparameters are used in the model to balance the mutual influence between different loss function components, ensuring that the model can achieve an optimal compromise in performing such operations as minimizing the classification error of the source domain, maximizing the difference between domains, minimizing the maximum mean difference between domains, and maximizing the benefits brought by the attention mechanism.
[0135] The experimental results show that the RNB-AFAM of the present invention has achieved the highest accuracy in almost all tasks, especially in the D1-D2 task, where RNB-AFAM has obvious advantages over other methods. At the same time, RNB-AFAM has achieved a high accuracy of more than 96% in the D1-D3 task, just like other methods, which shows that RNB-AFAM can not only effectively handle the challenges brought by domain distribution differences, but also surpass other methods in some cases.
[0136] The present invention uses a MEMS photoacoustic sensor with a capacitive resonator to analyze acoustic wave signals through photoacoustic spectroscopy and signal processing technology to collect dissolved gas in oil and form a parameter sequence. An adaptive fault attention mechanism is designed to improve the traditional attention module by adding an adaptive activation function and a fault index function, wherein the adaptive activation function reduces the attention to diagnosis-irrelevant features, and the fault index function enhances the attention weight of diagnosis-related features to generate diagnostic features that are highly correlated with fault features. The adaptive fault attention mechanism residual network model uses the local maximum mean difference and domain discriminator to adjust the source domain and target domain feature distribution, so that the model has generalization ability and domain adaptability. During the training process, the model performance is optimized by weighing hyperparameters, and the losses in various aspects are balanced to achieve dynamic balance. Experimental results show that the model has significant advantages in the similarity, diversity and effectiveness of generated samples. It can be effectively applied to transformer fault diagnosis, greatly improving the accuracy of diagnosis.
Claims
1. A transformer fault diagnosis method based on adaptive fault attention residual network, characterized in that: The following steps are involved: Step 1: Data acquisition: Using a MEMS photoacoustic sensor with a capacitive resonator, photoacoustic spectroscopy and signal processing technology are used to analyze the acoustic wave signal, collect the dissolved gas in the oil and form a parameter sequence; Step 2: Design of adaptive fault attention mechanism: Introduce adaptive activation function and fault index function. The adaptive activation function is used to reduce the attention to diagnosis irrelevant features, and the fault index function is used to enhance the attention weight of diagnosis relevant features to generate diagnostic features that are highly correlated with fault features. Step 3: Feature distribution adjustment: Use LMMD and domain discriminator to adjust the source domain and target domain feature distribution, where the data of the source domain and target domain come from different distributions but share the same label space. LMMD and domain discriminator are used to minimize the marginal distribution difference of features and conditions to achieve cross-domain invariant feature learning; Step 4: Model performance optimization: Optimize model performance by weighing hyperparameters and balancing the losses in various aspects to achieve dynamic balance; The adaptive fault attention mechanism includes two modules, sequence attention and channel attention, which are used to generate sequence attention weights and channel attention weights respectively, and form an improved feature map through adaptive activation function and residual connection; The adaptive activation function in step 2 modifies the activation function by an adaptive threshold to meet the characteristics of shrinkage and maintenance. The definition of the adaptive activation function and its derivative form are as shown in formula (1): Where x is the input value, y is the output value, and τ is the learnable adaptive threshold; The fault index function in step 2 is sensitive to changes in gas concentration and ratio in the transformer fault signal, and introduces the ratio FRG of fault-related gas components to non-fault-related gases as part of the loss function to increase the attention weight to fault-related features.
2. The transformer fault diagnosis method based on adaptive fault attention residual network according to claim 1 is characterized in that: The MEMS photoacoustic sensor in step 1 has a large number of comb-tooth structures to increase the sensitivity of capacitance detection, and reduces gas damping by increasing the gap between the movable structure and the substrate; when the resonator vibrates, the overlapping area between the movable teeth and the fixed teeth changes, and the movement direction of the resonator is perpendicular to the direction of the electric field lines, thereby realizing the separation of acoustic wave excitation and capacitance sensing.
3. The transformer fault diagnosis method based on adaptive fault attention residual network according to claim 1 is characterized in that: The value of FRG is determined by identifying the first local maximum corresponding to the fault-related gas component, as shown in equation (4): Among them, r′ x (τ max ) is r′ x The first local maximum of (τ) reflects the intensity of periodic changes of specific gas components in the fault signal; is the autocorrelation function r x The global maximum value r of (τ) x The normalized form of (0), is the sample mean, x(n) is the sample, n=1,2,…,N; N is the number of sample points.
4. The transformer fault diagnosis method based on adaptive fault attention residual network according to claim 1 is characterized in that: In step 3 represents the source domain, where is the i-th source sample and its corresponding label These samples are drawn from the distribution P s Extract from the middle, n s Indicates the number of source samples; source samples and |C s |Category associated, where C s Represents the label space of the source domain; represents the target domain, where is the i-th unlabeled sample in the target domain, which is drawn from the distribution P t Extract from the middle, n t represents the number of target samples; the target samples and |C t |Category associated, where C t Represents the label space of the target domain; Assume that the data in the source domain and the target domain are sufficient and obey different marginal probability distributions, that is, P s ≠P t ; The source domain and the target domain share the same label space, namely C s =C t , the number of samples in all categories in the source domain and the target domain is balanced; among them, design a function or neural network f * To minimize the difference between marginal distribution and conditional distribution, and learn invariant features across domains, and According to the covariate shift assumption, f * Maintain consistency in different domains; use LMMD and domain discriminator to minimize the conditional and marginal distribution differences of features. LMMD is inspired by the maximum mean difference MMD. MMD uses kernel embedding in the reproducing kernel Hilbert space RKHS to measure data distribution and is defined as follows: in The source domain D s The i-th sample in The target domain D t The i-th sample in n s and n t are the number of source samples and target samples, respectively. It is the Gaussian kernel function that maps samples to RKHS; LMMD embeds the class weights derived from pseudo labels, aligns the subdomain distributions, and uses the loss function L LMMD The definition is as follows: Where C=|C s |=|C t | is the number of categories in the dataset, and They are and The weight of the corresponding category, φ(x) is the Gaussian kernel function that maps the sample to the RKHS; The domain discriminator was originally used by the Domain Adversarial Neural Network (DANN) to generate domain-invariant features, where the feature extractor G makes the domain discriminator D unable to correctly classify the domain label until a Nash equilibrium is reached, and the loss function L d The definition is as follows: where g i is the true domain label of the i-th sample, g i ∈{0,1},D(G(x i ) is the domain label predicted for the i-th sample, D(G(x i ))∈{0,1}, m is the number of samples in the source data and the target data; Finally, for the classification loss of the source data It is obtained through the cross entropy function, as shown in formula (12): in is the i-th sample label in the source domain, is the kth element in the classifier output, n s is the number of source samples, and 1{·} represents the indicator function.
5. The transformer fault diagnosis method based on adaptive fault attention residual network according to claim 4 is characterized in that: In the training process, the step 4 combines the following four optimization goals: Goal 1: Minimize the classification error using source domain data; Goal 2: Maximize the loss of the domain discriminator or minimize the distinguishability between domains using source and target domain data; Goal 3: Minimize LMMD using source and target domain data; Goal 4: Maximize the benefits of the attention mechanism using source and target domain data; By combining the above four optimization objectives, as shown in formula (13): Among them, λ, μ and γ are trade-off hyperparameters, L tf Index losses for transformer faults; To ensure that the model reaches the optimal compromise state in terms of minimizing the source domain classification error, maximizing the difference between domains, minimizing the maximum mean difference between domains, and maximizing the benefits of the attention mechanism.