A Data Augmentation Method and System for Oil Discharge Impact Signals Based on Hybrid Time-Frequency Convolution

By using the VAE-GAN method with hybrid time-frequency convolution, the time-frequency features of discharge impact signals in transformer oil are extracted and generated, which solves the problem of insufficient small sample datasets, generates realistic and diverse data samples, and improves the robustness and recognition performance of the fault diagnosis model.

CN121479286BActive Publication Date: 2026-05-05YIBIN POWER SUPPLY COMPANY STATE GRID SICHUAN ELECTRIC POWER
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YIBIN POWER SUPPLY COMPANY STATE GRID SICHUAN ELECTRIC POWER
Filing Date
2026-01-09
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

The small sample size of discharge impact signals in transformer oil limits the development of deep learning-based fault diagnosis and early warning models. Existing data augmentation methods cannot effectively simulate time-frequency characteristic changes, and the generated data lacks authenticity and diversity, making it impossible to build highly robust diagnostic models.

Method used

A variational autoencoder generative adversarial network (VAE-GAN) based on hybrid time-frequency convolution is adopted. The time-domain and frequency-domain features of the time-frequency map are extracted and generated through the hybrid time-frequency convolution structure. The data augmentation is performed by combining the generative adversarial network to ensure the time-frequency authenticity of the samples and the stability of the model.

Benefits of technology

It effectively solves the problem of small sample and imbalanced datasets for discharge impact signals in oil. The generated samples have real time-frequency characteristics, which improves the diagnostic performance of deep learning models and makes them suitable for various application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121479286B_ABST
    Figure CN121479286B_ABST
Patent Text Reader

Abstract

This invention belongs to the field of high-voltage equipment monitoring and processing technology, and discloses a method and system for data enhancement of oil discharge impact signals based on hybrid time-frequency convolution. The method performs two-dimensional processing on the acquired oil discharge impact signals to obtain a two-dimensional time-frequency map; the two-dimensional time-frequency map is used to train a generative adversarial network (GAN); the GAN includes an encoder, a decoder, and a discriminator; the trained decoder is used as a generator to obtain a generated time-frequency map, thereby achieving data enhancement of the oil discharge impact signal. This invention significantly improves the fidelity of the time-frequency features of the generated samples while enhancing the oil discharge impact signal data, and can also achieve fault identification of discharge types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of high-voltage equipment monitoring and processing technology, and relates to data enhancement of discharge impulse signals in transformer oil, and particularly to a method and system for data enhancement of discharge impulse signals in oil based on hybrid time-frequency convolution. Background Technology

[0002] Oil-immersed transformers, as key equipment in modern power systems, undertake the important functions of voltage conversion and power transmission. Their safety and stability directly determine the overall operational reliability of the power system. With the continuous growth of electricity demand, the load on transformers is increasing daily. The strong coupling factors such as high voltage, strong electric field, high temperature, and strong vibration within the transformer are having a more significant impact on its insulation performance, leading to frequent internal arcing faults. When an arcing fault occurs, the rapid injection of external energy causes the arcing channel to expand rapidly and the temperature to rise sharply. The huge pressure difference formed in a short time causes a rapid increase in the internal pressure level of the oil tank. If the local stress exceeds the strength limit of the tank material, the tank will rupture, and the tank structure will be damaged.

[0003] Transformer oil discharge accidents are highly random and instantaneous, making it difficult to collect sufficient and complete samples of oil discharge impact signals for different fault causes and operating conditions. This "small sample" problem severely restricts the development of deep learning-based fault diagnosis and early warning models. Such models typically rely on massive amounts of high-quality data for sufficient training, to avoid overfitting, and to ensure their generalization ability.

[0004] Currently, research methods for addressing the small sample size problem of discharge impact signals in oil mainly include traditional data augmentation and generative models. Traditional methods, such as adding noise and time-series distortion, are simple and easy to implement, but the generated signals often only undergo simple transformations at the time domain level, making it difficult to effectively simulate the rich time-frequency feature changes in the discharge physical process. This results in insufficient data diversity and limited improvement in model performance. On the other hand, methods based on deep generative models such as generative adversarial networks use simple convolutional patterns that cannot independently extract or generate features based on the inherent differences between time-domain and frequency-domain features. Furthermore, their training process is inherently unstable, and under conditions of extremely limited data, they are prone to mode collapse or overfitting. The authenticity and diversity of the generated data cannot be guaranteed, and they cannot fully explore the deep fault feature information contained in small samples.

[0005] Furthermore, existing data augmentation methods fail to fully simulate the inherent patterns of the time-frequency distribution of the impact signal during arc discharge; they also fail to effectively and specifically extract the differentiated features of the discharge impact signal in the time and frequency domains. These factors will lead to discrepancies between the augmented data and the physical mechanism of the real discharge signal, making it difficult to use them to construct highly robust diagnostic models. Summary of the Invention

[0006] To address the problem that it is difficult to collect sufficient and complete signal features of arc fault discharge impulse signals in transformer oil under different typical faults and operating conditions, which fails to meet the demand for massive high-quality data in deep learning research on fault identification, this invention aims to propose a data enhancement method and system for discharge impulse signals in oil based on hybrid time-frequency convolution. This method takes into account the independent time-domain and frequency-domain features of the discharge impulse signal, and simultaneously reconstructs the time-domain and frequency-domain features of the time-frequency map during the data generation process through hybrid time-frequency convolution. Then, data enhancement is performed through generative adversarial networks.

[0007] The technical approach of this invention is as follows: Using a Variational Autoencoder Generative Adversarial Network (VAE-GAN) as the basic model, a hybrid time-frequency convolutional structure is introduced to extract and generate distinct time-domain and frequency-domain features from the time-frequency map. The hybrid time-frequency convolutional structure first preprocesses the time-frequency spectrum, using whitespace padding and interpolation padding in the time and frequency domains respectively. Subsequently, dedicated convolutional kernels are constructed and applied along the time and frequency domains for feature extraction. The extracted feature maps are then downsampled by pooling and output. This invention couples the hybrid time-frequency convolutional structure into the variational autoencoder, integrating its latent space distribution constraints with the adversarial training mechanism of the generative adversarial network to ensure the time-frequency authenticity of the samples and the stability of the model.

[0008] The present invention provides a method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution, comprising the following steps:

[0009] S1, The collected oil discharge impact signal is processed into two dimensions to obtain a two-dimensional time-frequency diagram;

[0010] S2, a generative adversarial network is trained using a two-dimensional time-frequency map; the generative adversarial network includes an encoder, a decoder, and a discriminator; the encoder takes the two-dimensional time-frequency map as input to obtain the corresponding latent spatial distribution; the decoder takes the output of the encoder and noise as input to obtain a reconstructed time-frequency map and a generated time-frequency map; the discriminator is used to determine the authenticity of the input image; both the encoder and the decoder are based on hybrid time-frequency convolution.

[0011] S3 utilizes the trained decoder as a generator to obtain the generated time-frequency map, thereby achieving data enhancement of the discharge impact signal in oil.

[0012] In step S1 above, the collected oil discharge impact signal is processed into two dimensions using short-time Fourier transform.

[0013] The continuous Fourier transform is defined as:

[0014] ;

[0015] The Discrete Fourier Transform (DFT) is defined as follows:

[0016] ;

[0017] In the formula, and These are continuous and discrete time-domain signals, respectively. and For window functions (usually Hamming or Hanning windows are chosen); and This is the transformed time-frequency representation, i.e., a two-dimensional time-frequency plot; f and k represent frequencies; N represents the number of discrete points in the signal. The time axis of the two-dimensional time-frequency plot (horizontally). (or ) and frequency axis (along the longitudinal direction) (or These respectively characterize the local time dynamics and instantaneous spectral characteristics of the signal.

[0018] In the preferred implementation, the collected oil discharge impact signal is processed into two dimensions using a fast Fourier transform (FFT).

[0019] Set an appropriate sampling frequency based on the frequency characteristics of the signal. Sampling signal length Window function length Number of overlapping points , FFT points After unfolding, a square shape for the spectrum is preferable, as it facilitates setting convolution parameters. A two-dimensional time-frequency graph X can be represented using a two-dimensional matrix. It means that, among them For the time dimension (number of time frames). The frequency dimension (the number of frequency points) is determined by the following formula:

[0020] ;

[0021] According to the Nyquist sampling theorem, the effective frequency components that can be identified are in the range [0, 1]. The larger the number of FFT frequency points, the denser the frequency axis, and the stronger the frequency domain resolution.

[0022] .

[0023] Window length The smaller the value, the denser the number of frames on the timeline, and the stronger the ability to capture instantaneous details.

[0024] According to the uncertainty principle, it is impossible to achieve infinite precision in both the time and spatial dimensions simultaneously. Therefore, finding parameter settings that balance time precision and frequency domain precision is crucial, and the number of overlap points is generally set accordingly. Window function length Half of the signal length is used, and other parameters are set according to the actual sampling frequency and signal length to ensure that the dimensions of the signal after two-dimensional unfolding are easy for subsequent calculations. Generally, it is set to half of the signal length. or Size: If the dimension of the two-dimensional spectrum sample is too large, it will affect subsequent operations.

[0025] In step S2 above, the Generative Adversarial Network (GAN) is a Variational Autoencoder (VAE) based on hybrid time-frequency convolution, which includes an encoder, a decoder, and a discriminator.

[0026] In one possible implementation, the encoder is composed of a hybrid time-frequency convolution module or two or more hybrid time-frequency convolution modules cascaded together; the directional features along the frequency axis and the directional features along the time axis in the time-frequency graph are extracted by heterogeneous convolution kernels respectively.

[0027] The hybrid time-frequency convolution module includes a first padding unit, a convolution unit, a first fusion unit, and a downsampling unit. The first padding unit performs time-frequency axis heterogeneous padding on the two-dimensional time-frequency graph using corresponding padding strategies along the time axis and frequency axis, respectively. The convolution unit includes a set of parallel heterogeneous convolution kernels, which perform convolution processing on the heterogeneously padded two-dimensional time-frequency graph using horizontal and vertical convolution kernels of different sizes, respectively. The first fusion unit performs weighted fusion of the output features of each convolution kernel to obtain fused time-frequency features. The downsampling unit is used to perform downsampling processing on the fused time-frequency features.

[0028] The first filling unit is based on heterogeneous filling along the time-frequency axis; for a two-dimensional time-frequency graph, different filling strategies are adopted along the time axis and the frequency axis to adapt to the inherent physical differences between time and frequency domain features.

[0029] In the specific implementation method:

[0030] Time axis extension: Zero values ​​are used to fill along the time axis of the two-dimensional time-frequency plot;

[0031] Frequency axis extension: Interpolation is used to fill along the frequency axis of the two-dimensional time-frequency plot.

[0032] By expanding the time axis as described above, zero-valued pixels are filled on both sides of the two-dimensional time-frequency map. This aims to expand the receptive field of the time dimension while avoiding the introduction of additional boundary noise, ensuring that the model focuses on learning dynamic features within the effective time interval. Similarly, by expanding the frequency axis as described above, new data points are inserted between the original frequency points using linear or spline interpolation algorithms. This aims to enhance the continuity and smoothness of the frequency distribution, thereby more naturally depicting the frequency domain energy distribution characteristics of the discharge signal. After this heterogeneous filling process, the original two-dimensional time-frequency map is expanded into an enhanced two-dimensional time-frequency map, which serves as the input time-frequency feature map for the subsequent generative adversarial network, laying the foundation for subsequent multi-scale feature extraction.

[0033] In the convolutional unit, the vertical convolutional kernel (with a height dimension larger than the width dimension, typically a 3×1 or 5×3 kernel) is used to scan the input time-frequency feature map (X) along the frequency axis to extract global spectral structure features; the horizontal convolutional kernel (with a width dimension larger than the height dimension, typically a 1×3 or 3×5 kernel) is used to scan in the local time-frequency region to capture the joint information of transient temporal features and global temporal patterns. The horizontal and vertical convolutional kernels can be one-dimensional or two-dimensional.

[0034] The first fusion unit performs weighted summation of the time-frequency feature maps output by each convolutional kernel using learnable weight coefficients, achieving adaptive fusion of multi-scale features. This coupling process ensures that temporal dynamics and frequency-domain structural performance are synergistically optimized and uniformly represented.

[0035] The downsampling unit includes a two-dimensional convolutional layer, a normalization layer, and an activation function. The two-dimensional convolutional layer convolves the fused time-frequency features and then processes them through the normalization layer and an activation function (such as ReLU) to output the enhanced deep time-frequency features, which are the latent spatial distribution of the two-dimensional time-frequency map.

[0036] The encoder's final output is the mean of the latent spatial variable distribution of the two-dimensional time-frequency plot. With log variance This maps the input time-frequency graph to a continuous and smooth latent distribution. .

[0037] ;

[0038] in, Indicates encoder; Represents a two-dimensional time-frequency graph;

[0039] From the potential distribution The sampled seed noise z is used as the input to the decoder, and the noise satisfies:

[0040] ;

[0041] Where I represents a unit diagonal matrix;

[0042] In this invention, the noise input to the decoder includes:

[0043] ;

[0044] ;

[0045] Among them, noise Used to obtain the reconstructed time-frequency map through the decoder; noise Used to obtain the generated time-frequency diagram through the decoder; Indicates standard deviation; ∠ represents standard normal distribution noise, and ⊙ represents element-wise multiplication.

[0046] In one implementation, the decoder receives noise sampled from the latent variable distribution using a reparameterization technique as input. The decoder is symmetrically positioned relative to the encoder and consists of one or more cascaded hybrid time-frequency deconvolution modules. The hybrid time-frequency deconvolution module includes a second padding unit, a deconvolution unit, a second fusion unit, and an upsampling unit. The second padding unit performs time-frequency heterogeneous padding on the noise input to the decoder using corresponding padding strategies along the time and frequency axes, respectively. The deconvolution unit includes a set of parallel heterogeneous deconvolution kernels (i.e., transposes of convolution kernels), reconstructing the heterogeneously padded noise using horizontal and vertical deconvolution kernels of different sizes. The second fusion unit performs weighted fusion of the output features of each deconvolution kernel to obtain fused reconstructed features. The upsampling unit is used to upsample the fused reconstructed features. The second padding unit uses the same padding strategy as the first padding unit. The horizontal and vertical deconvolution kernels can be one-dimensional or two-dimensional deconvolution kernels. The upsampling unit includes a two-dimensional deconvolution layer, a normalization layer, and an activation function. The two-dimensional deconvolution layer deconvolves the fused reconstructed features and then processes them through the normalization layer and an activation function (such as ReLU) to output the enhanced reconstructed features.

[0047] In this invention, input noise is addressed. and The time-frequency diagram is reconstructed through the decoder output. With the generation of time-frequency diagrams :

[0048] ;

[0049] ;

[0050] in, This indicates the decoder.

[0051] In one implementation, the discriminator includes a feature extraction module and a discriminant output module, with the goal of accurately distinguishing between genuine and fake input samples. The input sample for the feature extraction module is a real two-dimensional time-frequency graph. Reconstructing the time-frequency diagram and generating time-frequency diagrams The feature extraction module employs four sequentially arranged third convolutional units, each containing a convolutional layer, a normalization layer, and an activation function (such as the LeakyReLU activation function). The discriminant output module includes a first fully connected layer, used to output a score representing the authenticity of the input sample. The tensor output by the feature extraction module is flattened and then fed into the first fully connected layer.

[0052] Furthermore, the discriminator also predicts the type of the input samples. The discriminant output module also includes a second fully connected layer for outputting the predicted classification type. The tensor output by the feature extraction module is flattened and then simultaneously fed into the second fully connected layer.

[0053] This invention aims to collaboratively optimize the distribution matching capability of VAEs and the generation quality of GANs by setting a multi-objective joint loss function. Including variational lower bound loss Combating losses Auxiliary classification loss Three parts:

[0054] ;

[0055] ;

[0056] ;

[0057] ;

[0058] in, Represents the mathematical expectation of a variable; express KL The weighting coefficient of the divergence constraint term; This indicates the calculation of the KL divergence function; Denotes the prior distribution of the latent variable z; min G This indicates that the solution is minimized using the decoder's parameters as optimization variables; max D This means that the solution is maximized by optimizing the parameters of the discriminator. This describes the process by which the discriminator obtains the true or false scores of the input samples; This indicates the process by which the discriminator obtains the predicted classification. Indicates the predicted classification type. Indicates the true type label of the input sample; for or ; Indicates the distribution of real data The expected value of the samples obtained from sampling and their corresponding labels is calculated. This indicates the prior distribution of the latent variables. The mathematical expectation calculated from the latent variables sampled in the data and the type labels sampled from the real data distribution; , , This is a hyperparameter used to balance the proportion of each loss.

[0059] The core of the training method employed in this invention lies in a two-stage training strategy. First, the generative adversarial network (GAN) is globally trained using the original imbalanced sample set, enabling the decoder to learn the time-frequency distribution characteristics of discharge signals for different fault types, ensuring that the generated samples have the same time-frequency characteristics as the real samples. The trained decoder is then used as a generator to generate a sufficient number of time-frequency map samples for categories with scarce samples, constructing a balanced and adequate new dataset together with the real samples. Next, the discriminator undergoes fine-tuning: the encoder and decoder parameters are frozen, and the discriminator's classification function is trained solely using this balanced dataset, with the optimization objective being solely the auxiliary classifier loss. This strengthens the discriminator's auxiliary classification function, enabling a typical fault classifier that achieves optimal recognition performance in actual diagnosis.

[0060] The present invention also provides a system for implementing the above-mentioned method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution, comprising:

[0061] The preprocessing module is used to perform two-dimensional processing on the collected oil discharge impact signal to obtain a two-dimensional time-frequency diagram;

[0062] Generative Adversarial Networks (GANs) include an encoder, a decoder, and a discriminator. The encoder takes a two-dimensional time-frequency image as input to obtain the corresponding latent spatial distribution. The decoder takes the output of the encoder and noise as input to obtain a reconstructed time-frequency image and generate a time-frequency image. The discriminator is used to determine the authenticity of the input image. Both the encoder and the decoder are based on hybrid time-frequency convolution.

[0063] The preprocessing module described above performs the operation according to step S1 given above; the specific structures of the encoder, decoder and discriminator in the generated adversarial network are as described above.

[0064] This invention proposes a data enhancement method and system for oil discharge impact signals based on hybrid time-frequency convolution. Within a neural network framework, it introduces a hybrid time-frequency convolution structure based on a traditional generative adversarial network. During data generation, it simultaneously and independently reconstructs the time-domain and frequency-domain features of the time-frequency map, conforming to the time-frequency patterns of real samples. This method effectively addresses the challenges of small sample sizes and imbalanced datasets for oil discharge impact signals. By generating time-frequency map samples of oil discharge impact signals with realistic time-frequency features, it compensates for the deficiencies in class balance and data scale of the original dataset, providing a reliable data foundation for deep learning-driven intelligent diagnosis of high-voltage equipment. Furthermore, the network can also achieve fault identification functionality after transfer training, adapting to various application scenarios. Compared with existing technologies, this invention has the following advantages:

[0065] 1. Traditional convolution operations use the same sampling mode in both the time and frequency domains, which makes it difficult to adapt to the differences in characteristics between discrete distributions in the time domain and continuous distributions in the frequency domain. The encoder and decoder of the generative adversarial network proposed in this invention are based on hybrid time-frequency convolution. By designing zero-value padding along the time axis and interpolation padding along the frequency axis, combined with a heterogeneous convolution kernel structure, the decoupled extraction of time-frequency features is achieved, which significantly improves the fidelity of time-frequency features of the generated samples.

[0066] 2. This invention uses short-time Fourier transform to process the collected oil discharge impact signal into two dimensions, overcoming the inherent defects of traditional time-frequency analysis methods. Compared with the invalid region loss of pixels in wavelet transform and the irreversibility of Hilbert transform, it greatly preserves the transient characteristics and frequency domain structure characteristics of the signal, while ensuring the reversibility of the sample. Through reasonable parameter settings, the sample dimension can be effectively controlled, greatly improving the computational efficiency of the neural network.

[0067] 3. The generative adversarial network proposed in this invention combines the distributed learning capability of variational autoencoders with the high-quality generation advantage of generative adversarial networks. The random sampling of its latent space ensures the diversity of generated samples, while the adversarial training mechanism guarantees the visual realism of generated samples, and at the same time solves the problem of poor stability of adversarial networks.

[0068] 4. The discriminator network of this invention has the characteristics of multi-functional integration. The discriminator can serve as a true / false detector to provide adversarial loss discrimination for the generator, and can also serve as an auxiliary classifier to introduce sample category information. Especially in the deployment stage, the fully trained discriminator can be directly used as a high-performance fault classifier, reducing training costs and achieving multi-functional use. Attached Figure Description

[0069] Figure 1 This is a flowchart illustrating a data enhancement method for discharge impact signals in oil based on hybrid time-frequency convolution.

[0070] Figure 2 This is a schematic diagram of the time-frequency axis heterogeneous filling principle;

[0071] Figure 3 This is a schematic diagram illustrating the principle of time-frequency feature extraction and fusion;

[0072] Figure 4 This is a schematic diagram of the operation process of the verification experiment in Example 1;

[0073] Figure 5 It is a two-dimensional time-frequency diagram of the discharge impact signal in oil; among them, (a), (b), and (c) are the original two-dimensional time-frequency diagrams, and (d), (e), and (f) are the generated two-dimensional time-frequency diagrams. Detailed Implementation

[0074] The technical solutions of various embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0075] Example 1

[0076] This embodiment provides a method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution, such as... Figure 1 As shown, it includes the following steps:

[0077] S1. The collected oil discharge impact signal is processed into two dimensions to obtain a two-dimensional time-frequency diagram.

[0078] This step utilizes the short-time Fourier transform to perform two-dimensional processing on the acquired oil discharge impact signal. The short-time Fourier transform can be either a continuous Fourier transform or a discrete Fourier transform.

[0079] The continuous Fourier transform is defined as:

[0080] ;

[0081] The Discrete Fourier Transform (DFT) is defined as follows:

[0082] ;

[0083] In the formula, and These are continuous and discrete time-domain signals, respectively. and For window functions (usually Hamming or Hanning windows are chosen); and This is the transformed time-frequency representation, i.e., a two-dimensional time-frequency plot; f and k represent frequencies; N represents the number of discrete points in the signal. The time axis of the two-dimensional time-frequency plot (horizontally). (or ) and frequency axis (along the longitudinal direction) (or These respectively characterize the local time dynamics and instantaneous spectral characteristics of the signal.

[0084] In this embodiment, the Fast Fourier Transform (FFT) is used to perform two-dimensional processing on the acquired oil discharge impact signal.

[0085] Set an appropriate sampling frequency based on the frequency characteristics of the signal. Sampling signal length Window function length Number of overlapping points , FFT points When unfolded, the spectrum has a square shape, which is beneficial for setting convolution parameters. A two-dimensional time-frequency graph X can be represented using a two-dimensional matrix. It means that, among them For the time dimension (number of time frames). For frequency dimension (number of frequency points).

[0086] The dimension parameter is determined by the following formula:

[0087] .

[0088] According to the Nyquist sampling theorem, the effective frequency components that can be identified are in the range [0, 1]. The larger the number of FFT frequency points, the denser the frequency axis, and the stronger the frequency domain resolution.

[0089] .

[0090] Window length The smaller the value, the denser the number of frames on the timeline, and the stronger the ability to capture instantaneous details.

[0091] According to the uncertainty principle, it is impossible to achieve infinite precision in both the time and spatial dimensions simultaneously. Therefore, finding parameter settings that balance time precision and frequency domain precision is crucial, and the number of overlap points is generally set accordingly. Window function length Half of the signal length is used, and other parameters are set according to the actual sampling frequency and signal length to ensure that the dimensions of the signal after two-dimensional unfolding are easy for subsequent calculations. Generally, it is set to half of the signal length. or Size: If the dimension of the two-dimensional spectrum sample is too large, it will affect subsequent operations.

[0092] S2 uses a two-dimensional time-frequency graph to train the generative adversarial network.

[0093] Generative Adversarial Networks (GANs) are variational autoencoders (VAEs) based on hybrid time-frequency convolutions, comprising an encoder, a decoder, and a discriminator. The encoder takes a two-dimensional time-frequency image as input to obtain the corresponding latent spatial distribution; the decoder takes the encoder's output and noise as input to obtain a reconstructed time-frequency image and a generated time-frequency image; the discriminator is used to determine the authenticity of the input image. Both the encoder and decoder are based on hybrid time-frequency convolutions. During training, the input to the GAN is the two-dimensional time-frequency image X obtained in step S1 and the corresponding class label AC.

[0094] (1) The encoder is composed of two cascaded hybrid time-frequency convolution modules; the directional features along the frequency axis and the directional features along the time axis in the time-frequency map are extracted by heterogeneous convolution kernels respectively.

[0095] The hybrid time-frequency convolution module includes a first padding unit, a convolution unit, a first fusion unit, and a downsampling unit.

[0096] The first filling unit is based on heterogeneous filling along the time-frequency axis. For the two-dimensional time-frequency plot, different filling strategies are adopted along the time axis and the frequency axis to adapt to the inherent physical differences between the time and frequency domain characteristics. The first filling unit performs heterogeneous filling of the two-dimensional time-frequency plot along the time axis and the frequency axis respectively using corresponding filling strategies.

[0097] In specific implementation methods, such as Figure 2 As shown, the first filling unit:

[0098] Time axis extension: Zero values ​​are used to fill the time axis direction of the two-dimensional time-frequency plot;

[0099] Frequency axis extension: Interpolation is used to fill the frequency axis direction in the two-dimensional time-frequency plot.

[0100] The aforementioned time-axis expansion fills the two-dimensional time-frequency map with zero-valued pixels on both sides, aiming to expand the receptive field of the time dimension while avoiding the introduction of additional boundary noise, ensuring that the model focuses on learning dynamic features within the effective time interval. Similarly, the aforementioned frequency-axis expansion uses a linear interpolation algorithm to insert new data points between existing frequency points, aiming to enhance the continuity and smoothness of the frequency distribution, thereby more naturally depicting the frequency domain energy distribution characteristics of the discharge signal. After this heterogeneous filling process, the original two-dimensional time-frequency map is expanded into an enhanced two-dimensional time-frequency map, which serves as the input time-frequency feature map for the subsequent generative adversarial network, laying the foundation for subsequent multi-scale feature extraction.

[0101] The convolutional unit includes a set of parallel heterogeneous convolutional kernels, which perform convolution processing on the heterogeneously filled two-dimensional time-frequency map using horizontal and vertical convolutional kernels of different sizes. In this embodiment, both the horizontal and vertical convolutional kernels are one-dimensional convolutional kernels. Figure 3 As shown, the vertical convolution kernel (the height dimension is larger than the width dimension, such as a 2*1 convolution kernel) is used to scan the input time-frequency feature map (X) along the frequency axis to extract global spectral structure features; the horizontal convolution kernel (the width dimension is larger than the height dimension, such as a 1*2 or 1*4 convolution kernel) is used to scan in the local time-frequency region to capture the joint information of transient time-domain features and global time-domain patterns.

[0102] like Figure 3 As shown, the first fusion unit performs weighted fusion of the output features of each convolutional kernel to obtain fused time-frequency features. The output feature maps are weighted and summed using the learnable weight coefficients of each convolutional kernel (0.25, 0.25, 0.5 as shown in the figure) to achieve adaptive fusion of multi-scale features.

[0103] The downsampling unit is used to downsample the fused time-frequency features. The downsampling unit includes a two-dimensional convolutional layer, a normalization layer, and an activation function. The two-dimensional convolutional layer (in this embodiment, the two-dimensional convolutional layer is set to a 3*3 convolutional kernel) convolves the fused time-frequency features and then processes them through the normalization layer and the activation function (such as ReLU) to output the enhanced deep time-frequency features, which is the latent spatial distribution of the two-dimensional time-frequency map.

[0104] The encoder's final output is the mean of the latent spatial variable distribution of the two-dimensional time-frequency plot. With log variance This maps the input time-frequency graph to a continuous and smooth latent distribution. .

[0105] ;

[0106] in, Indicates encoder; This represents a two-dimensional time-frequency diagram.

[0107] From the potential distribution The sampled seed noise z is used as the input to the decoder, and the noise satisfies:

[0108] ;

[0109] Where I represents a unit diagonal matrix.

[0110] In this invention, the noise input to the decoder includes:

[0111] ;

[0112] ;

[0113] Among them, noise Used to obtain the reconstructed time-frequency map through the decoder; noise Used to obtain the generated time-frequency diagram through the decoder; Indicates standard deviation; ∠ represents standard normal distribution noise, and ⊙ represents element-wise multiplication.

[0114] (2) The decoder receives noise sampled from the latent variable distribution through a reparameterization technique as input.

[0115] The decoder is symmetrically positioned relative to the encoder and consists of two cascaded hybrid time-frequency deconvolution modules.

[0116] The hybrid time-frequency deconvolution module includes a second padding unit, a deconvolution unit, a second fusion unit, and an upsampling unit.

[0117] The second padding unit employs corresponding padding strategies along both the time and frequency axes to perform time-frequency heterogeneous padding of the noise input to the decoder. The padding strategy of the second padding unit is the same as that of the first padding unit.

[0118] The deconvolution unit comprises a set of parallel heterogeneous deconvolution kernels (i.e., transposes of convolution kernels), which reconstruct the noise after heterogeneous filling using horizontal and vertical deconvolution kernels of different sizes. Both the horizontal and vertical deconvolution kernels are one-dimensional deconvolution kernels, which are transposes of the aforementioned horizontal and vertical convolution kernels.

[0119] The second fusion unit performs weighted fusion of the output features of each deconvolution kernel to obtain fused reconstructed features.

[0120] The upsampling unit is used to upsample the fused reconstructed features. The upsampling unit includes a two-dimensional deconvolution layer, a normalization layer, and an activation function. The two-dimensional deconvolution layer (which is the transpose of the previous two-dimensional convolution layer) deconvolves the fused reconstructed features, which are then processed by the normalization layer and an activation function (such as ReLU) to output the enhanced reconstructed features.

[0121] In this embodiment, input noise is addressed. and The time-frequency diagram is reconstructed through the decoder output. With the generation of time-frequency diagrams :

[0122] ;

[0123] ;

[0124] in, This indicates the decoder.

[0125] (3) The discriminator includes a feature extraction module and a discriminant output module. Its goal is to accurately distinguish between the authenticity of the input samples and to predict the type of the input samples.

[0126] The input sample for the feature extraction module is a real two-dimensional time-frequency graph. Reconstructing the time-frequency diagram and generating time-frequency diagrams The feature extraction module uses four sequentially arranged third convolutional units. Each third convolutional unit contains a convolutional layer (in this embodiment, the convolutional layer is set to a 4*4 convolutional kernel), a normalization layer, and an activation function (such as the LeakyReLU activation function).

[0127] The discriminant output module includes a first fully connected layer and a second fully connected layer. The first fully connected layer outputs a score representing the authenticity of the input sample, and the second fully connected layer outputs the predicted classification type. The tensors output by the feature extraction module are flattened and then fed into the first fully connected layer and the second fully connected layer, respectively.

[0128] This embodiment sets up a multi-objective joint loss function to collaboratively optimize the distribution matching capability of VAE and the generation quality of GAN. Including variational lower bound loss Combating losses Auxiliary classification loss Three parts:

[0129] ;

[0130] ;

[0131] ;

[0132] ;

[0133] in, Represents the mathematical expectation of a variable; express KL The weighting coefficient of the divergence constraint term; This indicates the calculation of the KL divergence function; Denotes the prior distribution of the latent variable z; min G This indicates that the solution is minimized using the decoder's parameters as optimization variables; max D This means that the solution is maximized by optimizing the parameters of the discriminator. This describes the process by which the discriminator obtains the true or false scores of the input samples; This indicates the process by which the discriminator obtains the predicted classification. Indicates the predicted classification type. Indicates the true type label of the input sample; for or ; Indicates the distribution of real data The expected value of the samples obtained from sampling and their corresponding labels is calculated. This indicates the prior distribution of the latent variables. The mathematical expectation calculated from the latent variables sampled in the data and the type labels sampled from the real data distribution; , , This is a hyperparameter used to balance the proportion of each loss.

[0134] S3 utilizes the trained decoder as a generator to obtain the generated time-frequency map, thereby achieving data enhancement of the discharge impact signal in oil.

[0135] This embodiment uses discharge impact signals from typical arc discharge faults in transformer oil (four fault models: tip discharge, surface discharge, floating discharge, and inter-turn discharge) to verify the proposed data enhancement method for discharge impact signals in oil based on hybrid time-frequency convolution.

[0136] like Figure 4 As shown, according to the set fault model conditions, the AC uniform voltage boost method is used to break down the sample. The current transformer is used as the discharge impact sampling trigger signal for acquisition. Based on experience, the sampling frequency is set to 5MHz as the optimal sampling frequency. The discharge impact signal in the oil is captured by a piezoelectric pressure sensor placed in the transformer oil. The signal information is fully characterized using the minimum amount of data, forming an unbalanced impact signal sample dataset.

[0137] First, all samples in the discharge impulse signal sample dataset were preprocessed. The duration of the first wave of the discharge impulse signal is on the order of hundreds of seconds, corresponding to a data length of 520 points. The first wave segment of each signal was extracted as a signal sample segment. Training and test sets were then formed in a 6:4 ratio. The imbalanced data samples are shown in Table 1 below.

[0138] Table 1. Number of Samples in the Training and Testing Sets of Discharge Impact Signals

[0139]

[0140] The fault dataset is significantly imbalanced, and experience shows that imbalanced data can greatly impact the performance of the identification model. The sample signal, with a length of 520 bits and a sampling frequency of 5MHz, is processed using Fast Fourier Transform (FFT) according to step S1 to obtain a two-dimensional time-frequency graph dataset. Based on empirical parameters, the window function length is set to 8, and the number of FFT points is... The number of overlapping points is 126. The value is 4, corresponding to the sample dimension obtained after the Fourier transform of the sample. .

[0141] Following step S2 as described above, the training set with an imbalanced number of samples is input into the generative adversarial network for training. After each training round, the multi-objective joint loss function given earlier is applied. ( =1, =0.5, =0.1), and use the Adam optimizer (learning rate 0.001) to optimize the parameters of the generative adversarial network; repeat the above training process until the set number of training cycles (100 training cycles) is reached.

[0142] After training is complete, determine the total loss value. Whether it is less than 0.5 is used to test the quality of the generated sample; if the total loss value is... If the value is less than 0.5, the quality meets the standard. Then, proceed to step S3, based on the value determined by the encoder. and , construct noise The pre-trained decoder is used as a generator to generate different data samples based on different data types, enriching the training set with 500 samples of each fault type. If the quality of the generated samples is not up to standard, the generative adversarial network is improved by adjusting hyperparameters, the number of hybrid time-frequency convolutional modules, and the number of hybrid time-frequency deconvolutional modules. The training process is then repeated until the quality of the generated samples meets the standard.

[0143] Figure 5 The comparison results between the two-dimensional time-frequency plot generated by the generator trained by the method of this embodiment and the original two-dimensional time-frequency plot are given. As can be seen from the figure, the method of this invention can generate samples that are very close to the original two-dimensional time-frequency plot.

[0144] Then, the encoder and decoder parameters are frozen, and the discriminator's classification function is trained using only the constructed balanced dataset, with the optimization objective being solely the auxiliary classification loss. 50 training rounds were conducted to improve fault classification accuracy.

[0145] The performance of autoencoder generative adversarial networks (AGNs) based on hybrid time-frequency convolution in fault identification was analyzed. Recall and F1 score were used as evaluation metrics, and common identification methods, deep convolutional neural networks (DNNs) and support vector machines (SVMs), were selected as references for comparison. Each model was trained using both imbalanced original datasets and balanced datasets to evaluate the model's performance on small sample datasets and to verify the advantages of this invention in fault identification, as shown in Table 2.

[0146] Table 2 Data generation capability assessment and fault identification performance

[0147]

[0148] Table 2 shows the performance difference between the models trained on the original dataset and the balanced dataset. The performance of the balanced dataset is significantly improved compared to the original imbalanced dataset, indicating that the present invention has the function of simulating real samples, and the simulated time-frequency domain feature patterns are basically the same as those of real samples. At the same time, the discriminator recognition function of the present invention also has better performance than other methods.

[0149] Example 2

[0150] This embodiment provides a data enhancement system for oil discharge impact signals based on hybrid time-frequency convolution, including:

[0151] The preprocessing module is used to perform two-dimensional processing on the collected oil discharge impact signal to obtain a two-dimensional time-frequency diagram;

[0152] Generative Adversarial Networks (GANs) include an encoder, a decoder, and a discriminator. The encoder takes the input time-frequency feature map as input and obtains the corresponding latent spatial distribution. The decoder takes the encoder output and noise as input and obtains the reconstructed time-frequency map and generates a time-frequency map. The discriminator is used to determine the authenticity of the input image. Both the encoder and decoder are based on hybrid time-frequency convolution.

[0153] The preprocessing module described above performs the operation according to step S1 given in Example 1. The specific structures of the encoder, decoder, and discriminator in the generative adversarial network are as described in Example 1.

[0154] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution, characterized in that, Includes the following steps: S1, The collected oil discharge impact signal is processed into two dimensions to obtain a two-dimensional time-frequency diagram; S2, The generative adversarial network is trained using a two-dimensional time-frequency graph; the generative adversarial network includes an encoder, a decoder, and a discriminator; The encoder takes a two-dimensional time-frequency graph as input to obtain the corresponding latent spatial distribution. The encoder is composed of one or more cascaded hybrid time-frequency convolutional modules. The hybrid time-frequency convolutional module includes a first padding unit, a convolutional unit, a first fusion unit, and a downsampling unit. The first padding unit performs time-frequency axis heterogeneous padding on the two-dimensional time-frequency graph using corresponding padding strategies along the time and frequency axes. Specifically, the first padding unit uses zero-value padding along the time axis and interpolation padding along the frequency axis. The convolutional unit includes a set of parallel heterogeneous convolutional kernels, which perform convolution processing on the heterogeneously padded two-dimensional time-frequency graph using horizontal and vertical convolutional kernels of different sizes. The first fusion unit performs weighted fusion of the output features of each convolutional kernel to obtain fused time-frequency features. The downsampling unit is used to downsample the fused time-frequency features. The decoder takes the encoder's output and noise as input to obtain and generate a reconstructed time-frequency map. The decoder is symmetrically positioned relative to the encoder and consists of one or more cascaded hybrid time-frequency deconvolution modules. The hybrid time-frequency deconvolution module includes a second padding unit, a deconvolution unit, a second fusion unit, and an upsampling unit. The second padding unit uses corresponding padding strategies along the time and frequency axes to perform time-frequency heterogeneous padding on the noise input to the decoder. The deconvolution unit includes a set of parallel heterogeneous deconvolution kernels, which reconstruct the heterogeneously padded noise using horizontal and vertical deconvolution kernels of different sizes. The second fusion unit performs weighted fusion of the output features of each deconvolution kernel to obtain fused reconstructed features. The upsampling unit is used to upsample the fused reconstructed features. The padding strategy of the second padding unit is the same as that of the first padding unit. The discriminator is used to determine the authenticity of the input image; both the encoder and decoder are based on hybrid time-frequency convolution. S3 utilizes the trained decoder as a generator to obtain the generated time-frequency map, thereby achieving data enhancement of the discharge impact signal in oil.

2. The method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution according to claim 1, characterized in that, Step S1: The collected oil discharge impact signal is processed into two dimensions using short-time Fourier transform; the short-time Fourier transform is either continuous Fourier transform or discrete Fourier transform.

3. The method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution according to claim 1, characterized in that, The encoder's final output is the mean of the latent spatial variable distribution of the two-dimensional time-frequency plot. With log variance This maps the input time-frequency graph to a continuous and smooth latent distribution. : ; in, Indicates encoder; Represents a two-dimensional time-frequency graph; From the potential distribution The sampled seed noise z is used as the input to the decoder, and the noise satisfies: ; Where I represents a unit diagonal matrix; The noise in the input decoder includes: ; ; Among them, noise Used to obtain the reconstructed time-frequency map through the decoder; noise Used to obtain the generated time-frequency diagram through the decoder; Indicates standard deviation; ⊙ represents standard normal distribution noise, and ⊙ represents element-wise multiplication.

4. The method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution according to claim 3, characterized in that, For input noise and The time-frequency diagram is reconstructed through the decoder output. With the generation of time-frequency diagrams : ; ; in, This indicates the decoder.

5. The method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution according to claim 4, characterized in that, The discriminator includes a feature extraction module and a discriminant output module; the input sample for the feature extraction module is a real two-dimensional time-frequency graph. Reconstructing the time-frequency diagram and generating time-frequency diagrams The feature extraction module employs four sequentially arranged third convolutional units, each containing a convolutional layer, a normalization layer, and an activation function. The discriminant output module includes a first fully connected layer, used to output a score representing the authenticity of the input sample.

6. The method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution according to claim 5, characterized in that, The discrimination output module also includes a second fully connected layer for outputting the predicted classification type.

7. The method for enhancing oil discharge impact signal data based on hybrid time-frequency convolution according to claim 6, characterized in that, Train the generative adversarial network and set a multi-objective joint loss function. Including variational lower bound loss Combating losses Auxiliary classification loss Three parts: ; ; ; ; in, Represents the mathematical expectation of a variable; This represents the weighting coefficient of the KL divergence constraint term; This indicates the calculation of the KL divergence function; Denotes the prior distribution of the latent variable z; min G This indicates that the solution is minimized using the decoder's parameters as optimization variables; max D This means that the solution is maximized by optimizing the parameters of the discriminator. This describes the process by which the discriminator obtains the true or false scores of the input samples; This indicates the process by which the discriminator obtains the predicted classification. Indicates the predicted classification type. Indicates the true type label of the input sample; for or ; Indicates the distribution of real data The expected value of the samples obtained from sampling and their corresponding labels is calculated. This indicates the prior distribution of the latent variables. The mathematical expectation calculated from the latent variables sampled in the data and the type labels sampled from the real data distribution; , , This is a hyperparameter.

8. A system for implementing the oil discharge impact signal data enhancement method based on hybrid time-frequency convolution as described in any one of claims 1 to 7, characterized in that, include: The preprocessing module is used to perform two-dimensional processing on the collected oil discharge impact signal to obtain a two-dimensional time-frequency diagram; Generative adversarial networks include encoders, decoders, and discriminators; the encoder takes a two-dimensional time-frequency graph as input to obtain the corresponding latent spatial distribution; The decoder takes the encoder's output and noise as input to obtain the reconstructed time-frequency map and generate the time-frequency map; the discriminator is used to determine the authenticity of the input image; both the encoder and the decoder are based on hybrid time-frequency convolution.

Citation Information

Patent Citations

  • Speech enhancement method and system based on time-frequency graph convolutional network

    CN119517057A

  • Online monitoring method for product quality in solid waste recycling process

    CN120430518A