A method and system for identifying transformer acoustic fingerprint faults
By using an improved hybrid filter bank based on asymmetric time-varying overlapping wavelet packet decomposition and multi-granular dynamic multi-dimensional feature fusion, combined with a full-span sparse attention mechanism of structured factors, the problems of strong noise interference, weak feature discriminativeness, and poor model robustness in transformer fault acoustic fingerprint recognition are solved, achieving high accuracy and robust fault recognition in high-noise environments.
Patent Information
- Application Number
- CN202511375579.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing technologies have low accuracy and poor robustness in identifying transformer fault acoustic signatures in substations under high noise environments. Traditional Mel filter banks have insufficient resolution in the low-frequency band and are prone to losing nonlinear components. Furthermore, the fusion of single or static features makes it difficult to fully utilize complementary information from multi-dimensional features. Attention mechanisms are prone to indiscriminate attention across the entire range or lack of structured constraints.
A hybrid filter bank improved by asymmetric time-varying overlapping wavelet packet decomposition is constructed by combining multi-granularity dynamic multi-dimensional feature fusion and full-span sparse attention mechanism of structured factors. The power spectrum features are extracted by the hybrid filter bank improved by asymmetric time-varying overlapping wavelet packet decomposition, and the accuracy of feature extraction and recognition is improved by combining multi-granularity dynamic multi-dimensional feature fusion and sparse attention mechanism.
It effectively suppresses harmonic leakage and noise interference, improves the accuracy and robustness of transformer fault identification, especially in high-noise environments, it can more accurately distinguish fault modes, and improves the model's efficiency in capturing key fault features and its identification accuracy.
Smart Images

Figure CN120877783B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of transformer fault identification technology, and in particular relates to a method and system for identifying transformer acoustic fingerprint faults. Background Technology
[0002] The analysis and processing of transformer acoustic signature signals are of great significance for the health monitoring of power equipment. During long-term operation, transformers are subject to complex and diverse environments and different operating conditions, which can easily lead to problems such as material deterioration, fatigue wear of local components, and decreased insulation performance. This can induce abnormal vibrations, partial discharges, and other defects, and in severe cases, cause tripping and other faults, resulting in large-scale power outages. Due to factors such as magnetostriction, when there are potential mechanical structural defects in a transformer, the vibration state of components such as the core windings changes, often accompanied by the generation of abnormal acoustic signatures. Furthermore, based on the dissection of numerous damaged high-voltage substation equipment and many electrical aging tests, partial discharge is often one of the main causes of insulation damage. If not detected and investigated in time, it will gradually develop into arc discharge and spark discharge, and in severe cases, cause insulation breakdown.
[0003] In current methods for feature extraction of voiceprint signals, most studies choose to use MFCC (Medium Frequency Conversion) for feature extraction. The Mel frequency extraction method simulates the human ear's ability to distinguish between low and high frequencies, reflecting the nonlinear response of the human ear to frequency. However, the operating frequency of transformers is mainly concentrated in the 50 Hz harmonics below 1000 Hz. Most abnormal defects in transformers cause significant changes in these frequencies. The core fault characteristics of transformer voiceprint signals (such as DC bias and core vibration) are concentrated in the low-frequency range (1-1000 Hz), especially the 50 Hz power frequency and its odd harmonics (150 Hz, 250 Hz, etc.). Traditional Mel filter banks have insufficient resolution in this frequency band. Furthermore, because MFCC highly abstracts the voiceprint signal, it loses some of the original nonlinear components of the voiceprint signal. Therefore, there is an urgent need for a method and system for transformer voiceprint fault identification. Summary of the Invention
[0004] This invention provides a method and system for identifying transformer acoustic fingerprint faults, which solves the technical problems of low accuracy and poor robustness in identifying transformer acoustic fingerprint faults in substations under high noise environments.
[0005] In a first aspect, the present invention provides a method for identifying transformer acoustic fingerprint faults, comprising:
[0006] The transformer acoustic signature signal is acquired and preprocessed to obtain the target transformer acoustic signature signal, wherein the target transformer acoustic signature signal contains at least one frame of transformer acoustic signature sub-signal.
[0007] Based on the improved hybrid filter bank of asymmetric time-varying overlapping wavelet packet decomposition, the power spectrum feature vector of the hybrid filter is extracted from the acoustic fingerprint signal of the target transformer.
[0008] According to the preset multi-granularity dynamic multi-dimensional feature fusion algorithm, the power spectrum feature vector of the hybrid filter is fused with the preset multi-dimensional feature vector to obtain the MDF vector of the transformer acoustic text sub-signal of at least one frame, and the MDF vectors of each frame transformer acoustic text sub-signal are spliced together to generate a fused feature vector corresponding to the target transformer acoustic text signal.
[0009] A transformer voiceprint recognition model is constructed based on a full-span sparse attention mechanism using structured factors.
[0010] The fused feature vector is input into the transformer acoustic signature recognition model, and the transformer acoustic signature recognition model outputs a fault identification result corresponding to the transformer acoustic signature signal.
[0011] Secondly, the present invention provides a transformer acoustic fingerprint fault identification system, comprising:
[0012] The acquisition module is configured to acquire transformer acoustic fingerprint signals and preprocess the transformer acoustic fingerprint signals to obtain target transformer acoustic fingerprint signals, wherein the target transformer acoustic fingerprint signals contain at least one frame of transformer acoustic fingerprint sub-signals.
[0013] The extraction module is configured to extract the hybrid filter power spectrum feature vector from the acoustic signature signal of the target transformer based on the hybrid filter bank improved by asymmetric time-varying overlapping wavelet packet decomposition.
[0014] The splicing module is configured to fuse the power spectrum feature vector of the hybrid filter with the preset multidimensional feature vector according to the preset multi-granularity dynamic multidimensional feature fusion algorithm to obtain the MDF vector of the at least one frame of transformer acoustic text sub-signal, and splice the MDF vectors of each frame of transformer acoustic text sub-signal to generate a fused feature vector corresponding to the target transformer acoustic text signal.
[0015] The module is configured to build a transformer voiceprint recognition model based on a full-span sparse attention mechanism using structured factors.
[0016] The output module is configured to input the fused feature vector into the transformer acoustic signature recognition model, and the transformer acoustic signature recognition model outputs a fault identification result corresponding to the transformer acoustic signature signal.
[0017] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the transformer acoustic fingerprint fault identification method according to any embodiment of the present invention.
[0018] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the steps of the transformer acoustic fingerprint fault identification method according to any embodiment of the present invention.
[0019] The transformer acoustic fingerprint fault identification method and system of this application achieves significant advantages over existing technologies through the following specific technical means:
[0020] 1) Existing filter banks often employ fixed overlap widths, which can easily lead to harmonic leakage in the low-frequency critical band (close to the fundamental frequency) and high-frequency noise propagation. This application addresses these issues by using "asymmetric allocation" (wide overlap at low frequencies to protect critical bands, and narrow overlap at high frequencies to suppress noise) and "dynamic adjustment of overlap width based on local SNR" to precisely adapt to the signal characteristics of different frequency bands. This effectively suppresses harmonic leakage and noise interference, improves the accuracy of feature extraction, and provides higher-quality power spectrum features for subsequent identification.
[0021] 2) Existing technologies mostly employ single or static feature fusion, making it difficult to fully utilize the complementary information of multi-dimensional features (such as power, entropy, time-frequency characteristics, etc.). This application, through "multi-granularity dynamic fusion," dynamically integrates the power spectrum features of the hybrid filter with preset multi-dimensional features into an MDF vector, and concatenates them into a fused feature vector. This achieves deep synergy of features of different granularities and dimensions, enhances the distinguishability of features (such as the feature differences between fault and normal voiceprints), and solves the problem of low recognition accuracy caused by insufficient utilization of feature information.
[0022] 3) Existing attention mechanisms are prone to problems such as indiscriminate attention across the entire range (redundant computation) or lack of structured constraints (insufficient attention to key fault features). This application introduces a structured factor, enabling attention to focus on structured information that conforms to domain priors (such as the time-frequency continuous block of fault soundprints) across the entire range. At the same time, it suppresses noise interference through sparsity, improving the model's efficiency and robustness in capturing key fault features, especially in high-noise environments where it can more accurately distinguish fault modes.
[0023] The above-mentioned technical means optimize the entire process of "feature extraction-feature fusion-model recognition", effectively solving the problems of "strong noise interference, weak feature distinguishability and poor model robustness" in transformer fault acoustic recognition under high noise environment, and significantly improving the accuracy and robustness of fault recognition. Attached Figure Description
[0024] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 A flowchart of a transformer acoustic signature fault identification method provided in an embodiment of the present invention;
[0026] Figure 2 A schematic diagram of transformer near-sound field calculation simulation is provided for a specific embodiment of the present invention;
[0027] Figure 3 This is a structural block diagram of a transformer acoustic fingerprint fault identification system provided in an embodiment of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0030] Please see Figure 1 The diagram shows a flowchart of a transformer acoustic fault identification method according to this application.
[0031] like Figure 1 As shown, the transformer acoustic fingerprint fault identification method specifically includes the following steps:
[0032] Step S101: Obtain the transformer acoustic signature signal and preprocess the transformer acoustic signature signal to obtain the target transformer acoustic signature signal, wherein the target transformer acoustic signature signal contains at least one frame of transformer acoustic signature sub-signal.
[0033] In this step, a transformer acoustic signature acquisition platform is built, and on-site acoustic signature acquisition is conducted at a 220kV substation. The steps include: The transformer acoustic signature acquisition platform comprises the following components: acoustic signature signal sensor (microphone), preamplifier, data acquisition unit, and storage and analysis unit.
[0034] The acoustic signature monitoring device must meet IEC IP protection standards (IP55) and should also have a windproof cover, which should be inspected and replaced regularly to ensure reliable sound acquisition in outdoor environments. According to the international measurement standard IEC60651, acoustic signal measurement should cover the audible range of 20 Hz to 20 kHz; therefore, based on the Nyquist theorem, the device uses a sampling frequency of 48000 Hz. A signal transmission line resistant to strong magnetic field interference is used to effectively reduce external electromagnetic interference.
[0035] Since the installation location has a significant impact on the monitoring results, it is necessary to select acoustic signature monitoring points based on the acoustic field distribution characteristics of the transformer. The overall process is divided into three parts: signal measurement, acoustic field calculation, and measurement point analysis. The signal measurement part primarily provides the acoustic field reconstruction data foundation for acoustic field calculation, and the results of the acoustic field calculation, along with the collected sound source data, will be used together to calculate the evaluation indicators for the monitoring points.
[0036] Signal measurement section: 1) Measure the three-dimensional dimensions of the transformer, the distance between the transformer tank wall and the firewall, and other spatial parameters; 2) Divide the surface of the target transformer into several equivalent source regions and collect equivalent sound source signals using a microphone; 3) Perform phase correction on the equivalent sound source signals; 4) Perform fast Fourier transform on the equivalent sound source signals to calculate the frequency domain information of the equivalent sound source.
[0037] Sound field calculation part: 1) Based on the measured information, simplify and construct a 3D model of the computational domain, and then mesh it; 2) Define the parameters of the computational domain and define the environmental boundary conditions according to the simplified model; 3) Use the discontinuous Galerkin finite element method to calculate the transient change process of the sound field; 4) Extract the sound pressure time domain signal of the candidate monitoring point and perform fast Fourier transform to obtain its corresponding frequency domain information.
[0038] Measurement point analysis section: 1) Calculate the Pearson correlation matrix of the spectrum between all candidate measurement points and the equivalent sound source; 2) Calculate the Pearson equalization correlation coefficient based on the Pearson correlation matrix calculation results; 3) Select the candidate monitoring point with the highest Pearson equalization correlation coefficient as the optimal monitoring point.
[0039] Near-field acoustic field calculation and simulation of transformers, such as Figure 2 As shown in the figure, based on near-field calculation and simulation, the acoustic sensor should be installed on a floor stand at a distance of 0.6 to 2 meters from the device, and 2 to 3 sensors should be configured for one device.
[0040] For example, the field measurement points of the dual-microphone array are arranged as follows: the first microphone is located on the high-voltage side of the transformer, 100cm from the outer wall of the transformer tank and 35cm from the ground. The second microphone is located on the on-load tap changer side of the transformer, 60cm from the tap changer and 35cm from the ground.
[0041] The acquired transformer acoustic signature signal undergoes preprocessing, including the following steps: sampling quantization, pre-emphasis, and frame-by-frame windowing. Because audio signals exhibit time-varying and non-stationary characteristics, directly processing long, continuous audio signals may fail to capture their local features. Frame-by-frame windowing allows non-stationary signals to approximate as stationary signals within a short timeframe, facilitating spectral analysis. Windowing reduces boundary effects and spectral leakage. Pre-emphasis enhances the high-frequency components of the signal, improving its frequency characteristics and enhancing subsequent processing effectiveness.
[0042] Step S102: Extract the hybrid filter power spectrum feature vector from the acoustic signature signal of the target transformer based on the hybrid filter bank improved by asymmetric time-varying overlapping wavelet packet decomposition.
[0043] In this step, the Mel frequency extraction method simulates the human ear's ability to distinguish between low and high frequencies, reflecting the nonlinear response of the human ear to frequency. The conversion of frequency f to Mel frequency m is shown in the following equation:
[0044] ,
[0045] In the formula, This is the transformation function that converts frequency f to the corresponding Mel frequency.
[0046] However, the operating frequency of transformers is mainly concentrated in the 50 Hz harmonics below 1000 Hz. Most abnormal defects in transformers will cause significant changes in these frequencies. The core fault characteristics of transformer acoustic signatures (such as DC bias, core vibration, etc.) are concentrated in the low-frequency range (1-1000 Hz), especially the 50 Hz power frequency and its odd harmonics (150 Hz, 250 Hz, etc.). Traditional Mel filter banks have insufficient resolution in this frequency band, leading to the following problems:
[0047] Key harmonic ambiguity: The 50Hz harmonic frequency is easily smoothed by a broadband filter, weakening the fault characteristics.
[0048] Noise sensitive: Wideband is susceptible to power frequency interference and random noise pollution.
[0049] Therefore, preserving the resolution of key frequency bands is particularly important in frequency extraction.
[0050] The peak frequency of a Mel filter bank is difficult to align with the 50 Hz octave (critical band), so using a Mel filter bank may result in the loss of critical frequency characteristics below 1000 Hz.
[0051] Therefore, to address the problem of key frequency feature loss below 1kHz, a hybrid filter bank based on asymmetric time-varying overlapping wavelet packet decomposition is proposed. The power spectrum feature vector of the hybrid filter is extracted from the acoustic signature signal of the target transformer. Specifically:
[0052] The division of the low-frequency key band follows the principles of narrow-band high resolution and overlap noise reduction, and adopts the 20Hz frequency band overlap maximum sampling method. The low-frequency key band division is shown in the following formula:
[0053] ,
[0054] In the formula, Let be the starting frequency of the i-th frequency band. Let be the termination frequency of the i-th frequency band. It is a fixed single frequency band with a bandwidth of 20Hz. The step size for moving the starting point of the frequency band is determined by the overlap amount. , It is a time-varying anti-noise overlapping frequency band with a width of 5~15Hz, which changes dynamically with the signal-to-noise ratio; therefore, its length is 5~15Hz; the total span of the low frequency band is 1~1kHz.
[0055] To accurately cover frequency points that are integer multiples of 50Hz, the center frequency of the frequency band is forcibly adjusted as shown in the following formula:
[0056] ,
[0057] In the formula, round is the floor function. This ensures that key harmonics such as 50 Hz, 150 Hz, and 250 Hz are located in the center region of the frequency band.
[0058] Maximum sampling method, for each frequency band Extracting the maximum power spectrum , The maximum power spectral density within the i-th frequency band. Here is the power spectral density value at frequency point k;
[0059] The maximum sampling method can highlight the frequency components with the highest energy in the frequency band and enhance fault-related harmonics; and because the energy distribution of random noise is relatively uniform, it can suppress broadband noise.
[0060] An asymmetric time-varying overlapping wavelet packet decomposition method is proposed, which dynamically adjusts the time-varying anti-noise overlapping frequency band based on the local signal-to-noise ratio (SNR). To suppress harmonic leakage and noise interference, the wavelet packet uses the Daubechies 4 (db4) wavelet basis to process each frame of signal. , The decomposition level of the wavelet packet is [value]. The total number of frequency bands obtained after wavelet packet decomposition is given by the following formula, which generates the set of sub-band coefficient matrices:
[0061] ,
[0062] In the formula, Given the input frame signal, DWT() is the wavelet decomposition function. These are low-pass filters and high-pass filters, respectively. This is the set of sub-band coefficient matrices generated by the b-th sub-band after the j-th level wavelet packet decomposition. The decomposition level of the wavelet packet;
[0063] Asymmetric dynamic time-varying noise reduction overlap width As shown in the following formula:
[0064] ,
[0065] In the formula, Set the base overlap width to 10Hz. For local signal-to-noise ratio, The maximum signal-to-noise ratio threshold is set to 20dB. The lower the signal-to-noise ratio, the larger the overlap width, and the stronger the noise immunity.
[0066] The low-frequency key band filter bank and the Mel full-band hybrid filter are as follows:
[0067] ,
[0068] ,
[0069] The hybrid filter bank employs an asymmetric allocation method, resulting in wider overlap on the low-frequency side (closer to the fundamental frequency) to protect critical frequency bands, and narrower overlap on the high-frequency side to suppress noise propagation. This is illustrated in the following equation:
[0070] ,
[0071] In the formula, This refers to the overlap width on the low-frequency side of the hybrid filter bank. The overlap width on the high-frequency side;
[0072] Extracting sub-bands:
[0073] ,
[0074] In the formula, To extract the signal from the j-th sub-band and the b-th frequency band, The original length of the j-th subband. The truncation index range for subband signals (from) Starting position, to (Termination at position, retaining the valid middle segment);
[0075] Extruded overlapping subbands Energy-entropy joint feature extraction is performed on the overlapping subbands, and the extracted joint energy-entropy feature pairs are shown in the following equation:
[0076] ,
[0077] The joint feature pairs are as follows:
[0078] ,
[0079] Because Mel filters are prone to losing features in the high-frequency band, the IMFCC feature is used in the high-frequency band.
[0080] The final hybrid filter power is dynamically weighted and summed as shown in the following equation:
[0081] ,
[0082] In the formula, For the power of the hybrid filter, For joint feature pairs, The output power of the Mel full-band hybrid filter bank, As an energy characteristic, Entropy features The real number field containing the joint eigenvectors. For adaptive weights, This represents the total number of samples in the sub-band time series. The frequency domain power spectrum of the j-th sub-band in the b-th frequency band is... Let be the significance index of the j-th sub-band in the b-th frequency band. This is the slope adjustment parameter for the Sigmoid function. This represents the power percentage at frequency f. Let be the frequency domain signal of the j-th sub-band within the b-th frequency band. for The mean, The total number of low-frequency features. For the output power of the low-frequency critical band filter bank, The total number of high-frequency features. For the output power of the high-frequency Mel filter bank, The starting frequency, For frequency band overlap width, For the termination frequency, This is the input frequency.
[0083] Step S103: According to the preset multi-granularity dynamic multi-dimensional feature fusion algorithm, the power spectrum feature vector of the hybrid filter is fused with the preset multi-dimensional feature vector to obtain the MDF vector of the at least one frame of transformer acoustic text sub-signal, and the MDF vectors of each frame of transformer acoustic text sub-signal are spliced together to generate a fused feature vector corresponding to the target transformer acoustic text signal.
[0084] In this step, the transformer acoustic signature signal contains rich physical characteristics and nonlinear factors. In order to deeply analyze and mine the fault-related information contained in the transformer acoustic signature signal, a multi-dimensional physical quantity knowledge feature fusion method is proposed. At the same time, transformer physical formulas and empirical knowledge are introduced to extract and reorganize the transformer physical feature parameters to construct a multi-granularity physical knowledge feature vector, which provides a more accurate information representation for subsequent identification tasks, thereby enhancing the reliability and generalization of subsequent models. A fault identification system is built based on deep learning algorithms to train the extracted multi-granularity knowledge feature vector to achieve accurate identification of transformer fault types.
[0085] For the target transformer acoustic signature signal after framing, windowing, and pre-emphasis, the following features are extracted: 50Hz odd-even octave ratio feature vector (harmonic characteristics), high-frequency energy ratio feature vector (high-frequency characteristics), hybrid filter power spectrum feature vector, spectrum center feature vector, spectrum broadening feature vector (spectral characteristics), and vibration entropy feature vector (vibration characteristics). The MDF vector (multidimensional fusion feature vector) of the frame signal is obtained by concatenating the feature vectors. The fusion feature vector corresponding to the target transformer acoustic signature signal is generated by concatenating the MDF vectors of the frame-by-frame signal.
[0086] Specifically, the 50Hz odd-even octave ratio eigenvector:
[0087] Abnormal discharges in transformers (such as partial discharges) generate vibration and acoustic signals at specific frequencies. These signals typically contain harmonic components related to the transformer's power supply frequency (50Hz or 60Hz). By analyzing the ratio of odd to even harmonics at 50Hz (odd-even harmonic ratio) in the signal, the spectral characteristics caused by abnormal discharge can be revealed, thus enabling the detection and identification of abnormal transformer discharges. During normal operation, the vibration and acoustic signals of the transformer are mainly caused by the magnetostrictive effect of the iron core and electromagnetic forces. The spectral characteristics of these effects are mainly concentrated at the 50Hz fundamental frequency and its even harmonics (100Hz, 200Hz, etc.), with even harmonic components dominating and odd harmonic components relatively weak. Abnormal discharges (such as partial discharges and surface discharges) introduce nonlinear effects into the electric field, leading to a significant increase in the energy of odd harmonics (such as 150Hz, 250Hz). In addition, arcing and plasma oscillations during the discharge process may also introduce non-fundamental frequency components, altering the harmonic energy distribution. The formula for the eigenvector of the 50Hz odd-even harmonic ratio is:
[0088] ,
[0089] In the formula, This is the eigenvector of the 50Hz odd-even harmonic ratio. For a frame of transformer acoustic signature signal in the first... The power of the harmonic component;
[0090] High-frequency energy ratio eigenvector:
[0091] The high-frequency energy ratio eigenvector is an effective acoustic signature parameter. By analyzing the energy distribution of the acoustic signature signal across different frequency bands, specific fault patterns in transformers can be identified. The high-frequency components of the acoustic signature signal of a transformer show significant differences between normal operation and fault conditions. These differences primarily stem from changes in physical processes such as mechanical vibration, partial discharge, and electromagnetic forces. During normal operation, the main energy of the transformer's acoustic signature signal is concentrated in the low-frequency range (such as the 50Hz fundamental frequency and its even harmonics), due to periodic vibrations generated by the magnetostrictive effect of the iron core and the action of electromagnetic forces. High-frequency energy during normal operation is mainly caused by weak mechanical vibrations and local eddy current noise, resulting in a relatively low high-frequency energy component. Faults (such as partial discharge, loose windings, and abnormal core vibration) excite more nonlinear and non-periodic components, significantly increasing the high-frequency energy. For example, the high-frequency acoustic wave energy generated by partial discharge is concentrated in the ultrasonic and high-frequency bands from 1kHz to 1MHz. Abnormal mechanical vibrations, such as loose windings, lead to high-frequency vibrations and impact noise; iron core faults, such as magnetic saturation effects, increase high-frequency harmonic components.
[0092] The high-frequency energy ratio describes the proportion of high-frequency energy in a signal relative to the total energy across the entire frequency band. It is an important parameter characterizing the frequency distribution of a signal. The expression for calculating the high-frequency energy ratio eigenvector is as follows:
[0093] ,
[0094] In the formula, This is the high-frequency energy ratio eigenvector. This represents the upper limit of the high-frequency energy range. This represents the lower limit of the high-frequency energy range. This is the upper limit of the entire frequency band. This is the lower limit of the full frequency band range. The power of the acoustic signature signal of the target transformer at frequency f;
[0095] Vibrational entropy eigenvector:
[0096] The vibration entropy eigenvector reflects the complexity and uncertainty of transformer vibration characteristics. By incorporating vibration entropy into the signal composition of the MDF eigenvector, the noise immunity of the MDF eigenvector can be improved. The expression for the vibration entropy eigenvector is:
[0097] ,
[0098] In the formula, The frequency domain signal of the framed signal. The vibration entropy eigenvector, Total number of frames;
[0099] Spectrum center eigenvector and spectrum broadening eigenvector:
[0100] The spectral center eigenvector characterizes the spectral distribution characteristics of the voiceprint signal, i.e., the energy concentration range of the voiceprint signal. The spectral broadening eigenvector, on the other hand, describes the width of the audio signal's spectral energy distribution across frequency ranges. The expression for calculating the spectral center eigenvector is:
[0101] ,
[0102] The expression for calculating the spectral broadening eigenvector is:
[0103] ,
[0104] In the formula, The signal amplitude, The eigenvector of the spectrum center This is the spectral broadening feature vector.
[0105] The final MDF vector is obtained, expressed as:
[0106]
[0107] In the formula, For MDF vectors, This represents the power spectrum eigenvector of the hybrid filter. For connection operations.
[0108] It should be noted that concatenating the MDF vectors of the various frame transformer acoustic signature sub-signals to generate a fused feature vector with the target transformer acoustic signature specifically includes:
[0109] Calculate the local signal-to-noise ratio (SNR) and generate dynamic weights based on the SNR. The expression for the dynamic weights is as follows:
[0110] ,
[0111] ,
[0112] In the formula, For dynamic weights, It is the Sigmoid activation function. For learnable weight matrix, For local signal-to-noise ratio, For learnable bias terms, For the m-th frame of the sound signal, the first... The power of the point, The set of locations determined to be "signal-dominant" The set of locations determined to be "noise-dominant";
[0113] The MDF vector is Top-K sparsified based on dynamic weights, retaining the top K high-weight features and setting the rest to zero, resulting in the target MDF vector. The expression for the target MDF vector is:
[0114] ,
[0115] ,
[0116] In the formula, For the target Top-K sparse MDF feature vector, The number of feature vectors to retain. The total dimension of the features. The global signal-to-noise ratio reflects the overall noise level of the signal. The slope parameter controls the steepness of the Sigmoid curve. This is a threshold parameter used to control the position of the midpoint of the Sigmoid curve. This indicates taking the top K MDF feature vectors after sorting them based on their dynamic weights.
[0117] The MDF vectors of each target are concatenated to generate the MDF spectrum of the target transformer acoustic signature signal. The expression of the MDF spectrum is as follows:
[0118] ,
[0119] ,
[0120] ,
[0121] In the formula, This is an MDF spectrum. For noise-aware masking, This is the Hadamard product (element-by-element multiplication), and T is the transpose operation. This is the attention weight matrix (which combines sparsity and attention weights). The query vector corresponding to the target Top-K sparse MDF feature vector. The key vectors corresponding to the target Top-K sparse MDF feature vectors. For the dimensions of the query / key vector, The mask generation function takes the target MDF vector and the local signal-to-noise ratio (SNR) m as input and outputs the mask fragment of that local region. Finally, by concatenating all local masks, the global mask M is obtained.
[0122] Step S104: Construct a transformer voiceprint recognition model based on a full-span sparse attention mechanism using structured factors.
[0123] In this step, the transformer voiceprint recognition model architecture is specifically as follows:
[0124] To address the Transformer's poor ability to extract fine-grained local feature patterns and its tendency to lose local features, a power transformer fault acoustic signature identification method based on an improved Transformer is proposed. Leveraging the advantages of CNNs in local feature extraction, the new model can simultaneously consider both local features and global representations. A dynamic sparse-full attention mechanism enhances the ability to model local information and effectively models long-distance dependencies, optimizing computational efficiency and information interaction while maintaining low computational and memory complexity. The model's powerful feature extraction and learning capabilities further strengthen its noise resistance.
[0125] The transformer voiceprint recognition model, also known as the improved Conformer model, consists of multiple identical Conformer Blocks. The Conformer Block network architecture expression is as follows:
[0126] ,
[0127] In the formula, LayerNorm() is the layer normalization function. For full-span sparse attention output based on structured factors, For the output of the convolution module, Input to the model, It is a feedforward network;
[0128] Unlike the Transformer, which has only one feedforward network, the Conformer has two feedforward networks, one before and one after self-attention. These feedforward modules use linear transformations to map and transform the input, increasing the model's expressive power. This improvement in feedforward networks enhances the model's expressive power, while also improving its ability to model local features and global dependencies, making it superior to the Transformer in speech recognition and NLP tasks. Furthermore, the introduction of the Swish activation function incorporates non-linear relationships and helps the model converge faster. The Swish activation function is as follows:
[0129] ,
[0130] In the formula, Input to the model, Use the Sigmoid activation function;
[0131] The Swish activation function is a smooth, non-linear activation function that combines the characteristics of the siqmoid function. Through its self-adjusting mechanism, it can outperform the traditional ReLU function in certain tasks. Its smoothness and differentiability also give it good gradient propagation properties in deep neural networks, making it particularly suitable for training deep models.
[0132] The convolutional module performs convolution operations on the input over time to extract local features, helping to capture the local structure and patterns of the input sequence. The convolutional module first uses a gating mechanism, specifically implemented with pointwise convolution and gated linear units (GLUs). This gating mechanism effectively controls the information flow. Following this, the model employs a one-dimensional depthwise separable convolutional layer to improve the efficiency and effectiveness of the convolutional operation. To facilitate the training of deep models, batch normalization is applied immediately after the convolutional layer.
[0133] Full-span sparse attention mechanism based on structured factors:
[0134] From the perspective of conditional expectation, for position i, the attention mechanism formula can be rewritten as:
[0135] ,
[0136] ,
[0137] In the formula, The input feature vector at the i-th position in the sequence Attention output, Let j be the value vector at the j-th position in the sequence. For attention weights, Attention weights right Conditional expectation, For the dimensions of the query / key vector, For query vector, This is the transpose of the key vector;
[0138] A full-span sparse attention mechanism based on structured factors introduces support factor variables. conditional probability The structured decomposition is performed using the following formula:
[0139] ,
[0140] In the formula, Let be the joint conditional probability, representing the relationship between position j and the r-th class support factor given position i. The probability of simultaneous association;
[0141] Based on the above decomposition process, the full-span sparse attention mechanism based on the structure factor is as follows:
[0142] ,
[0143] In the formula, For structured sparse attention at location The output at that location, The input feature is the i-th position of the sequence. Let i be the set of global sparse features associated with the i-th position. Let be the global sparse attention weight between position i and position j. The number of structured sets associated with the i-th position. For the r-th type of support factor, Let be the structural preference probability, representing the probability of the r-th class of support factors appearing at a given position i. Let be the intrastructural association probability, representing the inherent probability that position j belongs to the corresponding structure given a support factor of class r. Let j be the value vector at the j-th position in the sequence. Let be the global sparse feature set at position i. This is an indicator function used to indicate whether position j is within the range of global sparsity.
[0144] This demonstrates how a full-span sparse attention mechanism based on structured factors can form the final conditional expectation by combining direct attention and local attention, achieving full attention while maintaining low computational and memory complexity.
[0145] Using a fully connected layer as the classifier in the Conformer model, the formula is as follows:
[0146] ,
[0147] In the formula, For the classifier output of the Conformer model, To A fully connected operation is performed, and the final fault identification result is obtained through the fully connected layer.
[0148] Step S105: Input the fused feature vector into the transformer acoustic signature recognition model, and the transformer acoustic signature recognition model outputs the fault identification result corresponding to the transformer acoustic signature signal.
[0149] Please see Figure 3 The diagram shows a structural block diagram of a transformer acoustic fingerprint fault identification system according to this application.
[0150] like Figure 3 As shown, the transformer acoustic fingerprint fault identification system 200 includes an acquisition module 210, an extraction module 220, a splicing module 230, a construction module 240, and an output module 250.
[0151] The acquisition module 210 is configured to acquire a transformer acoustic signature signal and preprocess the signal to obtain a target transformer acoustic signature signal, wherein the target signal contains at least one frame of transformer acoustic signature sub-signal; the extraction module 220 is configured to extract a hybrid filter power spectrum feature vector from the target signal based on a hybrid filter bank improved by asymmetric time-varying overlapping wavelet packet decomposition; the splicing module 230 is configured to fuse the hybrid filter power spectrum feature vector with a preset multi-dimensional feature fusion algorithm to obtain the MDF vector of the at least one frame of transformer acoustic signature sub-signal, and splice the MDF vectors of each frame to generate a fused feature vector corresponding to the target signal; the construction module 240 is configured to construct a transformer acoustic signature recognition model based on a full-span sparse attention mechanism of structured factors; and the output module 250 is configured to input the fused feature vector into the model and output a fault identification result corresponding to the signal.
[0152] It should be understood that Figure 3 The modules and references described in the document Figure 1 The steps described in the text correspond to those in the method described above. Therefore, the operations, features, and corresponding technical effects described above also apply to the method described in the text. Figure 3 The various modules in the document will not be described in detail here.
[0153] In other embodiments, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the transformer acoustic fingerprint fault identification method in any of the above method embodiments.
[0154] In one embodiment, the computer-readable storage medium of the present invention stores computer-executable instructions, which are configured as follows:
[0155] The transformer acoustic signature signal is acquired and preprocessed to obtain the target transformer acoustic signature signal, wherein the target transformer acoustic signature signal contains at least one frame of transformer acoustic signature sub-signal.
[0156] Based on the improved hybrid filter bank of asymmetric time-varying overlapping wavelet packet decomposition, the power spectrum feature vector of the hybrid filter is extracted from the acoustic fingerprint signal of the target transformer.
[0157] According to the preset multi-granularity dynamic multi-dimensional feature fusion algorithm, the power spectrum feature vector of the hybrid filter is fused with the preset multi-dimensional feature vector to obtain the MDF vector of the transformer acoustic text sub-signal of at least one frame, and the MDF vectors of each frame transformer acoustic text sub-signal are spliced together to generate a fused feature vector corresponding to the target transformer acoustic text signal.
[0158] A transformer voiceprint recognition model is constructed based on a full-span sparse attention mechanism using structured factors.
[0159] The fused feature vector is input into the transformer acoustic signature recognition model, and the transformer acoustic signature recognition model outputs a fault identification result corresponding to the transformer acoustic signature signal.
[0160] Computer-readable storage media may include a stored program area and a stored data area, wherein the stored program area may store an operating system and an application program required for at least one function; the stored data area may store data created based on the use of the transformer acoustic fingerprint fault identification system, etc. Furthermore, the computer-readable storage medium may include high-speed random access memory, and may also include memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the computer-readable storage medium may optionally include memory remotely configured relative to a processor, which can be connected to the transformer acoustic fingerprint fault identification system via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0161] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 4 As shown, the device includes a processor 310 and a memory 320. The electronic device may also include an input device 330 and an output device 340. The processor 310, memory 320, input device 330, and output device 340 can be connected via a bus or other means. Figure 4 Taking a bus connection as an example, the memory 320 is the computer-readable storage medium described above. The processor 310 executes various server functions and data processing by running non-volatile software programs, instructions, and modules stored in the memory 320, thereby implementing the transformer acoustic fingerprint fault identification method described in the above embodiment. The input device 330 can receive input digital or character information and generate key signal inputs related to user settings and function control of the transformer acoustic fingerprint fault identification system. The output device 340 may include a display screen or other display device.
[0162] The aforementioned electronic device can execute the method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in the embodiments of the present invention.
[0163] In one implementation, the above-described electronic device is applied to a transformer acoustic fingerprint fault identification system for a client, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to:
[0164] The transformer acoustic signature signal is acquired and preprocessed to obtain the target transformer acoustic signature signal, wherein the target transformer acoustic signature signal contains at least one frame of transformer acoustic signature sub-signal.
[0165] Based on the improved hybrid filter bank of asymmetric time-varying overlapping wavelet packet decomposition, the power spectrum feature vector of the hybrid filter is extracted from the acoustic fingerprint signal of the target transformer.
[0166] According to the preset multi-granularity dynamic multi-dimensional feature fusion algorithm, the power spectrum feature vector of the hybrid filter is fused with the preset multi-dimensional feature vector to obtain the MDF vector of the transformer acoustic text sub-signal of at least one frame, and the MDF vectors of each frame transformer acoustic text sub-signal are spliced together to generate a fused feature vector corresponding to the target transformer acoustic text signal.
[0167] A transformer voiceprint recognition model is constructed based on a full-span sparse attention mechanism using structured factors.
[0168] The fused feature vector is input into the transformer acoustic signature recognition model, and the transformer acoustic signature recognition model outputs a fault identification result corresponding to the transformer acoustic signature signal.
[0169] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of various embodiments or some parts of embodiments.
[0170] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A transformer voiceprint fault recognition method, characterized in that, The method comprises the following steps: obtaining a transformer voiceprint signal, and preprocessing the transformer voiceprint signal to obtain a target transformer voiceprint signal, wherein the target transformer voiceprint signal comprises at least one transformer voiceprint sub-signal frame; extracting a hybrid filter power spectrum feature vector from the target transformer voiceprint signal according to an improved hybrid filter bank based on an asymmetric time-varying overlapping wavelet packet decomposition, wherein the expression of the hybrid filter power spectrum feature vector is: , , , , , wherein, is the mixed filter power, is the joint feature pair, is the mel full-band mixed filter bank output power, is the energy feature, is the entropy feature, is the real number domain where the joint feature vector is located, is the adaptive weight, is the total number of samples of the sub-band time series, is the frequency domain power spectrum of the jth sub-band in the bth frequency band, is the saliency index of the jth sub-band in the bth frequency band, is the slope adjustment parameter of the Sigmoid function, is the power proportion at frequency f, is the frequency domain signal of the jth sub-band in the bth frequency band, is the mean of is the total number of low-frequency features, is the low-frequency key band filter bank output power, is the total number of high-frequency features, is the high-frequency band mel filter bank output power, is the start frequency, is the frequency band overlap width, is the end frequency, is the input frequency; fusing the hybrid filter power spectrum feature vector and a preset multi-dimensional feature vector according to a preset multi-granularity dynamic multi-dimensional feature fusion algorithm to obtain an MDF vector of the at least one transformer voiceprint sub-signal frame, and splicing the MDF vectors of the transformer voiceprint sub-signal frames to generate a fusion feature vector corresponding to the target transformer voiceprint signal; constructing a transformer voiceprint recognition model based on a full-span sparse attention mechanism of a structured factor; inputting the fusion feature vector into the transformer voiceprint recognition model, and outputting a fault recognition result corresponding to the transformer voiceprint signal by the transformer voiceprint recognition model.
2. The transformer voiceprint fault identification method of claim 1, wherein, Before the step of fusing the hybrid filter power spectrum feature vector and a preset multi-dimensional feature vector according to a preset multi-granularity dynamic multi-dimensional feature fusion algorithm to obtain an MDF vector of the at least one transformer voiceprint sub-signal frame, the method further comprises: obtaining a multi-dimensional feature vector, wherein the multi-dimensional feature vector comprises a 50Hz odd-even frequency ratio feature vector, a high-frequency energy ratio feature vector, a vibration entropy feature vector, a spectral center feature vector, and a spectral spread feature vector; wherein the expression of the 50Hz odd-even frequency ratio feature vector is: , In the formula, is a 50 Hz odd power bit feature vector, is a power of a 50 Hz odd power component of a transformer voiceprint sub-signal in a frame; is a 50 Hz even power bit feature vector, is a power of a 50 Hz even power component of a transformer voiceprint sub-signal in a frame; the expression of the high-frequency energy ratio feature vector is: , wherein, is a high frequency energy range upper limit, is a high frequency energy range lower limit, is a high frequency energy range lower limit, is a full frequency range upper limit, is a full frequency range lower limit, is a power of the target transformer voiceprint signal at frequency f; the expression of the vibration entropy feature vector is: , In the formula, is a frequency domain signal of the frame signal, is a vibration entropy feature vector, is the total number of frames; the expression of the spectral center feature vector is: , the expression of the spectral spread feature vector is: , wherein is the signal amplitude, is the spectral center eigenvector, is the spectral spread eigenvector.
3. The transformer voiceprint fault identification method of claim 2, wherein, the expression of the MDF vector is: wherein is the MDF vector, is the hybrid filter power spectrum feature vector, is the concatenation operation.
4. The transformer voiceprint fault identification method of claim 3, wherein, the step of splicing the MDF vectors of the transformer voiceprint sub-signal frames to generate a fusion feature vector corresponding to the target transformer voiceprint signal comprises: calculating a local signal-to-noise ratio, and generating a dynamic weight according to the local signal-to-noise ratio, wherein the expression of the dynamic weight is: , , wherein, is a dynamic weight, is a Sigmoid activation function, is a learnable weight matrix, is a local signal-to-noise ratio, is a learnable bias term, and is a set of positions determined to be "signal dominant", is a set of positions determined to be "noise dominant"; performing Top-K sparsification on the MDF vector according to the dynamic weight, retaining the first K high-weight features, and setting the remaining features to zero to obtain a target MDF vector, wherein the expression of the target MDF vector is: , , In the formula, is the target Top-K sparse MDF feature vector, is the number of reserved feature vectors, is the total dimension of features, is the global signal-to-noise ratio, reflecting the noise level of the whole signal, is the slope parameter, used to control the steepness of the Sigmoid curve, is the threshold parameter, used to control the midpoint position of the Sigmoid curve, represents taking the first K MDF feature vectors based on the dynamic weight size ranking; splicing the target MDF vectors to generate an MDF spectrogram of the target transformer voiceprint signal, wherein the expression of the MDF spectrogram is: , , , wherein, is the MDF spectrogram, is the noise-aware mask, is the Hadamard product, T is the transpose operation, is the attention weight matrix, is the query vector corresponding to the target Top-K sparsified MDF feature vector, is the key vector corresponding to the target Top-K sparsified MDF feature vector, is the dimension of the query / key vector, is the mask generation function.
5. The transformer voiceprint fault identification method of claim 1, wherein, the expression of the full-span sparse attention mechanism is: , wherein, is the output of the structured sparse attention at position , is the input feature of the i-th position in the sequence, is the global sparse feature set associated with the i-th position, is the global sparse attention weight of the i-th position on the j-th position, is the number of structured sets associated with the i-th position, is the r-th support factor, is the structural preference probability, representing the preference probability of the occurrence of the r-th support factor given the i-th position, is the intra-structure correlation probability, representing the inherent probability of the j-th position belonging to the corresponding structure given the r-th support factor, is the value vector of the j-th position in the sequence, is the global sparse feature set of the i-th position, is the indicator function, used to mark whether the j-th position is within the scope of the global sparse.
6. A transformer voiceprint fault recognition system, characterized in that, The method comprises the following steps: an obtaining module configured to obtain a transformer voiceprint signal, and preprocess the transformer voiceprint signal to obtain a target transformer voiceprint signal, wherein the target transformer voiceprint signal comprises at least one transformer voiceprint sub-signal frame; The extraction module is configured to extract a hybrid filter power spectrum feature vector in the target transformer voiceprint signal according to an improved hybrid filter bank based on an asymmetric time-varying overlapping wavelet packet decomposition, wherein an expression of the hybrid filter power spectrum feature vector is: , , , , , wherein, is the mixed filter power, is the joint feature pair, is the mel all-band mixed filter bank output power, is the energy feature, is the entropy feature, is the real number domain where the joint feature vector resides, is the adaptive weight, is the total number of samples of the sub-band time series, is the frequency domain power spectrum of the jth sub-band in the bth frequency band, is the salience index of the jth sub-band in the bth frequency band, is the slope adjustment parameter of the Sigmoid function, is the power proportion at frequency f, is the frequency domain signal of the jth sub-band in the bth frequency band, is the mean of is the total number of low frequency features, is the low frequency key band filter bank output power, is the total number of high frequency features, is the high frequency band mel filter bank output power, is the start frequency, is the frequency band overlap width, is the end frequency, is the input frequency; The splicing module is configured to fuse the hybrid filter power spectrum feature vector and a preset multi-dimensional feature vector according to a preset multi-granularity dynamic multi-dimensional feature fusion algorithm, to obtain an MDF vector of the at least one frame transformer voiceprint sub-signal, and to splice the MDF vectors of the respective frame transformer voiceprint sub-signals to generate a fusion feature vector corresponding to the target transformer voiceprint signal. The construction module is configured to construct a transformer voiceprint recognition model based on a structured factor full-span sparse attention mechanism. The output module is configured to input the fusion feature vector into the transformer voiceprint recognition model, and the transformer voiceprint recognition model outputs a fault recognition result corresponding to the transformer voiceprint signal.
7. An electronic device, comprising: The program is executed by the processor to implement the method of any one of claims 1 to 5. The program is executed by the processor to implement the method of any one of claims 1 to 5.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that,
Citation Information
Patent Citations
Fault automatic detection and repair method for self-healing intelligent power line
CN118739184A
Transformer partial discharge defect voiceprint feature extraction method and system
CN119517085A