Transformer abnormity identification method based on voiceprint feature analysis

By deploying acoustic sensors on the transformer to collect acoustic signals, and combining noise recognition and denoising technologies, and using convolutional neural networks for feature extraction and diagnosis, the problem of online monitoring of transformer core loosening faults has been solved, achieving high-precision non-contact diagnosis, which is suitable for fault identification in complex environments.

CN121542580APending Publication Date: 2026-02-17YANCHENG POWER SUPPLY CO STATE GRID JIANGSU ELECTRIC POWER CO
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511543045.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-27
Publication Date
2026-02-17

AI Technical Summary

Technical Problem

Existing transformer fault diagnosis technologies struggle to balance diagnostic accuracy and convenience when faced with environmental noise interference. In particular, online monitoring and diagnosis of core loosening faults suffer from poor real-time performance and insufficient anti-interference capabilities.

Method used

A method based on acoustic signature analysis is adopted. Acoustic signals are collected by deploying acoustic sensors at different locations on the transformer. Combined with noise recognition and adaptive noise reduction technology, a hierarchical convolutional neural network is used for signal processing and fault diagnosis, including density peak clustering and CEEMDAN-wavelet thresholding for noise reduction. Combined with MFCC, energy spectrum and spectrogram feature extraction, 2D-CNN and 3D-CNN are used for recognition.

Benefits of technology

It achieves high-precision, non-contact diagnosis of transformer core loosening faults under complex operating conditions, improves the signal-to-noise ratio and identification accuracy, and is suitable for online monitoring of substations and other electromechanical equipment. It has good versatility and engineering application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121542580A_ABST
    Figure CN121542580A_ABST
Patent Text Reader

Abstract

The invention discloses a transformer abnormity identification method based on voiceprint feature analysis, and belongs to the field of power equipment state monitoring and intelligent diagnosis. The method comprises the following steps: firstly, analyzing an iron core acoustic mechanism based on a magnetostrictive effect, and establishing a three-dimensional model through finite element simulation to obtain vibration and sound field characteristics; in a complex substation environment, a hybrid noise reduction method combining density peak clustering and a CEEMDAN-wavelet threshold is provided, and the signal-to-noise ratio is effectively improved. Then extracting Mel-frequency cepstrum coefficients (MFCC) and spectrum features, and performing local linear embedding (LLE) dimension reduction to form a compact feature set; in the recognition stage, a convolutional neural network framework is designed, specifically, a spectrogram and an energy spectrum are modeled through a two-dimensional CNN, an MFCC tensor obtained after dimensionality reduction is modeled through a three-dimensional CNN, and accurate diagnosis of mechanical faults such as core looseness is achieved. The method has the advantages of being non-contact, anti-noise and high in recognition precision, and real-time diagnosis and early warning of mechanical abnormity of the transformer can be achieved under complex working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system equipment condition monitoring and intelligent fault diagnosis technology, specifically to a transformer anomaly identification method based on acoustic signature analysis. Background Technology

[0002] With the improvement of social productivity and the continuous growth of electricity demand, the scale of power systems is constantly expanding, placing higher demands on the safety and reliability of equipment operation. As the core equipment of the power system, transformers play a crucial role in the power transmission and distribution process. However, in recent years, the failure rate of transformers has shown an upward trend. Among them, mechanical abnormal defects such as core loosening, if not detected and dealt with in a timely manner, can easily lead to more serious equipment damage and power grid accidents. Therefore, researching efficient and reliable transformer condition monitoring and fault diagnosis methods has significant engineering implications.

[0003] Currently, domestic and international research methods for transformer fault diagnosis mainly include oil chromatography, vibration detection, ultrasonic detection, and preliminary diagnostic methods based on acoustic signatures. Oil chromatography analyzes the composition and concentration of dissolved gases in transformer oil to identify potential electrical faults, enabling early detection of abnormalities such as discharge and overheating. However, it is an offline detection method with poor real-time performance and insufficient sensitivity to mechanical faults such as core loosening and winding loosening. Vibration detection uses sensors placed on the transformer casing to collect vibration signals and performs spectral characteristic analysis to diagnose core or winding abnormalities. While it has achieved some success in core loosening research, this method requires direct contact with the equipment, is susceptible to electromagnetic interference, and has complex measurement point arrangements; large transformers require multiple sensors, leading to high costs and poor stability. Ultrasonic detection is mainly used to detect electrical defects such as partial discharge and air bubbles in the oil, but it is not sensitive to mechanical faults such as core loosening, limiting its application scope. In recent years, researchers have begun to explore the use of voiceprint signals for diagnosis. By extracting Mel-frequency cepstral coefficients (MFCC), time-frequency features, or i-vector features, and combining them with support vector machines (SVM) and deep neural networks (DNN), this method has shown some effectiveness in experimental environments. However, most of these methods are based on quiet conditions and lack noise processing mechanisms for complex substation environments. The acoustic signal-to-noise ratio is low under noise interference from fans, switches, etc., feature extraction is unstable, and the model does not make sufficient use of the spatiotemporal distribution of voiceprint features, resulting in poor robustness.

[0004] Existing methods generally suffer from poor real-time performance, complex deployment, insufficient anti-interference capabilities, and limited identification accuracy and robustness, making it difficult to meet the online monitoring and diagnosis needs of transformer mechanical faults in actual power grid operation. Therefore, there is an urgent need for a non-contact transformer fault diagnosis technology with good noise immunity and high accuracy and robustness under complex operating conditions. Based on this, this invention proposes a transformer anomaly identification method based on acoustic signature analysis. This method can acquire acoustic signals in real time without affecting normal equipment operation, improve the signal-to-noise ratio through noise identification and adaptive noise reduction, and then combine it with a deep learning model to achieve high-precision diagnosis of mechanical faults such as core loosening, providing a new approach for the safe and stable operation of power systems. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing transformer core loosening fault diagnosis technology, such as "large environmental noise interference and difficulty in balancing diagnostic accuracy and convenience", and to provide a transformer core loosening fault diagnosis device and method based on acoustic print. By accurately identifying and removing environmental noise and constructing a hierarchical convolutional neural network diagnostic model, efficient and accurate diagnosis of transformer core loosening faults can be achieved.

[0006] This invention proposes a transformer anomaly identification method based on voiceprint feature analysis, which is implemented through the following technical solution: Step (1) Acquire acoustic signals from the transformer, deploy acoustic sensors at different locations on the transformer, establish a physical model, and collect acoustic signals generated by core vibration under no-load and different operating conditions. Step (2) performs normalization preprocessing on the signals acquired in step (1), and combines it with the time-frequency density peak clustering method for noise identification. Then, the identified noise pollution signals are processed. Denoising processing based on CEEMDAN-wavelet thresholding method is performed. Step (3) involves extracting acoustic features from the samples identified as noisy in step (2), performing framing, windowing, and spectral analysis on the signal, and extracting features such as MFCC, energy spectrum, and spectrogram. Then, LLE is used for dimensionality reduction and a dataset is created. Step (4) construct recognition inputs for the obtained feature datasets according to different needs, input the extracted time-frequency feature two-dimensional images into 2D-CNN for recognition; input the feature datasets obtained in step (3) into 3D-CNN for recognition, and then train and optimize the parameters of the convolutional neural network until the model converges and anomaly recognition is completed. Finally, step (5) combines the model recognition results from step (4) to determine whether there is any loosening abnormality in the transformer core and outputs the corresponding loosening level to achieve real-time diagnosis and early warning of the operating status.

[0007] Further, step (1) includes: Step (11) Multiple acoustic sensors are arranged on the side, top surface, and front of the oil tank to collect acoustic signals. Preferably, 2-4 sensors are evenly distributed on the side with the largest sound pressure amplitude to ensure multi-directional capture of acoustic signature information generated by core vibration. The position coordinates of each sensor are recorded as follows:

[0008] in, This is the sensor's serial number identifier, with a value range of [value range missing]. , , This represents the coordinate components of the acoustic sensor in a three-dimensional Cartesian coordinate system. Step (12) After the sensor deployment is completed, a physical model needs to be established to simulate, verify, and correct the acoustic signal propagation mechanism. This invention utilizes the multiphysics finite element method to establish a three-dimensional transformer model, which includes main components such as the core, windings, tank, and air domain. The core material is set as lossless soft iron, and a density is assigned... Young's modulus and magnetization curve Features: The winding part is made of copper with high conductivity; the air domain boundary is set as a sound-absorbing boundary condition to approximate a free sound field. In order to balance computational efficiency and simulation accuracy, the core region adopts a fine mesh, while the oil tank and air domain adopt a relatively coarse mesh. Step (13) After establishing the model, it is necessary to simulate the electromagnetic excitation during actual operation. In this invention, under no-load operation, a three-phase symmetrical AC voltage is applied to the primary winding as the excitation source, and its expression is:

[0009] in , , These are three-phase voltage symbols, representing the instantaneous voltages across the windings of phases A, B, and C, respectively, and are time-dependent. t The function, Rated voltage, The frequency is the power supply frequency. This excitation generates an alternating magnetic field distribution in the iron core, and further induces periodic mechanical vibrations due to the magnetostrictive effect, thereby forming a perceptible acoustic signal; Finally, in step (14), the displacement of the iron core surface caused by electromagnetic excitation is... and acceleration As an acoustic source term, it is coupled into the sound field control equation to solve for the sound pressure distribution:

[0010] in, Sound pressure level, represented in spatial coordinates place, time The amount of air pressure fluctuation at that time; The displacement of the iron core surface represents the displacement of a point on the iron core surface. At any moment t The amount of vibration displacement; air density, The speed of sound.

[0011] Further, step (2) includes: Step (21) involves collecting the acoustic signals. Frame processing is performed using a fixed frame length to avoid the impact of non-stationary features in long signals on analysis accuracy. The length of each frame is selected as follows: Each frame has 50 ms sampling points (corresponding to 50 kHz sampling frequency), and 20%–50% overlap is set between frames to ensure a smooth transition. To reduce spectral leakage, each frame signal is multiplied by a Hamming window, the expression of which is:

[0012] in, The window function is in the first position. The values ​​of each sampling point, Frame length; Step (22) extracts four typical time-domain features from the windowed signal frame to characterize the signal's energy magnitude, volatility, and impulsiveness. The specific definitions are as follows: Step (221) Root Mean Square Reflects the overall energy level of the signal.

[0013] in, For the signal at the 1st The amplitude of each sampling point The number of sampling points per frame; Step (222) Variance : Measures the degree of fluctuation between the signal and the mean.

[0014] in, This is the mean value of the signal in that frame; Step (223) Kurtosis : Reflects the degree of signal spikes, often used to characterize impulsiveness and non-Gaussianness.

[0015] in, This represents the standard deviation of the signal. A kurtosis value significantly greater than 3 indicates that the signal may contain impulse or abrupt changes. Step (224) Skewness (Skewness): Describes the symmetry of signal distribution.

[0016] If the skewness is positive, the signal distribution is biased to the left; if the skewness is negative, the distribution is biased to the right. Step (23) Based on the time-domain features, it is necessary to further extract the frequency-domain energy features to reveal the energy distribution pattern of the signal in different frequency bands. The specific method is to perform... Layered wavelet packet decomposition can decompose the original signal. Decomposed into The reconstruction formula for each sub-band signal is:

[0017] in, Indicates the first Layer Each sub-band signal component; corresponding band energy The calculation formula is:

[0018] in, Indicates the first The sub-band is in the first The value of each sampling point, The total number of sampling points; to facilitate comparison of the contributions of different frequency bands, we introduce... Energy ratio:

[0019] in, This allows us to determine if there is an abnormally high concentration of energy in a specific frequency band, thus aiding in noise identification. Step (24) involves extracting the time-domain and frequency-domain features, combining them into a feature vector matrix, and then inputting it into the Density Peaks Clustering (DPC) algorithm for noise identification. This algorithm uses graphs to draw... Decision graphs are used to visually determine cluster centers, using points that are "locally high in density and far from other high-density points" as cluster centers, thereby distinguishing effective signals from abnormal noise; (1) First, let the sample points be set. The feature vector is Sample points The feature vector is Then the sample points and Euclidean distance The expression is:

[0020] in The dimension of the feature vector represents the number of features contained in each sample point. In this invention, four features—root mean square, variance, kurtosis, and skewness—were extracted during the time-domain feature extraction of the transformer acoustic signal. Therefore, the dimension of the feature vector for each sample is... ; m An index for a single feature; (2) Next, calculate the local density of each sample point. This is used to measure the degree of clustering of a point in the feature space. Its mathematical definition is:

[0021] in, This is to cut off the distance. Therefore, we can conclude that... The larger the value, the more neighbors the point has, and the higher the local density. (3) Subsequently, define the sample points. minimum distance Its representation point With all densities greater than The closest distance between sample points:

[0022] (4) If a point has the highest local density globally, i.e., there is no point with a higher density, then its minimum distance is... Defined as the maximum distance between all sample points;

[0023] (5) Furthermore, in order to select cluster centers more intuitively, this invention introduces a comprehensive index. Defined as:

[0024] in, The value of reflects both the "density advantage" and "distance advantage" of the sample points. When a point simultaneously has a high density advantage... and At that time, its If the value is also relatively large, this point is determined to be the class center. Finally, by drawing... Decision chart, combined with comprehensive indicators It can identify points with high local density and large distance as cluster centers of effective signals, while points with low local density but large distance are identified as abnormal noise points, thereby realizing the automatic identification and removal of noisy samples. Step (25) in the signal Superimposed Gaussian white noise , and repeat The mean of the first-order IMF components obtained from the EMD decomposition is used as the final IMF:

[0025] in, Indicates the first The first-order IMF obtained from the decomposition. This represents the final first-order IMF. Then, EMD decomposition is performed. Finally, the residual components are defined:

[0026] For residual information Add noise again And decompose it to obtain a new IMF:

[0027] in, Indicates the first The EMD operation is performed once. The decomposition is iterated until the remaining signal is obtained. No longer containing decomposable components, at this point:

[0028] In the formula, This represents the remaining quantity signal that cannot be further decomposed. Indicates the first Each intrinsic mode component The final number of IMF components obtained from the decomposition; Step (26) Calculate the kurtosis value for each of the IMF components obtained from the decomposition. Kurtosis is used to measure the peaks and impulsiveness of a signal. According to experimental results, normal transformer acoustic signature signals are relatively stable, with kurtosis generally in the range of 2.6–3.0; while environmental noise often exhibits impulsive characteristics, with kurtosis values ​​much greater than 3. When the kurtosis of a certain IMF component significantly deviates from 3, the component can be determined to be a noisy component and enter the wavelet threshold denoising process. Step (27) involves wavelet decomposition of the IMF components identified as noise to obtain wavelet coefficients. Set threshold And it is processed using a soft threshold function:

[0029] Where w represents the processed wavelet coefficients, and λ is the threshold. This method retains the large amplitude coefficients (corresponding to the effective signal) and removes the small amplitude coefficients (corresponding to noise), thereby achieving denoising of the IMF components; Step (28) involves re-superimposing the IMF components after wavelet threshold denoising with the normal IMF components to obtain the final denoised signal. :

[0030] To avoid ambiguity, the summation index is used separately. and ;in, Let S represent the total number of IMF components obtained from EMD decomposition, and let S represent the IMF components that undergo wavelet threshold denoising, while R represent the IMF components that do not require processing. Indicates the first The result after applying the denoising operator to each IMF; Step (29) introduces the signal-to-noise ratio (SNR) metric to quantitatively evaluate the noise reduction effect:

[0031] in, Indicates the effective signal power. This represents the noise power. By comparing the changes in SNR before and after noise reduction, the noise reduction effect can be intuitively reflected.

[0032] Further, step (3) includes: Step (31) First, the denoised signal is divided into frames and windowed (50 ms per frame, 50% overlap, Hamming window), and then a Discrete Fourier Transform (DFT) is performed on each frame to obtain the time-frequency matrix. Used to describe the energy distribution of a signal in time and frequency, it can draw a spectrogram, thus visually displaying significant peaks in the low-frequency range (such as 100 Hz, 200 Hz, and 300 Hz), reflecting the operating characteristics of the transformer. The specific formula for the time-frequency matrix is ​​as follows:

[0033] in, For the first The signal amplitude at each sampling point The total number of sampling points per frame. For the first The complex amplitude of each frequency component, Represents the imaginary unit; Step (32) introduces the energy spectrum as a key feature in the frequency domain features. After performing a Fourier transform on each frame of the signal, the energy distribution is calculated:

[0034] in, For the first Complex amplitude at a frequency point The energy at the corresponding frequency point; Step (33) will generate the energy spectrum pass Filter bank processing yields the energy output of each filter. Then, the logarithm of the filter energy is taken to obtain the logarithmic spectral energy. To enhance the nonlinear characteristics of human hearing perception, a discrete cosine transform (DCT) is then performed on the logarithmic energy to obtain the MFCC coefficients.

[0035] Where M is the number of Mel filters (usually 26). For the first cepstral coefficients, This is the order of the cepstral coefficients (usually taken as 13). To enhance timing information, the first-order difference of MFCC can also be calculated. and second-order difference Finally, the MFCC feature matrix is ​​obtained:

[0036] in, The feature matrix is ​​a real number matrix with 19 rows and 39 columns. The 19 rows correspond to 19 consecutive signal frames (each frame is 50ms, with an overlap rate of 50%, covering a signal segment of about 0.95 seconds), and the 39 columns correspond to the feature dimensions of each frame (13 static + 13 first-order difference + 13 second-order difference). Step (34): Due to the high feature dimension of MFCC, in order to reduce redundancy and improve computational efficiency, the Local Linear Embedding (LLE) method is used for dimensionality reduction. The processing flow is as follows: Step (341) assumes the MFCC feature sample set is Calculate the Euclidean distance for each sample and find The nearest neighbors:

[0037] in, For the signal at the 1st The amplitude of each sampling point The number of sampling points per frame. The dimension of the original feature space. The index variable is a feature dimension variable, and its value range is... ; Representing samples respectively and In the p Feature values ​​in each dimension; Step (342) minimizes the reconstruction error for each sample.

[0038]

[0039] in, This represents the total number of samples, i.e., the number of sample points contained in the MFCC feature set. The number of nearest neighbors for each sample. Weight matrix , Indicates the first The nearest neighbor sample pair of the nth The contribution weight of each sample, Indicates the first The first sample One nearest neighbor sample; Step (343) retains the weight matrix obtained by solving formula (29). Without changing the dimensions, the high-dimensional sample data is mapped to a low-dimensional space. The output vector is obtained. :

[0040] Before selection by feature decomposition We obtain the dimensionality-reduced MFCC features from the eigenvectors. The sample set after mapping to a lower-dimensional space. It is the first i The vector representation of each sample in the low-dimensional space is the output feature after dimensionality reduction; It is the first The first sample The vector representation of each nearest neighbor sample in a low-dimensional space.

[0041] Further, step (4) includes: Step (41) For two-dimensional static time-frequency feature matrices such as spectrograms (dimension [465×40]) and energy spectra (dimension [19×8], 8 frequency bands after 3-layer wavelet packet decomposition), 2D-CNN is used for local feature extraction, focusing on capturing the differences in frequency component distribution and the features of energy concentration areas. Its core operation is two-dimensional convolution operation:

[0042] in, Indicates the first Layer Each output feature map represents an abstract representation of the local time-frequency pattern by that layer. Indicates the relationship with the first A set of input feature maps connected together; Indicates the first Layer connection input feature map With output feature map A two-dimensional convolutional kernel (3×3 size, stride 1, with zero padding at the edges to maintain the feature map size) is used to extract local time-frequency texture. Indicates the first Layer The bias term of each feature map is used to adjust the overall offset of the feature map. This indicates that the activation function is ReLU, and the formula is: This is to alleviate the gradient vanishing problem and enhance the model's nonlinear expressive power. Step (42) involves compressing the data size, reducing computational complexity, and enhancing the model's robustness to sensor position shifts. A max-pooling layer (2×2 kernel size, stride 2) is added after every two convolutional layers to downsample local features. After four convolutional layers (channel counts change from 64 to 128 to 256 to 512) and two pooling layers, the extracted deep feature map is flattened into a one-dimensional vector (e.g., [512×5×58] flattened to 14848 dimensions). This vector is then input into two fully connected layers (256 to 64 nodes) for feature integration, outputting a 2D-CNN feature vector. F 2 D (Dimension 64); Step (43) involves extracting spatiotemporal joint features from the LLE-reduced MFCC features (including first-order and second-order difference coefficients) into a three-dimensional tensor feature (dimension [4×19×18×1], where 4 is the number of time segments, 19 is the number of frames, 18 is the frequency dimension after dimensionality reduction, and 1 is the number of channels). The focus is on capturing the dynamic changes of frequency components within different time segments. The convolution operation formula is as follows:

[0043] Indicates the first Layer output feature tensor at position The value at that location, Corresponding time segmentation dimensions Corresponding frame time dimension Corresponding frequency dimension, The weights of the l-th layer 3D convolutional kernel, P, Q, and R are the size of the convolutional kernel in time, frequency, and channel, respectively, and m is the input channel index of the previous layer; For the first Global bias terms for the layer; Step (44) finally inputs the obtained feature vector into the classification module (containing one fully connected layer and a Softmax activation function) to classify the degree of core loosening. In the fully connected layer, the 1088-dimensional fused feature is mapped into a 4-dimensional vector (corresponding to 4 types of fault states: not loose, loose 40%, loose 80%, loose 100%), and the original scores of each category are output. Then, the original scores are transformed into a probability distribution using the Softmax activation function, as shown in the formula:

[0044] Where C=4 is the number of categories. This indicates that the input sample belongs to the first... The probability of each type of fault state is calculated, and the class with the highest probability is ultimately selected as the model output label.

[0045] Further, step (6) includes: based on the Softmax output probability distribution obtained in step (61), selecting the category with the highest probability as the prediction result, thereby determining whether there is any loosening abnormality in the transformer core, and distinguishing different levels (not loose, 40% loose, 80% loose, 100% loose). Step (62) Finally, based on the final loosening level, a real-time diagnostic report is generated, which includes status identifiers, feature details, and early warning suggestions.

[0046] Compared with the prior art, the present invention has the following significant advantages: (1) The present invention can realize multi-directional acoustic signal acquisition by arranging a small number of acoustic sensors on the surface of the transformer tank. It does not require direct contact with the iron core or winding, thus avoiding the complex layout and electromagnetic interference problems in the traditional vibration detection method, and greatly improving the ease of installation and monitoring reliability. (2) In response to background noise such as fans, human voices, and vehicles in the complex environment of substations, this invention introduces a hybrid noise reduction method based on density peak clustering (DPC) and CEEMDAN wavelet threshold, which effectively eliminates environmental interference and significantly improves the signal-to-noise ratio (SNR) of the voiceprint signal, providing high-quality input for subsequent feature extraction. (3) This invention combines multiple features such as spectrogram, energy spectrum and MFCC to comprehensively characterize the time-frequency distribution and cepstral features of voiceprints, and achieves dimensionality reduction optimization through the local linear embedding (LLE) method, which not only preserves key difference information, but also reduces redundancy and improves feature compactness and computational efficiency. (4) This invention designs recognition methods based on both 2D-CNN and 3D-CNN. 2D-CNN is suitable for modeling two-dimensional time-frequency features such as spectrograms and energy spectra, and has high computational efficiency; 3D-CNN is designed for MFCC tensors after LLE dimensionality reduction, and can characterize spatiotemporal evolution features, resulting in more refined recognition. The two methods have their own focuses and can be flexibly selected or combined according to actual working conditions and computing resources, so as to achieve efficient and reliable fault diagnosis in different application scenarios; (5) The method of the present invention relies on acoustic signal analysis and belongs to the non-contact acoustic monitoring framework. It is not only applicable to online monitoring of substations and large power transformers, but can also be extended to non-contact fault diagnosis scenarios of other electromechanical equipment, such as large motors, compressors and rotating machinery. It has good versatility and engineering application value. Attached Figure Description

[0047] Figure 1 This is a flowchart of a transformer anomaly identification method based on voiceprint feature analysis proposed in this invention; Figure 2This invention provides a three-dimensional simulation diagram of sensor layout and transformer core-winding; Figure 3 The present invention proposes a signal processing flow. Detailed Implementation

[0048] The following detailed description, in conjunction with the accompanying drawings, illustrates a specific implementation method for transformer anomaly identification based on voiceprint feature analysis according to the present invention.

[0049] The present invention adopts the following technical solution: In order to complete the task of real-time diagnosis and early warning of transformer core loosening abnormality, the present invention proposes a transformer abnormality identification method based on acoustic feature analysis.

[0050] Please refer to Figure 1 This invention provides a transformer anomaly identification method based on voiceprint feature analysis, the specific process of which is as follows: Please refer to step (1). Figure 2 Acoustic signals of the transformer are collected by deploying acoustic sensors at different locations on the transformer, establishing a physical model, and collecting acoustic signals generated by core vibration under no-load and different operating conditions. Step (2) performs normalization preprocessing on the signals acquired in step (1), and combines it with the time-frequency density peak clustering method for noise identification. Then, the identified noise pollution signals are processed. Denoising processing based on CEEMDAN-wavelet thresholding method is performed. Step (3) involves extracting acoustic features from the samples identified as noisy in step (2), performing framing, windowing, and spectral analysis on the signal, and extracting features such as MFCC, energy spectrum, and spectrogram. Then, LLE is used for dimensionality reduction and a dataset is created. Step (4) constructs recognition inputs from the obtained feature datasets according to different needs, and inputs the extracted time-frequency feature 2D images into a 2D-CNN for recognition; please refer to Figure 3 The feature dataset obtained in step (3) is input into 3D-CNN for recognition. Then, the convolutional neural network is trained and its parameters are optimized until the model converges and anomaly recognition is completed. Finally, step (5) combines the model recognition results from step (4) to determine whether there is any loosening abnormality in the transformer core and outputs the corresponding loosening level to achieve real-time diagnosis and early warning of the operating status.

[0051] Further, step (1) includes: Step (11) Multiple acoustic sensors are arranged on the side, top surface, and front of the oil tank to collect acoustic signals. Preferably, 2-4 sensors are evenly distributed on the side with the largest sound pressure amplitude to ensure multi-directional capture of acoustic signature information generated by core vibration. The position coordinates of each sensor are recorded as follows:

[0052] in, This is the sensor's serial number identifier, with a value range of [value range missing]. (positive integer), For the number of sensors, This represents the coordinate components of the acoustic sensor in a three-dimensional Cartesian coordinate system. Step (12) After the sensor deployment is completed, a physical model needs to be established to simulate, verify, and correct the acoustic signal propagation mechanism. This invention utilizes the multiphysics finite element method to establish a three-dimensional transformer model, which includes main components such as the core, windings, tank, and air domain. The core material is set as lossless soft iron, and a density is assigned... Young's modulus and magnetization curve Features: The winding part is made of copper with high conductivity; the air domain boundary is set as a sound-absorbing boundary condition to approximate a free sound field. In order to balance computational efficiency and simulation accuracy, the core region adopts a fine mesh, while the oil tank and air domain adopt a relatively coarse mesh. Step (13) After establishing the model, it is necessary to simulate the electromagnetic excitation during actual operation. In this invention, under no-load operation, a three-phase symmetrical AC voltage is applied to the primary winding as the excitation source, and its expression is:

[0053] in , , These are three-phase voltage symbols, representing the instantaneous voltages across the windings of phases A, B, and C, respectively, and are time-dependent. t The function, Rated voltage, The frequency is the power supply frequency. This excitation generates an alternating magnetic field distribution in the iron core, and further induces periodic mechanical vibrations due to the magnetostrictive effect, thereby forming a perceptible acoustic signal; A4) Finally, the displacement of the iron core surface caused by electromagnetic excitation and acceleration As an acoustic source term, it is coupled into the sound field control equation to solve for the sound pressure distribution:

[0054] in, Sound pressure level, represented in spatial coordinates place, time The amount of air pressure fluctuation at that time; The displacement of the iron core surface represents the displacement of a point on the iron core surface. At any moment t The amount of vibration displacement; air density, The speed of sound.

[0055] Further, step (2) includes: Step (21) involves collecting the acoustic signals. Frame processing is performed using a fixed frame length to avoid the impact of non-stationary features in long signals on analysis accuracy. The length of each frame is selected as follows: Each frame has 50 ms sampling points (corresponding to 50 kHz sampling frequency), and 20%–50% overlap is set between frames to ensure a smooth transition. To reduce spectral leakage, each frame signal is multiplied by a Hamming window, the expression of which is:

[0056] in, The window function is in the first position. The values ​​of each sampling point Frame length; Step (22) extracts four typical time-domain features from the windowed signal frame to characterize the signal's energy magnitude, volatility, and impulsiveness. The specific definitions are as follows: (1) Root mean square Reflects the overall energy level of the signal.

[0057] in, For the signal at the 1st The amplitude of each sampling point The number of sampling points per frame; (2) Variance : Measures the degree of fluctuation between the signal and the mean.

[0058] in, This is the mean value of the signal in that frame; (3) kurtosis : Reflects the degree of signal spikes, often used to characterize impulsiveness and non-Gaussianness.

[0059] in, This represents the standard deviation of the signal. A kurtosis value significantly greater than 3 indicates that the signal may contain impulse or abrupt changes. (4) Skewness : Describes the symmetry of signal distribution,

[0060] If the skewness is positive, the signal distribution is biased to the left; if the skewness is negative, the distribution is biased to the right. Step (23) Based on the time-domain features, it is necessary to further extract the frequency-domain energy features to reveal the energy distribution pattern of the signal in different frequency bands. The specific method is to perform... Layered wavelet packet decomposition can decompose the original signal. Decomposed into The reconstruction formula for each sub-band signal is:

[0061] in, Indicates the first Layer Each sub-band signal component; corresponding band energy The calculation formula is:

[0062] in, Indicates the first The sub-band is in the first The value of each sampling point, The total number of sampling points; to facilitate comparison of the contributions of different frequency bands, we introduce... Energy ratio:

[0063] in, This allows us to determine if there is an abnormally high concentration of energy in a specific frequency band, thus aiding in noise identification. Step (24) involves extracting the time-domain and frequency-domain features, combining them into a feature vector matrix, and then inputting it into the Density Peaks Clustering (DPC) algorithm for noise identification. This algorithm uses graphs to draw... Decision graphs are used to visually determine cluster centers, using points that are "locally high in density and far from other high-density points" as cluster centers, thereby distinguishing effective signals from abnormal noise; Step (241) First, set the sample points The feature vector is Sample points The feature vector is Then the sample points and Euclidean distance The expression is:

[0064] in The dimension of the feature vector represents the number of features contained in each sample point. In this invention, four features—root mean square, variance, kurtosis, and skewness—were extracted during the time-domain feature extraction of the transformer acoustic signal. Therefore, the dimension of the feature vector for each sample is... ; m An index for a single feature; Step (242) then calculates the local density for each sample point. This is used to measure the degree of clustering of a point in the feature space. Its mathematical definition is:

[0065] in, This is to cut off the distance. Therefore, we can conclude that... The larger the value, the more neighbors the point has, and the higher the local density. Step (243) then defines the sample points. minimum distance Its representation point With all densities greater than The closest distance between sample points:

[0066] Step (244): If a point has the highest local density globally, i.e., there is no point with a higher density, then its minimum distance is... Defined as the maximum distance between all sample points;

[0067] Step (245) Further, in order to select cluster centers more intuitively, this invention introduces a comprehensive index. Defined as:

[0068] in, The value of reflects both the "density advantage" and "distance advantage" of the sample points. When a point simultaneously has a high density advantage... and At that time, its If the value is also relatively large, this point is determined to be the class center. Finally, by drawing... Decision chart, combined with comprehensive indicators It can identify points with high local density and large distance as cluster centers of effective signals, while points with low local density but large distance are identified as abnormal noise points, thereby realizing the automatic identification and removal of noisy samples. Step (25) in the signal Superimposed Gaussian white noise , and repeat The mean of the first-order IMF components obtained from the EMD decomposition is used as the final IMF:

[0069] in, Indicates the first k The first-order IMF obtained from the decomposition. This represents the final first-order IMF. Then, EMD decomposition is performed. Finally, the residual components are defined:

[0070] For residual information Add noise again And decompose it to obtain a new IMF:

[0071] in, Indicates the first The EMD operation is performed once. The decomposition is iterated until the remaining signal is obtained. No longer containing decomposable components, at this point:

[0072] In the formula, This represents the residual signal that cannot be further decomposed. Indicates the first Each intrinsic mode component The final number of IMF components obtained from the decomposition; Step (26) Calculate the kurtosis value for each of the IMF components obtained from the decomposition. Kujic is used to measure the peaks and impulsiveness of a signal. According to experimental results, normal transformer acoustic signature signals are relatively stable, with a kujicity generally in the range of 2.6–3.0; while environmental noise often exhibits impulsive characteristics, with kujicity values ​​much greater than 3. When the kujicity of a certain IMF component significantly deviates from 3, the component can be determined to be a noisy component and enter the wavelet threshold denoising process. Step (27) involves performing wavelet decomposition on the IMF components identified as noise to obtain wavelet coefficients. Set threshold And it is processed using a soft threshold function:

[0073] Where w represents the processed wavelet coefficients, The threshold is used. This method retains the large amplitude coefficient (corresponding to the effective signal) and removes the small amplitude coefficient (corresponding to noise), thereby achieving denoising of the IMF component; Step (28) involves re-superimposing the IMF components after wavelet threshold denoising with the normal IMF components to obtain the final denoised signal. :

[0074] To avoid ambiguity, the summation index is used separately. and ;in, Let S represent the total number of IMF components obtained from EMD decomposition, and let S represent the IMF components that undergo wavelet threshold denoising, while R represent the IMF components that do not require processing. Indicates the first The result after applying the denoising operator to each IMF; Step (29) introduces the signal-to-noise ratio (SNR) metric to quantitatively evaluate the noise reduction effect:

[0075] in, Indicates the effective signal power. This represents the noise power. By comparing the changes in SNR before and after noise reduction, the noise reduction effect can be intuitively reflected.

[0076] Further, step (3) includes: Step (31) First, the denoised signal is divided into frames and windowed (50 ms per frame, 50% overlap, Hamming window), and then a Discrete Fourier Transform (DFT) is performed on each frame to obtain the time-frequency matrix. Used to describe the energy distribution of a signal in time and frequency, it can draw a spectrogram, thus visually displaying significant peaks in the low-frequency range (such as 100 Hz, 200 Hz, and 300 Hz), reflecting the operating characteristics of the transformer. The specific formula for the time-frequency matrix is ​​as follows:

[0077] in, For the first The signal amplitude at each sampling point The total number of sampling points per frame. For the first The complex amplitude of each frequency component, Represents the imaginary unit; Step (32) introduces the energy spectrum as a key feature in the frequency domain features. After performing a Fourier transform on each frame of the signal, the energy distribution is calculated:

[0078] in, For the first Complex amplitude at a frequency point, The energy at the corresponding frequency point; Step (33) will generate the energy spectrum pass Filter bank processing yields the energy output of each filter. Then, the logarithm of the filter energy is taken to obtain the logarithmic spectral energy. To enhance the nonlinear characteristics perceived by the human ear, the logarithmic energy is then subjected to a discrete cosine transform (DCT) to obtain the MFCC coefficients:

[0079] Where M is the number of Mel filters (usually 26). For the first cepstral coefficients, This is the order of the cepstral coefficients (usually taken as 13). To enhance timing information, the first-order difference of MFCC can also be calculated. and second-order difference Finally, the MFCC feature matrix is ​​obtained:

[0080] in, The feature matrix is ​​a real number matrix with 19 rows and 39 columns. The 19 rows correspond to 19 consecutive signal frames (each frame is 50ms, with an overlap rate of 50%, covering a signal segment of about 0.95 seconds), and the 39 columns correspond to the feature dimensions of each frame (13 static + 13 first-order difference + 13 second-order difference). Step (34): Due to the high feature dimension of MFCC, in order to reduce redundancy and improve computational efficiency, the Local Linear Embedding (LLE) method is used for dimensionality reduction. The processing flow is as follows: Step (341) assumes the MFCC feature sample set is Calculate the Euclidean distance for each sample and find The nearest neighbors:

[0081] in, For the signal at the 1st The amplitude of each sampling point The number of sampling points per frame. The dimension of the original feature space. The index variable is a feature dimension variable, and its value range is... ; Representing samples respectively and In the Feature values ​​in each dimension; Step (342) minimizes the reconstruction error for each sample.

[0082]

[0083] in, This represents the total number of samples, i.e., the number of sample points contained in the MFCC feature set. The number of nearest neighbors for each sample. Weight matrix , Indicates the first The nearest neighbor sample pair of the nth The contribution weight of each sample, Indicates the first The first sample One nearest neighbor sample; Step (343) retains the weight matrix obtained by solving formula (29). Without changing the dimensions, the high-dimensional sample data is mapped to a low-dimensional space. The output vector is obtained. :

[0084] Before selection by feature decomposition We obtain the dimensionality-reduced MFCC features from the eigenvectors. The sample set after mapping to a lower-dimensional space. It is the first The vector representation of each sample in the low-dimensional space is the output feature after dimensionality reduction; It is the first The first sample The vector representation of each nearest neighbor sample in a low-dimensional space.

[0085] Further, step (4) includes: Step (41) For two-dimensional static time-frequency feature matrices such as spectrograms (dimension [465×40]) and energy spectra (dimension [19×8], 8 frequency bands after 3-layer wavelet packet decomposition), 2D-CNN is used for local feature extraction, focusing on capturing the differences in frequency component distribution and the features of energy concentration areas. Its core operation is two-dimensional convolution operation:

[0086] in, Indicates the first Layer Each output feature map represents an abstract representation of the local time-frequency pattern by that layer. Indicates the relationship with the first A set of input feature maps connected together; Indicates the first Layer connection input feature map With output feature map A two-dimensional convolutional kernel (3×3 size, stride 1, with zero padding at the edges to maintain the feature map size) is used to extract local time-frequency texture. Indicates the first Layer The bias term of each feature map is used to adjust the overall offset of the feature map. This indicates that the activation function is ReLU, and the formula is: To alleviate the gradient vanishing problem and enhance the nonlinear expression capability of the model, step (42) is to compress the data scale, reduce the computational complexity and enhance the robustness of the model to sensor position shift. After every 2 convolutional layers, a max pooling layer (pooling kernel size 2×2, stride 2) is set to downsample the local features. After 4 convolutional layers (the number of channels is 64→128→256→512) and 2 pooling layers, the extracted deep feature map is flattened into a one-dimensional vector (e.g., [512×5×58] flattened into 14848 dimensions), input into 2 fully connected layers (number of nodes 256→64) for feature integration, and output 2D-CNN feature vector. F 2 D (Dimension 64); Step (43) involves extracting spatiotemporal joint features from the LLE-reduced MFCC features (including first-order and second-order difference coefficients) into a three-dimensional tensor feature (dimension [4×19×18×1], where 4 is the number of time segments, 19 is the number of frames, 18 is the frequency dimension after dimensionality reduction, and 1 is the number of channels). The focus is on capturing the dynamic changes of frequency components within different time segments. The convolution operation formula is as follows:

[0087] Indicates the first Layer output feature tensor at position The value at that location, Corresponding time segmentation dimensions Corresponding frame time dimension Corresponding frequency dimension, The weights of the l-th layer 3D convolutional kernel These represent the dimensions of the convolutional kernel in terms of time, frequency, and channels, respectively. For the first Global bias terms for the layer; Step (44) finally inputs the obtained feature vector into the classification module (containing one fully connected layer and a Softmax activation function) to classify the degree of core loosening. In the fully connected layer, the 1088-dimensional fused feature is mapped into a 4-dimensional vector (corresponding to 4 types of fault states: not loose, loose 40%, loose 80%, loose 100%), and the original scores of each category are output. Then, the original scores are transformed into a probability distribution using the Softmax activation function, as shown in the formula:

[0088] Where C=4 is the number of categories. This indicates that the input sample belongs to the first... The probability of each type of fault state is calculated, and the class with the highest probability is ultimately selected as the model output label.

[0089] Further, step (6) includes: Based on the Softmax output probability distribution obtained in step (61), the category with the highest probability is selected as the prediction result, thereby determining whether there is any loosening abnormality in the transformer core and distinguishing different levels (not loose, 40% loose, 80% loose, 100% loose). Step (62) Finally, based on the final loosening level, a real-time diagnostic report is generated, which includes status identifiers, feature details, and early warning suggestions.

[0090] Finally, it should be noted that the above embodiments are merely illustrative of the technical solutions of the present invention and not intended to limit it. Those skilled in the art should understand that modifications or equivalent substitutions can be made to the specific embodiments of the present invention, but such modifications or alterations are all within the scope of protection of the pending claims.

Claims

1. A transformer abnormality identification method based on voiceprint feature analysis, characterized in that, Comprise the following steps: Step (1) collecting the acoustic signal of the transformer, arranging acoustic sensors at different positions of the transformer, establishing a physical model, and collecting the acoustic signal generated by the core vibration under no-load and different operating states; Step (2) normalizes and pretreats the signal collected in step (1), and combines a time and frequency density peak value clustering method to identify noise, and then performs CEEMDAN-wavelet threshold method denoising processing on the identified noise contaminated signal Step (3) extracts acoustic features for the sample determined as containing noise in step (2), frames, windows and analyzes the spectrum of the signal, and extracts MFCC, energy spectrum and spectrogram features. Then dimension reduction is performed by using LLE and a data set is made; Step (4) the obtained feature data set is respectively constructed into an identification input according to different requirements, the extracted two-dimensional image of time-frequency features is input into 2D-CNN for identification, the feature data set obtained in step (3) is input into 3D-CNN for identification, then the convolutional neural network is trained and the parameters are optimized until the model converges and the abnormality recognition is completed; Step (5) finally, the model recognition results of step (4) are integrated to judge whether the transformer core has looseness abnormality, and the corresponding looseness degree grade is output, realizing real-time diagnosis and early warning of the operating state.

2. The transformer abnormality recognition method based on voiceprint feature analysis according to claim 1, characterized in that, Step (1) comprises: Step (11) arranging multiple acoustic sensors on the side, upper surface and front of the oil tank respectively to collect acoustic signals, and the position coordinates of each sensor are recorded as: wherein, is a serial number of the sensor, is a number of sensors, represents a coordinate component of the acoustic sensor in a three-dimensional rectangular coordinate system; Step (12) after the completion of sensor layout, the need to establish a physical model, in order to simulate the propagation mechanism of acoustic signal verification and correction; the use of multi-physical field finite element method to establish a three-dimensional transformer model, the model contains the core, winding, oil tank and air domain and other main components; wherein, the core material is set to lossless soft iron, and is endowed with density , Young's modulus and magnetization curve characteristics; winding part is set to high conductivity copper material; air domain boundary is set to sound absorbing boundary condition, in order to approximate the free sound field, in order to balance the calculation efficiency and simulation accuracy, the core area is divided by fine grid, and the oil tank and air domain are relatively coarse grid. Step (13) simulates the electromagnetic excitation in actual operation after the model is established; in the no-load operation state, three-phase symmetrical alternating voltage is applied to the primary winding as the excitation source, and its expression is: wherein , , is the three-phase voltage symbol, respectively representing the instantaneous voltage across the A-phase, B-phase, C-phase windings, is a function of time t , is the rated voltage, is the power supply frequency; Step (14) finally, the electromagnetic excitation caused by the core surface displacement and acceleration As an acoustic source term, coupled to the acoustic field control equation for solving the sound pressure distribution: wherein is the sound pressure, representing the air pressure fluctuation at a spatial coordinate at time ; is the core surface displacement, representing the vibration displacement of a point on the core surface at time ; t ; is the air density, is the sound speed.

3. The transformer abnormality recognition method based on voiceprint feature analysis according to claim 2, characterized in that, Step (2) comprises: Step (21) involves collecting the acoustic signals. Frame processing is performed using a fixed frame length to avoid the impact of non-stationary features in long signals on analysis accuracy; the length of each frame is selected... Each frame has 50 ms sampling points (corresponding to 50 kHz sampling frequency), and 20%–50% overlap is set between frames to ensure a smooth transition. To reduce spectral leakage, each frame signal is multiplied by a Hamming window, the expression of which is: wherein, denotes the value of the window function at the sample point, is the frame length; Step (22) four typical time-domain features are extracted from the windowed signal frame to represent the energy size, volatility and impact of the signal; Step (23) further extracts the frequency energy feature on the basis of the time domain feature to reveal the energy distribution law of the signal in different frequency bands. The specific method is to perform layer wavelet packet decomposition on the signal to decompose the original signal into sub-band signals, and the reconstruction formula is: ​ wherein, represents the layer the signal component of the sub-band; the corresponding band energy The calculation formula is: wherein, represents the value of the th sub-band at the th sample point, is the total number of sample points; for the purpose of comparing the contribution of different frequency bands, the energy ratio is introduced: ​ wherein ; Step (24) combines the time domain and frequency domain features into a feature vector matrix after extracting them, and inputs the density peak clustering algorithm for noise recognition; the algorithm visually determines the clustering center by drawing a decision graph, and uses a point with "high local density and far from other high density points" as the clustering center, thereby distinguishing effective signals from abnormal noise. Step (24) combines the time domain and frequency domain features into a feature vector matrix after extracting them, and inputs the density peak clustering algorithm for noise recognition; the algorithm visually determines the clustering center by drawing a decision graph, and uses a point with "high local density and far from other high density points" as the clustering center, thereby distinguishing effective signals from abnormal noise. Step (25) in the signal Superimposed Gaussian white noise , and repeat The mean of the first-order IMF components obtained from the EMD decomposition is used as the final IMF: wherein, represents the first order IMF obtained by the first decomposition, k represents the first order IMF obtained by the first decomposition, ​ residual signal again superimposed noise and decomposed, resulting in a new IMF: wherein, denotes the EMD operation; decomposition is iterated until the residual signal has no more decomposable components, at which point: wherein represents the non-decomposable residual signal, represents the first intrinsic mode component, is the number of IMF components finally decomposed. Step (26) calculates the kurtosis value of each IMF component ; wherein, the kurtosis is used to measure the spikiness and impact of the signal; when the kurtosis of a certain IMF component deviates from 3 significantly, it can be determined that the component contains noise, and enters the wavelet threshold denoising processing link; Step (27) wavelet decomposition is performed on the IMF component determined as noise to obtain wavelet coefficients , a threshold value is set , and a soft threshold function is used for processing: wherein w denotes the wavelet coefficients after processing, is a threshold value; Step (28) re-stacks the IMF component after wavelet threshold denoising processing with normal IMF component to obtain the final denoised signal : To avoid ambiguity, the summation indices are used with ; wherein, is the total number of IMF components obtained by EMD decomposition, set S represents the IMF components that enter the wavelet threshold denoising, set R represents the IMF components that do not need to be processed, represents the result after applying a denoising operator to the first IMF component. Step (29) is to quantitatively evaluate the noise reduction effect, and introduce the signal-to-noise ratio (SNR) index: wherein, represents the effective signal power, represents the noise power; by comparing the change of SNR before and after noise reduction, the noise reduction effect is intuitively reflected.

4. The transformer abnormality recognition method based on voiceprint feature analysis according to claim 3, characterized in that, Step (22) is specifically defined as follows: (1) root mean square : reflects the overall energy level of the signal, wherein, is the amplitude of the signal at the th sample point, is the number of sample points per frame; (2) variance : measures the degree of fluctuation between the signal and the mean, wherein is the mean of the frame signal; (3) kurtosis : reflects the degree of signal spikes, often used to represent the impact and non-Gaussian, wherein, is the standard deviation of the signal; a kurtosis value significantly greater than 3 indicates that the signal can have a pulse or abrupt component; (4) skewness : describes the symmetry of the signal distribution, If the skewness is positive, the signal distribution is biased to the left; if the skewness is negative, the distribution is biased to the right.

5. The transformer abnormality recognition method based on voiceprint feature analysis according to claim 4, characterized in that, Step (24) specifically comprises: Step (241) first, set the feature vector of sample point , the feature vector of sample point , , the Euclidean distance between sample point and is , the expression is: wherein denotes the dimension of the feature vector, i.e. the number of features contained in each sample point; when extracting the time domain features of the transformer acoustic signal, 4 features, i.e. root mean square, variance, kurtosis, skewness, are extracted, so the dimension of the feature vector of each sample is ; m is the index of a single feature; Step (242) then computes the local density of each sample point , which measures the degree of concentration of the point in the feature space; its mathematical definition is: wherein, is the cut-off distance. It is known that, The greater the value of the point is located in the region, the more the number of neighbors, the higher the local density. Step (243) then defines the minimum distance of a sample point representing the closest distance between the point and all sample points with a density greater than : Step (244) the minimum distance of a point is defined as the maximum distance to all sample points. defined as the maximum distance to all sample points. Step (245) introduces the composite index defined as: in, The value of reflects both the "density advantage" and "distance advantage" of the sample points: when a point has both high density and distance advantages... and At that time, its If the value is also relatively large, the point is determined to be the class center; finally, by drawing... Decision chart, combined with comprehensive indicators It can identify points with high local density and large distance as cluster centers of effective signals, while points with low local density but large distance are identified as abnormal noise points, thereby realizing the automatic identification and removal of noisy samples.

6. The transformer abnormality recognition method based on voiceprint feature analysis according to claim 5, characterized in that, Step (3) comprises: Step (31) firstly frames and windows the noise-reduced signal, and then performs discrete Fourier transform on each frame to obtain a time-frequency matrix The energy distribution of the signal in time and frequency can be described, and a spectrogram can be drawn to directly show the significant peak value in the low-frequency part and reflect the operating state characteristics of the transformer. The specific formula of the time-frequency matrix is as follows: wherein is the signal amplitude of the th sampling point, is the total number of sampling points per frame, is the complex amplitude of the th frequency component, denotes the imaginary unit; Step (32) introduces the energy spectrum as a key feature in the frequency domain feature; after Fourier transform of each frame of signal, the energy distribution is calculated: wherein, is the complex amplitude of the th frequency point, is the energy of the corresponding frequency point. Step (33) obtains the energy spectrum By Filter bank processing, obtaining the energy output of each filter , then taking the logarithm of the filter energy to obtain the logarithmic spectrum energy To enhance the nonlinear characteristics perceived by the human ear, then taking the discrete cosine transform of the logarithmic energy to obtain the MFCC coefficients: where M is the number of Mel filters, is the first order cepstral coefficient, is the order of the cepstral coefficient; To enhance the timing information, the first order difference of MFCC is calculated and the second order difference and finally the MFCC feature matrix is obtained: wherein, represents that the feature matrix belongs to a real number matrix of 19 rows and 39 columns, 19 rows correspond to 19 continuous signal frames, and 39 columns correspond to the feature dimension of each frame; Step (34) because the MFCC feature dimension is high, local linear embedding (LLE) method is used for dimension reduction to reduce redundancy and improve calculation efficiency.

7. The transformer anomaly identification method based on voiceprint feature analysis according to claim 6, characterized in that, Step (34) because the MFCC feature dimension is high, local linear embedding (LLE) method is used for dimension reduction to reduce redundancy and improve calculation efficiency, and the processing procedure is as follows: Step (341) assumes a set of MFCC feature samples , computes the Euclidean distance of each sample, and finds nearest neighbors: wherein, is the amplitude of the signal at the th sample point, is the number of sample points per frame, is the dimension of the original feature space, is the index variable of the feature dimension, taking values 1, 2, …, D ; respectively represent the feature values of the sample and in the p th dimension. Step (342) minimizes the reconstruction error for each sample, the reconstruction error being minimized ; ; wherein, is the total number of samples, i.e. the number of sample points contained in the MFCC feature set, is the number of neighbors for each sample, is the weight matrix , denotes the contribution weight of the th neighbor sample to the th sample, denotes the th sample of the a nearest neighbor sample; Step (343) holds the weight matrix Invariance, mapping high-dimensional sample data to low-dimensional space , get the output vector : The first several eigenvectors are selected by eigen decomposition to obtain the reduced dimension MFCC features; For the sample set mapped to the low-dimensional space, is the vector representation of the th sample in the low-dimensional space, and is the output feature after dimension reduction; is the vector representation of the th neighbor sample of the th sample in the low-dimensional space.

8. The transformer anomaly identification method based on voiceprint feature analysis according to claim 7, characterized in that, Step (4) comprises: Step (41) uses 2D-CNN to extract local features for the two-dimensional static time-frequency feature matrix of the spectrogram, focuses on capturing the frequency component distribution difference and energy concentration area features, and the core operation is two-dimensional convolution operation: ; wherein, represents the first layer output feature map, representing the abstract representation of the local time-frequency pattern of this layer, represents the input feature map set connected to the first feature map; represents the first layer connection input feature map and output feature map two-dimensional convolution kernel, used to extract local time-frequency texture, represents the first layer bias term of the first feature map, used to adjust the overall offset of the feature map, represents the activation function adopts , the formula is to alleviate the gradient vanishing problem, while enhancing the nonlinear expression ability of the model;​ Step (42) is to compress the data size, reduce the calculation complexity and enhance the robustness of the model to the position offset of the sensor. Max pooling layer is set after every 2 convolution layers to down-sample the local features. After 4 convolution layers and 2 pooling layers, the extracted deep feature map is flattened into a one-dimensional vector, which is input into 2 fully connected layers for feature integration, and a 2D-CNN feature vector is output F 2 D ; Step (43) adopts 3D-CNN to extract spatio-temporal joint features for the three-dimensional tensor feature composed of LLE-dimensioned MFCC features, focuses on capturing the dynamic changes of frequency components in different time segments, and the convolution operation formula is: ; Indicates the first Layer output feature tensor at position The value at that location, Corresponding time segmentation dimensions Corresponding frame time dimension Corresponding frequency dimension, The weights of the l-th layer 3D convolutional kernel These represent the size of the convolutional kernel in terms of time, frequency, and channels, respectively, and m is the index of the input channel of the previous layer. For the first Global bias terms for the layer; Step (44) finally inputs the obtained feature vector into a classification module to realize classification of the core looseness degree; in a full connection layer, 1088-dimensional fusion features are mapped into a 4-dimensional vector, and original scores of each category are output Then, the original scores are converted into a probability distribution by using a Softmax activation function, and the formula is as follows: Wherein, C=4 is the number of categories, represents the probability that the input sample belongs to the i-th fault state, and finally, the category with the maximum probability is selected as the model output label.

9. The transformer anomaly identification method based on voiceprint feature analysis according to claim 8, characterized in that, Step (5) comprises: Step (51) on the basis of the obtained Softmax output probability distribution, the class with the maximum probability is selected as the prediction result, so as to determine whether the transformer core has looseness abnormality and to distinguish different grades; Step (52) finally, according to the final looseness grade, a real-time diagnosis report containing state identification, feature details and early warning suggestions is generated.

Citation Information

Cited By

  • Wind turbine generator blade efficiency diagnosis method and device based on field operation data

    CN121760895A

  • Wind turbine blade efficiency diagnosis method and device based on field operation data

    CN121760895B