A method and device for power equipment defect identification based on acoustic imaging and acoustic fingerprint fusion

By using acoustic imaging and acoustic fingerprint fusion technology, the frequency information of acoustic data of power equipment is restored and spatial distribution features are extracted. Combined with modal response coupling monitoring, adaptive fusion is achieved, which solves the problems of data integrity and multi-dimensional feature capture in acoustic detection of power equipment, and realizes accurate identification and condition assessment of various defects.

CN122409862APending Publication Date: 2026-07-17STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO +1

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
Filing Date
2026-06-22
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing acoustic detection technologies in power equipment suffer from problems such as environmental noise interference, sensor installation position deviation, and signal transmission path obstruction, resulting in missing frequency components in the collected acoustic data. This makes it difficult to fully capture multi-dimensional defect features. Furthermore, insufficient feature extraction and rigid fusion strategies during defect identification make it difficult to accurately determine the confidence level of multiple defect candidate results.

Method used

By using acoustic imaging and voiceprint fusion technology, missing frequency information is recovered. Acoustic imaging technology is combined to extract spatial distribution features and voiceprint analysis technology to extract spectral evolution features. Modal response coupling monitoring is used to achieve adaptive fusion of image features and voiceprint features. Furthermore, a weight transfer enhancement mechanism is used to improve the identification ability of low-confidence candidates.

Benefits of technology

It enables accurate identification of various power equipment defects such as transformer core loosening, winding deformation, and partial discharge of switchgear, providing a reliable basis for condition assessment and preventive maintenance, and solving the problems of data integrity and multi-dimensional feature capture of acoustic detection technology in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122409862A_ABST
    Figure CN122409862A_ABST
Patent Text Reader

Abstract

This invention discloses a method and apparatus for identifying defects in power equipment based on acoustic imaging and acoustic signature fusion. It locates missing spectral segments through frequency domain integrity checks, and generates a complete acoustic flow by establishing a missing-compensation coupling relationship using energy attenuation nodes and compensation energy references. The complete acoustic flow is then subjected to acoustic imaging conversion to extract spatial sound pressure distribution, and multi-scale image feature channels are established based on energy transition band intervals through scale layering. Simultaneously, time-frequency decomposition is performed, identifying silent frequency band intervals as the basis for spectral segmentation, generating segmented acoustic signature feature clusters, and performing cross-frequency band correlation analysis. Modal response coupling monitoring identifies the distribution differences between image features and acoustic signature features, triggering a weighted modulation mechanism to achieve adaptive fusion. Low-confidence candidates are directionally allocated using weight transfer sources and weight demand maps to generate a high-confidence candidate set for defect identification, thus improving the accuracy and reliability of power equipment defect identification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power equipment condition monitoring technology, and in particular to a method and apparatus for identifying power equipment defects based on acoustic imaging and acoustic fingerprint fusion. Background Technology

[0002] During long-term operation, power equipment is susceptible to various potential defects due to factors such as mechanical wear, electrical aging, and insulation deterioration. Traditional power equipment inspection methods mainly rely on manual inspection and contact sensor monitoring, which suffer from drawbacks such as low detection efficiency, limited coverage, and difficulty in capturing early, weak defect signals. In recent years, non-contact detection technology based on acoustic signals has gradually gained attention. This technology identifies equipment status by analyzing the sound signals generated during operation, offering advantages such as convenient installation, real-time monitoring, and no downtime required.

[0003] Existing acoustic detection technologies still have many limitations in practical applications. The operating environment of power equipment is complex, and sound signals are easily affected by environmental noise interference, sensor installation deviations, and signal transmission path obstructions during propagation. This often results in missing frequency components and incomplete information in the acquired acoustic data. The acoustic features generated by different types of defects are distributed across different frequency bands and spatial locations, making it difficult for single-dimensional signal analysis to comprehensively capture the multi-dimensional features of defects. Furthermore, in the defect identification stage, due to insufficient feature extraction and rigid fusion strategies, multiple defect candidate results often have similar confidence levels, making accurate judgment difficult. Summary of the Invention

[0004] This invention discloses a method and apparatus for identifying defects in power equipment based on acoustic imaging and acoustic fingerprint fusion. It aims to recover missing frequency information through acoustic feature restoration technology, extract spatial distribution features by combining acoustic imaging technology and spectral evolution features by combining acoustic fingerprint analysis technology, achieve adaptive fusion of image features and acoustic fingerprint features by using modal response coupling monitoring, and improve the identification capability of low-confidence candidates through a weight transfer enhancement mechanism. Ultimately, it can accurately identify various defects in power equipment such as transformer core loosening, winding deformation, switchgear partial discharge, and poor busbar contact, providing a reliable basis for power equipment condition assessment and preventive maintenance.

[0005] The first aspect of this invention proposes a method for identifying defects in power equipment based on acoustic imaging and acoustic signature fusion, comprising the following steps: Collect the operating sound signal of the power equipment, perform frequency domain integrity verification on the operating sound signal to identify the missing spectrum segment, and use the missing spectrum segment to perform acoustic feature repair analysis to generate a complete acoustic flow; The complete acoustic flow is subjected to acoustic imaging conversion analysis to extract a spatial sound pressure distribution map. The spatial sound pressure distribution map is used to establish a multi-scale image feature channel. Spatial gradient analysis is performed on the multi-scale image feature channel to generate image features. The complete acoustic flow is subjected to time-frequency decomposition to extract the spectral energy distribution matrix. The silent frequency band intervals in the spectral energy distribution matrix are identified. The spectral energy is segmented and segmented using the silent frequency band intervals to generate segmented acoustic signature clusters. Cross-band correlation analysis is performed on the segmented acoustic signature clusters to generate acoustic signature features. Modal response coupling monitoring is performed on the image features and the voiceprint features to identify differences in feature distribution consistency. Based on the differences in feature distribution consistency, a weighted modulation mechanism is triggered to perform weighted fusion of the image features and the voiceprint features to form an enhanced fusion feature set. Based on the enhanced fusion feature set, a defect candidate set is generated. The defect candidate set is then subjected to competitive screening to identify the low-confidence candidate weight distribution. The low-confidence candidate weight distribution is then transferred and enhanced to generate a high-confidence candidate set. Based on the high-confidence candidate set, a defect type determination is performed to complete the defect identification.

[0006] A second aspect of the present invention proposes a power equipment defect identification device based on acoustic imaging and acoustic fingerprint fusion, comprising: The sound acquisition module is used to acquire the sound signals of power equipment operation, perform frequency domain integrity verification on the sound signals to identify missing spectral segments, and use the missing spectral segments to perform acoustic feature repair analysis to generate a complete acoustic stream; The acoustic imaging processing module is used to perform acoustic imaging conversion analysis on the complete acoustic flow to extract a spatial sound pressure distribution map, establish a multi-scale image feature channel using the spatial sound pressure distribution map, and perform spatial gradient analysis on the multi-scale image feature channel to generate image features. The voiceprint extraction module is used to perform time-frequency decomposition processing on the complete acoustic stream to extract the spectral energy distribution matrix, identify the silent frequency band intervals in the spectral energy distribution matrix, use the silent frequency band intervals to perform spectral energy segmentation to generate segmented voiceprint feature clusters, and perform cross-frequency band correlation analysis on the segmented voiceprint feature clusters to generate voiceprint features. The feature fusion module is used to perform modal response coupling monitoring and identification of feature distribution consistency differences between the image features and the voiceprint features, and trigger a weight modulation mechanism based on the feature distribution consistency differences to perform weighted fusion of the image features and the voiceprint features to form an enhanced fusion feature set; The defect identification module is used to generate a defect candidate set based on the enhanced fusion feature set, perform competitive screening on the defect candidate set to identify low-confidence candidate weight distribution, transfer and enhance the low-confidence candidate weight distribution to generate a high-confidence candidate set, and perform defect type determination based on the high-confidence candidate set to complete defect identification.

[0007] The beneficial effects of this invention are reflected in the following points: First, by identifying missing spectral segments through frequency domain integrity verification, energy attenuation nodes are located in the missing segments, and compensation energy references are obtained. Based on these two, a missing-compensation coupling point is generated to excite a spectral repair response, constructing a complete acoustic flow. Acoustic imaging conversion is performed on the complete acoustic flow to extract spatial sound pressure distribution. The sound pressure peak distribution is scale-layered using energy transition band intervals, establishing multi-scale image feature channels and extracting spatial gradient features. This technical approach solves the problem of missing frequency information in acoustic data due to sensor deviation or transmission obstruction, restoring the integrity of defective acoustic features and extracting sound pressure distribution patterns at different scales from a spatial dimension. Second, time-frequency decomposition is performed on the complete acoustic flow to extract the spectral energy distribution matrix. Silent frequency band intervals are identified as the basis for spectral segmentation. The spectrum is divided into multiple frequency band subsets, and voiceprint feature fragments are extracted. Clustering and integration are performed to construct segmented voiceprint feature clusters. Inter-band similarity calculation is performed on the segmented voiceprint feature clusters. Weakly correlated intervals are extracted from the correlation strength matrix as frequency band difference boundaries. Based on these difference boundaries, differential recombination is performed to form cross-frequency band feature chains. This technical approach segments spectral energy according to the natural intervals of acoustic characteristics, extracts unique acoustic patterns for each frequency band, and reveals the coupling relationships and differences between different frequency bands. Finally, modal response coupling monitoring is performed on image features and acoustic features to identify distribution differences and evaluate their levels. A weighted modulation strategy is triggered based on these differences, and weights are dynamically allocated according to modal reliability to achieve adaptive fusion. Confidence strength features are extracted from low-confidence candidates to form a weight transfer source. A weight demand map is constructed using confidence weakness features, and intensity-level mapping is performed to form hierarchical weight channels. Enhancement weights are generated through targeted allocation to improve low-confidence candidates. This technical approach achieves adaptive fusion of image features and acoustic features, avoiding the limitations of fixed weight strategies, and allocates surplus weights from high-confidence candidates to the weakness features of low-confidence candidates through a weight transfer mechanism. Attached Figure Description

[0008] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.

[0009] Unless otherwise specified, the same reference numerals in different figures represent the same or similar technical features, and different reference numerals may be used to represent the same or similar technical features.

[0010] Figure 1 This is a flowchart illustrating the power equipment defect identification method based on acoustic imaging and acoustic fingerprint fusion according to the present invention.

[0011] Figure 2 This is a structural block diagram of the power equipment defect identification device based on acoustic imaging and voiceprint fusion of the present invention. Detailed Implementation

[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0013] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0014] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0015] The technical solutions of the embodiments of this application will be described below.

[0016] like Figure 1 As shown, this embodiment of the invention provides a method for identifying defects in power equipment based on acoustic imaging and acoustic fingerprint fusion, including the following steps S110-S150: Step S110: Collect the operating sound signal of the power equipment, perform frequency domain integrity verification on the operating sound signal to identify the missing spectrum segment, and use the missing spectrum segment to perform acoustic feature repair analysis to generate a complete acoustic stream.

[0017] Specifically, the system collects sound signals from the power equipment during operation. High-sensitivity acoustic sensors are deployed at key locations such as the transformer casing, switchgear doors, and busbar trays. These sensors utilize MEMS condenser microphones with a frequency response range of 20Hz-20kHz. The acoustic sensors collect various sound signals generated by the power equipment in real time during operation, including mechanical vibration, electromagnetic noise, partial discharge, and cooling fan noise. The collected raw sound signals are preprocessed, employing a 50Hz notch filter to eliminate power frequency electromagnetic interference and a high-pass filter to remove low-frequency ambient noise, with the filter cutoff frequency set to 100Hz. During normal transformer operation, the operating sound signal is predominantly a low-frequency humming sound, with the spectral energy concentrated in the 100-500Hz frequency band. When the transformer core becomes loose or the windings vibrate abnormally, abnormal energy peaks appear in the mid-frequency band of 800-2000Hz. When partial discharge occurs inside the switchgear, pulse-type spike signals appear in the high-frequency band of 5-15kHz. The acquisition device continuously records the running sound signal at a sampling rate of 48kHz, and each sampling point is quantized with 16 bits to ensure that the time domain and amplitude characteristics of the sound signal are completely preserved.

[0018] Frequency domain integrity checks are performed on the operating sound signals to identify missing spectral segments. Through frequency domain integrity checks, a Fast Fourier Transform (FFT) is performed on the acquired operating sound signals to convert the time-domain signal into a frequency-domain representation, obtaining the spectral distribution of the sound signal. The spectrum is divided into three frequency bands: low-frequency (20-500Hz), mid-frequency (500-5kHz), and high-frequency (5-20kHz), corresponding to different types of equipment defect characteristics. The energy density distribution within each frequency band is calculated, obtained through the power spectral density function, with units of dB / Hz. Normally operating power equipment should exhibit a continuous energy distribution within the complete spectrum. Frequency ranges where energy density suddenly drops or disappears completely are identified. When the energy density of a frequency band is less than 40% of the average of adjacent frequency bands, that frequency band is marked as having energy deficiency. The start frequency, end frequency, and bandwidth of the energy deficiency frequency bands are statistically analyzed. Energy deficiency regions with bandwidth greater than 200Hz are defined as spectral deficiency segments. The transformer sensor's signal attenuation in the 2-4kHz frequency band was caused by installation misalignment, with the energy density in this band plummeting from -45dB / Hz to -70dB / Hz, identifying a missing spectral range in the 2-4kHz range. The switchgear's acoustic channel was partially blocked by an insulating partition, preventing effective propagation of 10-15kHz high-frequency signals; this band's energy was almost entirely lost, forming a 5kHz bandwidth missing spectral range. The center frequency, missing depth, and energy levels of adjacent intact frequency bands for each missing spectral range were recorded to provide baseline data for subsequent repair analysis.

[0019] In some embodiments, the step of generating a complete acoustic flow by performing acoustic feature repair analysis using the missing spectral segment includes: locating energy attenuation nodes in the missing spectral segment and simultaneously obtaining a compensation energy reference; generating a missing-compensation coupling point by performing frequency domain correlation between the energy attenuation nodes and the compensation energy reference; using the missing-compensation coupling point to excite a spectral repair response to generate a compensation spectral signal; and constructing a complete acoustic flow based on the compensation spectral signal.

[0020] In the spectrum missing segment, energy attenuation nodes are located and compensation energy references are obtained. At the start and end boundaries of the spectrum missing segment, the transition points where energy density shifts from normal to missing levels are identified; these transition points are the energy attenuation nodes. Energy attenuation nodes correspond to the critical frequency points where spectral energy begins to drop significantly or recovers to normal levels. In the 2-4 kHz spectrum missing segment, the initial energy attenuation node is located at 1.9 kHz, where energy density drops from -45 dB / Hz to -70 dB / Hz; the final energy attenuation node is located at 4.1 kHz, where energy density recovers from -68 dB / Hz to -43 dB / Hz. Energy characteristics of the complete frequency bands on both sides of the spectrum missing segment are extracted as compensation energy references. The compensation energy references include the average energy density of the complete frequency band to the left of the missing segment, the average energy density of the complete frequency band to the right of the missing segment, and the energy change trends of the frequency bands on both sides. For the missing 2-4kHz band, the average energy density of the complete 1.5-1.9kHz band on the left is -46dB / Hz, and the average energy density of the complete 4.1-4.5kHz band on the right is -44dB / Hz. The energy levels on both sides are similar and show a gradual trend; therefore, the average value of -45dB / Hz is used as the compensation energy reference. When the energy levels on both sides of the missing band differ significantly, a weighted average is used to calculate the compensation energy reference. The weights are determined based on the frequency distance between the two bands and the missing band; the closer the distance, the greater the weight.

[0021] Missing-compensation coupling points are generated based on frequency domain correlation between energy attenuation nodes and compensation energy references. Frequency domain correlation maps the frequency positions of energy attenuation nodes to the energy values ​​of the compensation energy references. Within the spectral missing segment, multiple repair target frequencies are set at equal intervals of 100Hz. For the 2-4kHz missing segment, 21 repair target frequencies are set at positions from 2.0kHz, 2.1kHz, 2.2kHz to 4.0kHz. A frequency distance correlation is established between each repair target frequency and the starting and ending energy attenuation nodes, defined as the frequency difference between the target frequency and the attenuation node. The expected energy density of each repair target frequency is calculated using the interpolation relationship between the compensation energy reference and the frequency distance. A linear interpolation method is used, where the expected energy density of the repair target frequency is calculated between the left and right compensation energy references based on its relative position within the missing segment. The frequency positions of the repair target frequencies and their expected energy densities are combined to form missing-compensation coupling points through frequency domain correlation. The coupling point contains information in two dimensions: frequency coordinates and energy coordinates. The frequency coordinates indicate the frequency location that needs to be repaired, and the energy coordinates indicate the energy level that should be restored to at that location.

[0022] A compensated spectral signal is generated by exciting a spectral restoration response using missing-compensation coupling points. A frequency domain restoration template is constructed based on the frequency and energy coordinates of each missing-compensation coupling point. The restoration template sets a target energy value at the frequency position of each coupling point, exciting the spectral restoration response. A smooth interpolation algorithm fills in a continuous energy curve between adjacent coupling points, avoiding abrupt energy jumps in the restored spectrum. A cubic spline interpolation algorithm is used to ensure that the restored spectral energy curve accurately matches the desired energy value at the coupling point and maintains continuity of the second derivative throughout the entire missing segment. The frequency domain representation of the compensated spectral signal is generated based on the spectral restoration response. The compensated spectral signal has an energy distribution consistent with the restoration template within the missing spectral segment, and its energy is set to zero in the frequency range outside the missing segment. An inverse Fourier transform is performed on the frequency domain representation of the compensated spectral signal to convert it into a time-domain waveform. Phase information is extracted from the complete frequency bands on both sides of the missing segment and applied to the compensated spectral signal to ensure phase consistency with the original signal. For the compensated spectrum signal generated in the 2-4kHz missing segment, its energy density in this frequency band is restored to the range of -45 to -48dB / Hz, the energy curve transitions smoothly, and it seamlessly connects with the complete frequency band of the original signal at the boundaries of 1.9kHz and 4.1kHz.

[0023] A complete acoustic stream is constructed based on the compensated spectral signal. The compensated spectral signal undergoes frequency domain expansion processing, extending it from the missing segment to the complete 20Hz-20kHz frequency range. Within the missing spectral segment, the repaired spectral components of the compensated spectral signal are preserved, and the energy density of this portion has been restored to normal levels. Outside the missing segment, the compensated spectral signal fills in the original complete spectral components of the acquired signal, maintaining the original energy distribution of these frequency bands. A frequency domain transition band is set near the boundary of the missing segment, with a width of 5% of the missing segment bandwidth. Within the transition band, a weighted fusion algorithm is applied to smoothly connect the repaired components with the complete components. The weighting coefficients change linearly according to the distance of the frequency point from the boundary, achieving a smooth transition. The expanded and fused compensated spectral signal is then subjected to an inverse Fourier transform, converting it into a time-domain sound signal, which is the complete acoustic stream. The complete acoustic stream maintains the same sampling rate (48kHz) and quantization precision (16-bit) as the original acquired signal in the time domain, and covers the complete 20Hz-20kHz frequency range in the frequency domain, with continuous energy distribution in each frequency band without significant gaps. After acoustic repair, the complete acoustic flow of the transformer showed that the energy density in the 2-4kHz frequency band recovered from the missing state of -70dB / Hz to the normal level of -46dB / Hz. The winding vibration characteristic signals contained in this frequency band were successfully recovered, providing complete acoustic characteristic data for subsequent defect identification.

[0024] Step S120: Perform acoustic imaging conversion analysis on the complete acoustic flow to extract the spatial sound pressure distribution map, use the spatial sound pressure distribution map to establish a multi-scale image feature channel, and perform spatial gradient analysis on the multi-scale image feature channel to generate image features.

[0025] Specifically, acoustic imaging conversion analysis is performed on the complete acoustic flow to extract the spatial sound pressure distribution map. Through acoustic imaging conversion analysis, the complete acoustic flow is input into an acoustic beamforming algorithm. This algorithm calculates the spatial location distribution of the sound source based on the time delay and phase difference between sound signals collected by multiple acoustic sensors. An 8×8 array of 64 sensors is deployed on the surface of the transformer casing, with a sensor spacing of 15 cm and an array coverage area of ​​1.05 m × 1.05 m. A delay-sum beamforming algorithm is used to accumulate the sound pressure contribution of each frequency component in the complete acoustic flow at various points in space. The spatial region is divided into a 256×256 pixel grid, with each pixel corresponding to a 5 mm × 5 mm actual spatial area. The sound pressure amplitude at each pixel location is calculated; the sound pressure amplitude is obtained by coherently superimposing the signals from each sensor at that location after time delay compensation. The calculated sound pressure amplitude is normalized, with the maximum sound pressure value normalized to 255 and the minimum sound pressure value normalized to 0, forming a two-dimensional image with a grayscale range of 0-255. This two-dimensional image is a spatial sound pressure distribution map, where the pixel grayscale values ​​represent the sound pressure intensity at the corresponding spatial location. High grayscale areas in the spatial sound pressure distribution map correspond to the main radiation locations of the sound source, while low grayscale areas correspond to locations with weaker sound energy. In the distribution map, transformer core loosening defects appear as localized high sound pressure patches in the core area, with grayscale values ​​reaching 200-230 and patch sizes of approximately 8-12 centimeters.

[0026] In some embodiments, establishing a multi-scale image feature channel using the spatial sound pressure distribution map includes: acquiring the sound pressure peak distribution features of the spatial sound pressure distribution map; locating the energy transition zone in the spatial sound pressure distribution map; performing scale layering on the sound pressure peak distribution features based on the energy transition zone to generate layered sound pressure signals; and performing channelization processing on the layered sound pressure signals to establish a multi-scale image feature channel.

[0027] The peak distribution characteristics of the spatial sound pressure distribution map are obtained. Local peak detection is performed on the spatial sound pressure distribution map to identify pixels with gray values ​​higher than their surrounding neighborhoods. A 3×3 sliding window is used to traverse the entire distribution map. When the gray value of the central pixel is greater than the gray values ​​of the other 8 pixels within the window, that central pixel is marked as a local peak point. The gray values ​​and spatial coordinates of all local peak points are statistically analyzed. The gray value represents the sound pressure intensity at that location, and the spatial coordinates, in pixels, represent the position of the peak point in the image. The spatial distance between each local peak point is calculated, defined as the Euclidean distance between the coordinates of two peak points. The density distribution of peak points is analyzed. Density is expressed as the number of peak points per unit area, in units of points / square centimeter. The peak point density is high in the transformer defect area, with 15-25 peak points detected within a 10 cm × 10 cm area, and the gray value range of the peak points is 180-230. The peak point density is low in the normal operation area, with only 3-8 peak points detected within the same area, and the gray value range of the peak points is 90-130. Statistical features of the peak points are extracted, including the total number of peak points, the average gray value, and the standard deviation of the gray value. These statistical features are then integrated into a sound pressure peak distribution feature, which includes information on three dimensions: the number of peaks, their intensity, and their spatial distribution.

[0028] Locate the energy transition zone in the spatial sound pressure distribution map. Calculate the gray-level gradient between each pixel and its neighboring pixels in the spatial sound pressure distribution map. The gray-level gradient is defined as the difference between the gray-level value of the center pixel and the average gray-level value of its neighboring pixels. The Sobel operator is used to calculate the gray-level gradients in the horizontal and vertical directions. The horizontal gradient Gx and the vertical gradient Gy are obtained through convolution operations. Calculate the gradient magnitude for each pixel. The gradient magnitude reflects the degree of grayscale change at a given location. A gradient magnitude threshold of 30 is set; when a pixel's gradient magnitude exceeds this threshold, it is marked as a high-gradient pixel. High-gradient pixels correspond to locations where sound pressure changes rapidly in space, typically at the boundary between high and low sound pressure regions. Spatially continuous high-gradient pixels are aggregated into gradient band regions, defined as energy transition zones. Energy transition zones mark the spatial range of sound pressure energy transitioning from high to low values. In the sound pressure distribution of transformer defect radiation, the grayscale value of the sound pressure at the defect center (220) gradually decreases to 100 towards the periphery, forming an energy transition zone in a ring-shaped area 3-5 cm from the center. The starting and ending boundary coordinates and average gradient magnitude of each energy transition zone are recorded to provide spatial positioning data for subsequent scale stratification.

[0029] Layered sound pressure signals are generated by scaling the peak sound pressure distribution characteristics based on the energy transition band. According to the spatial location of the energy transition band, the spatial sound pressure distribution map is divided into three layers: a core high-pressure region, a transition region, and a peripheral low-pressure region. The core high-pressure region corresponds to the high sound pressure area inside the transition band, where the gray value of all pixels is higher than the average gray value of the transition band plus one standard deviation. The transition region corresponds to the energy transition band itself, where pixel gray values ​​change rapidly. The peripheral low-pressure region corresponds to the low sound pressure area outside the transition band, where pixel gray values ​​are lower than the average gray value of the transition band minus one standard deviation. Peak points in the sound pressure peak distribution characteristics are classified according to their spatial layer: peak points in the core high-pressure region are classified as high-intensity peaks, peak points in the transition region as medium-intensity peaks, and peak points in the peripheral low-pressure region as low-intensity peaks. Feature parameters of the peak points in each layer are extracted, including the number of peak points and the average gray value. For the sound pressure distribution of transformer defects, the core high-voltage zone contains 8-12 peak points with an average gray value of 210, the transition zone contains 10-18 peak points with an average gray value of 155, and the peripheral low-voltage zone contains 20-35 peak points with an average gray value of 95. The peak characteristic parameters of each level are combined with the corresponding spatial sound pressure data to generate a layered sound pressure signal through scale layering. The layered sound pressure signal contains three independent signal components, each corresponding to a spatial level. The amplitude of a component is determined by the average gray value of that level, and the texture characteristics of a component are determined by the peak distribution characteristics of that level.

[0030] Multi-scale image feature channels are established by channelizing the layered sound pressure signals. Through channelization, the three components of the layered sound pressure signals are processed as independent data channels. A high-resolution feature extraction algorithm is applied to the core high-pressure region to extract pixel-level detailed texture features, which are calculated using the gray-level co-occurrence matrix. A medium-resolution feature extraction algorithm is applied to the transition region to extract region-level edge and gradient features. Edge features are detected using the Canny operator, and gradient features are described using the Histogram of Oriented Gradients (HOG). The edge density in the transition region is 15-25 edges / cm², and the principal gradient direction coincides with the defect radiation direction. A low-resolution feature extraction algorithm is applied to the peripheral low-pressure region to extract the overall shape and distribution features. The feature extraction results of the three components are stored in three independent feature channels, each retaining complete feature information at its scale. This channelization process establishes a multi-scale image feature channel structure. The multi-scale image feature channel structure includes high-resolution, medium-resolution, and low-resolution channels, each corresponding to sound pressure features at different spatial scales. The data format of the multi-scale image feature channels is a three-dimensional tensor. The first dimension corresponds to 256 pixels in the X coordinate space, the second dimension corresponds to 256 pixels in the Y coordinate space, and the third dimension corresponds to 3 feature channels.

[0031] Spatial gradient analysis is performed on multi-scale image feature channels to generate image features. Spatial gradient analysis is used to calculate the spatial gradient for each channel in the multi-scale image feature channels. The central difference algorithm is used to calculate the gradient. For a pixel at position (x, y) in a channel, its horizontal gradient is [f(x+1, y) - f(x-1, y)] / 2, and its vertical gradient is [f(x, y+1) - f(x, y-1)] / 2. Gradient magnitude and gradient direction maps are calculated for each channel. The gradient magnitude map reflects the intensity of feature changes, and the gradient direction map reflects the direction of feature changes. The distribution characteristics of gradient magnitude for each channel are statistically analyzed, including gradient mean, gradient variance, proportion of high-gradient pixels, and gradient distribution entropy. High-resolution channels have higher gradient mean values, typically in the range of 35-50, reflecting drastic feature changes in the core region; low-resolution channels have lower gradient mean values, typically in the range of 10-20, reflecting gentler feature changes in the peripheral region. The consistency of gradient direction for each channel is analyzed, and a gradient direction histogram is calculated to statistically analyze the distribution of gradients in each direction within the 0-360 degree range. The gradient direction of the radiated sound pressure from the defect exhibits a radial distribution from the center outwards, with the dominant gradient directions concentrated at 4-8 discrete angles. Cross-channel gradient correlation features are extracted, and the correlation coefficient between the high-resolution and medium-resolution channel gradients is calculated. This correlation coefficient reflects the degree of correlation between features at different scales. Through spatial gradient analysis, the gradient statistical features, directional distribution features, and cross-channel correlation features of each channel are integrated to form an image feature vector. This image feature vector contains 36 dimensions: 12 dimensions for gradient statistical features, 18 dimensions for directional distribution features, and 6 dimensions for cross-channel correlation features. This vector comprehensively describes the spatial structure and multi-scale characteristics of the sound pressure distribution.

[0032] Step S130: Perform time-frequency decomposition processing on the complete acoustic flow to extract the spectral energy distribution matrix, identify the silent frequency band intervals in the spectral energy distribution matrix, use the silent frequency band intervals to segment the spectral energy to generate segmented acoustic signature feature clusters, and perform cross-band correlation analysis on the segmented acoustic signature feature clusters to generate acoustic signature features.

[0033] Specifically, time-frequency decomposition is performed on the complete acoustic stream to extract the spectral energy distribution matrix. A Short-Time Fourier Transform (STFT) is performed on the complete acoustic stream to convert the time-domain sound signal into a time-frequency domain representation. The time window length is set to 2048 sampling points, corresponding to a time resolution of 42.7 milliseconds. A Hanning window is used to reduce spectral leakage. The overlap rate between adjacent windows is set to 75%, i.e., sliding 512 sampling points each time, to ensure the continuity of the time-frequency representation. A Fourier transform is performed on the sound signal within each time window to obtain the spectral distribution of that time window, with a spectral resolution of 23.4 Hz. The energy density of each frequency component is calculated; the energy density is equal to the square of the spectral amplitude. The energy densities of all time windows are arranged in chronological order to construct a two-dimensional spectral energy distribution matrix. The rows of the matrix correspond to different time frames, and the columns correspond to different frequency channels. The matrix element values ​​represent the energy intensity at the corresponding time and frequency position. For a complete acoustic stream with a sampling rate of 48 kHz and a duration of 10 seconds, the spectral energy distribution matrix contains approximately 938 time frames and 1025 frequency channels. A logarithmic transformation is performed on the matrix elements to convert the energy values ​​to decibels. The calculation formula is E_dB = 10 × log10(E), where E is the original energy value. The dynamic range of the logarithmically transformed matrix is ​​from -80dB to 0dB. During normal operation of the transformer, the spectral energy distribution matrix exhibits a continuous high energy distribution in the low-frequency band of 100-500Hz, with energy values ​​remaining within the range of -10 to -5dB.

[0034] Identify the silent frequency bands in the spectral energy distribution matrix. Calculate the average energy of each frequency channel in the spectral energy distribution matrix over the time dimension. The average energy is obtained by averaging the energy values ​​of all time frames for that frequency channel. Plot a frequency-average energy curve, with the horizontal axis representing the frequency range of 20Hz-20kHz and the vertical axis representing the average energy range of -80dB to 0dB. Identify the frequency ranges in the curve where the energy is consistently below the background noise level, set to -55dB. When the average energy of a frequency channel is below -55dB and the bandwidth of this low-energy state is greater than 500Hz, that frequency band is marked as a silent frequency band. Silent frequency bands correspond to frequency ranges where the energy is extremely weak or completely absent in the sound signal; these bands are typically located between the radiation bands of different types of sound sources. The spectral energy distribution matrix of the transformer operating sound shows that the average energy of the 600-800Hz band is -62dB, significantly lower than the bands on either side; this band is identified as a silent frequency band with a bandwidth of 200Hz. In the 3-5kHz frequency band, the average energy drops to -68dB, forming a relatively wide silent band interval of 2kHz, which separates mid-frequency mechanical vibration noise from high-frequency electromagnetic noise. The statistical spectral energy distribution matrix shows the start frequency, end frequency, average energy, and bandwidth of all silent band intervals.

[0035] In some embodiments, the step of segmenting the spectrum energy using the silent frequency band interval to generate segmented voiceprint feature clusters includes: locating and identifying segmentation points based on the spectrum boundary of the silent frequency band interval; dividing the spectrum energy region using the segmentation points to form a frequency band subset; extracting voiceprint features from the frequency band subset to generate feature segments; and performing clustering and integration on the feature segments to construct segmented voiceprint feature clusters.

[0036] Spectral boundary location and segmentation point identification are performed based on silent frequency bands. The center frequency of each silent frequency band is extracted, and the center frequency is equal to the arithmetic mean of the start and end frequencies. Segmentation points are set at the center frequency positions of the silent frequency bands, serving as boundary markers for the spectral energy regions. For the 600-800Hz silent frequency band, the center frequency is 700Hz, and a segmentation point is set at this frequency. For the 3-5kHz silent frequency band, the center frequency is 4kHz, and another segmentation point is set at this location. The frequency spacing between adjacent segmentation points is checked. When the spacing is less than 1kHz, the two silent frequency bands are merged, and a single segmentation point is set at the center of the merged band to avoid overly fragmented frequency band division. For a spectral energy distribution matrix with a frequency range of 20Hz-20kHz, spectral boundary location can typically identify 5-8 segmentation points, dividing the entire spectrum into 6-9 active frequency bands. The frequency position of each segmentation point and the corresponding energy level of the silent frequency band are recorded. The sequence of segmentation points for the transformer audio signal is [700Hz, 1800Hz, 4000Hz, 8500Hz, 15000Hz]. These 5 segmentation points divide the spectrum into 6 frequency band subsets.

[0037] Frequency band subsets are formed by dividing the spectral energy region using segmentation points. Based on the frequency location of the segmentation points, the frequency dimension of the spectral energy distribution matrix is ​​divided into multiple continuous frequency bands. The first frequency band subset extends from the starting frequency of the spectrum (20Hz) to the first segmentation point frequency; the second frequency band subset extends from the first segmentation point to the second segmentation point, and so on, with the last frequency band subset extending from the last segmentation point to the ending frequency of the spectrum (20kHz). For the segmentation point sequence [700Hz, 1800Hz, 4000Hz, 8500Hz, 15000Hz], the resulting frequency band subsets are: Subset 1 covers 20-700Hz, Subset 2 covers 700-1800Hz, Subset 3 covers 1800-4000Hz, Subset 4 covers 4000-8500Hz, Subset 5 covers 8500-15000Hz, and Subset 6 covers 15000-20000Hz. Extract local regions of the spectral energy distribution matrix corresponding to each frequency band subset. These local regions contain energy data from all frequency channels and all time frames within that frequency band. Calculate the statistical characteristics of each frequency band subset, including bandwidth, average energy, and peak energy. Subset 1 has a bandwidth of 680Hz, an average energy of -12dB, and a peak energy of -6dB, corresponding to low-frequency mechanical vibration noise from a transformer. Subset 3 has a bandwidth of 2200Hz, an average energy of -25dB, and a peak energy of -15dB, corresponding to mid-frequency electromagnetic noise. Integrate each frequency band subset and its statistical characteristics into independent data units.

[0038] Feature segments are generated by extracting voiceprint features from frequency band subsets. The Mel frequency-cephalic coefficient (MFCC) extraction algorithm is applied to the energy distribution data of each frequency band subset. The frequency axis of the frequency band subset is converted to a Mel scale, which better reflects the human ear's frequency perception characteristics; the conversion formula is Mel = 2595 × log10(1 + f / 700). Twenty-six triangular filter banks are set on the Mel scale, each filter covering a different part of the frequency range of the subset. The energy output of each filter is calculated, and after taking the logarithm of the energy, a Discrete Cosine Transform (DCT) is performed. The first 13 DCT coefficients are extracted as MFCC features. A 13-dimensional MFCC feature vector is extracted for each time frame in the frequency band subset. For a 10-second signal containing 938 time frames, this subset generates a 938 × 13 feature matrix. The statistics of this feature matrix in the time dimension are calculated, including the mean, standard deviation, maximum, and minimum values ​​of each dimension, forming a 52-dimensional static MFCC statistical feature. The first and second differences of MFCC are calculated to reflect the dynamic trend of voiceprint features. The difference features are also extracted using 52-dimensional statistics. The 52-dimensional static features, the 52-dimensional first-order difference, and the 52-dimensional second-order difference are combined to form a 156-dimensional feature segment vector.

[0039] Clustering and merging of feature segments are performed to construct segmented voiceprint feature clusters. Feature segments generated from all frequency band subsets are collected, resulting in six 156-dimensional feature segment vectors for each of the six subsets. The Euclidean distance between the feature segments is calculated using the following formula: Let xi and yi be the values ​​of the two feature segments in the i-th dimension, respectively. A 6×6 distance matrix is ​​constructed, where the matrix elements represent the similarity between corresponding feature segment pairs; the smaller the distance, the more similar the voiceprint patterns of the two segments. A hierarchical clustering algorithm is used to group the feature segments, with the cluster distance threshold set to 1.5 times the median distance. When the distance between two feature segments is less than the threshold, they are grouped into the same cluster. For the transformer sound signal, the feature segment distance of subset 1 and subset 2 is 8.5, which is less than the threshold of 12.3, and they are grouped into the low-frequency cluster; the feature segment distance of subset 3 and subset 4 is 9.8, and they are grouped into the mid-frequency cluster; subsets 5 and subset 6 form separate high-frequency clusters. Representative features of each cluster are extracted. The representative features are calculated by weighted averaging of all feature segments within the cluster, with the weights determined based on the average energy of each frequency band subset. The representative features of each cluster and the frequency band range of the members within the cluster are integrated to construct segmented voiceprint feature clusters through cluster integration. The segmented voiceprint feature clusters consist of three clusters: low-frequency, mid-frequency, and high-frequency. Each cluster contains a 156-dimensional representative feature vector and a corresponding frequency range label.

[0040] In some embodiments, performing cross-band association analysis on the segmented voiceprint feature clusters to generate voiceprint features includes: calculating the inter-band similarity of the segmented voiceprint feature clusters to generate an association strength matrix; extracting weak association intervals from the association strength matrix as frequency band difference boundaries; performing differential recombination on the segmented voiceprint feature clusters based on the frequency band difference boundaries to form a cross-band feature chain; and performing vectorized reconstruction on the cross-band feature chain to generate voiceprint features.

[0041] Inter-band similarity calculation is performed on segmented voiceprint feature clusters to generate an association strength matrix. Through inter-band similarity calculation, representative feature vectors of each cluster in the segmented voiceprint feature clusters are extracted. For the case containing three clusters, the feature vectors for the low-frequency cluster (F_low), mid-frequency cluster (F_mid), and high-frequency cluster (F_high) are obtained, each with a dimension of 156. The cosine similarity between any two cluster feature vectors is calculated using the following formula: × [ ] where xi and yi are the values ​​of the two vectors in the i-th dimension, respectively. The cosine similarity ranges from -1 to 1. The closer the value is to 1, the more similar the voiceprint patterns of the two clusters are; the closer the value is to -1, the greater the difference in patterns. The similarity S_low-mid between low-frequency and mid-frequency clusters, S_low-high between low-frequency and high-frequency clusters, and S_mid-high between mid-frequency and high-frequency clusters are calculated. For transformer sounds, S_low-mid = 0.62, indicating a certain correlation between low-frequency vibrations and mid-frequency noise; S_low-high = 0.28, indicating a large difference between low-frequency and high-frequency features; and S_mid-high = 0.45, indicating a moderate correlation between mid-frequency and high-frequency sounds. A 3×3 correlation strength matrix is ​​constructed. Diagonal elements of the matrix are 1, indicating that a cluster is completely similar to itself; off-diagonal elements are filled with the corresponding similarity values. The rows and columns of the correlation strength matrix correspond to low-frequency, mid-frequency, and high-frequency clusters, respectively, and the matrix elements visually represent the degree of correlation between voiceprint features in each frequency band.

[0042] For example, extracting weak correlation intervals from the correlation strength matrix as frequency band difference boundaries includes: performing threshold segmentation on the correlation strength matrix to identify the distribution of low-intensity elements; mapping the low-intensity element distribution to a frequency band coordinate system to generate weak correlation location markers; merging the weak correlation location markers into intervals to form continuous weak correlation bands; and selecting the boundary nodes of the continuous weak correlation bands to determine the frequency band difference boundaries.

[0043] Threshold segmentation was performed on the association strength matrix to identify the distribution of low-intensity elements. A threshold of 0.40 was set, determined statistically based on a large amount of power equipment sound data. This threshold indicates that clusters with similarity below this value are considered to have weak association and significant voiceprint differences. All off-diagonal elements of the association strength matrix were traversed. The matrix size was N×N, where N is the number of segmented voiceprint feature clusters. For the case of 3 clusters, the matrix was 3×3 containing 6 off-diagonal elements. When an element value was less than 0.40, it was marked as a low-intensity element, and its row and column coordinates in the matrix were recorded. For the transformer association strength matrix, S_low-high=0.28 was less than the threshold of 0.40 and was marked as a low-intensity element with coordinates (1, 3) and the symmetrical position (3, 1); S_low-mid=0.62 and S_mid-high=0.45 were both greater than the threshold and remained normal association elements. The total number and distribution pattern of low-intensity elements in the matrix were statistically analyzed. Record the specific similarity values ​​of each low-intensity element and its corresponding cluster pair identifier. The closer the similarity value is to 0, the more orthogonal and essential the voiceprint patterns of the corresponding cluster pair are. Integrate the coordinates, values, and cluster pair information of the low-intensity elements into a low-intensity element distribution.

[0044] The distribution of low-intensity elements is mapped to a frequency band coordinate system to generate weak association location markers. A mapping table between cluster indices and frequency ranges is established, recording each cluster number, cluster name, and corresponding frequency interval. Low-frequency cluster number 1 corresponds to the frequency range of 20-1800Hz, mid-frequency cluster number 2 corresponds to 1800-8500Hz, and high-frequency cluster number 3 corresponds to 8500-20000Hz. Cluster pair identifiers are extracted from the low-intensity element distribution. For a low-intensity element at coordinate (1, 3), it corresponds to a weak association between a low-frequency cluster and a high-frequency cluster. The frequency ranges of the two clusters are queried according to the mapping table: the low-frequency cluster 20-1800Hz and the high-frequency cluster 8500-20000Hz. The interval region of 1800-8500Hz between the two frequency bands is identified. This interval region corresponds to the projection position of the weak association in the frequency domain in the association strength matrix, and a weak association location marker is set in this region. The tags are represented in a structured format, containing three fields: start frequency, end frequency, and association strength value. For weak low-frequency-high-frequency associations, the tags are set to start at 1800Hz, end at 8500Hz, and have an intensity of 0.28. The same mapping operation is performed on each element in the low-intensity element distribution, transforming the cluster pairs in the matrix abstract space into interval positions in the specific frequency domain. All weak association position tags are collected and sorted from low to high start frequency to form an ordered tag sequence.

[0045] Weakly correlated location markers are merged into continuous weakly correlated bands. The frequency intervals of adjacent weakly correlated location markers are checked for overlap, inclusion, or adjacency. A rule for determining the relationship between two marker intervals is defined: if the ending frequency of the preceding interval is greater than or equal to the starting frequency of the following interval, or if the interval interval is less than 500Hz, it is considered mergingable. The merging operation takes the smaller starting frequency and the larger ending frequency of the two intervals, and the correlation strength is the weighted average of the two marker strengths, with the weights determined based on the bandwidth of each interval. The sorted marker sequence is checked and merged sequentially from front to back until all mergingable adjacent intervals are integrated to form continuous weakly correlated bands. The number of merged continuous weakly correlated bands, the frequency range of each band, the bandwidth, and the average correlation strength are counted. The structured information of each continuous weakly correlated band is recorded, including four fields: starting frequency, ending frequency, bandwidth, and average strength.

[0046] Boundary nodes of consecutive weakly correlated bands are selected as frequency band difference boundaries. The start and end frequencies of each consecutive weakly correlated band are extracted; these two frequency positions constitute the left and right boundary nodes of that weakly correlated band. For a consecutive weakly correlated band starting at 1800Hz and ending at 8500Hz, with a bandwidth of 6700Hz and an average strength of 0.30, its left boundary node is at 1800Hz, and its right boundary node is at 8500Hz. The validity of the boundary nodes is verified by checking whether the node is located at the intersection of two highly correlated clusters. A valid boundary node should satisfy the condition that the correlation strength between its left and right frequency bands is significantly higher than the correlation strength crossing the node. The average correlation strength within a 500Hz range on both sides of the boundary node is calculated. When the correlation strength within both sides is greater than 0.60 and the correlation strength across nodes is less than 0.40, the node is confirmed as a valid frequency band difference boundary. For the 1800Hz boundary, the correlation strength within the low-frequency band (20-1800Hz) to its left is 0.75, and the correlation strength within the mid-low-frequency band (1800-4000Hz) to its right is 0.68. However, the low-to-mid-frequency correlation strength across 1800Hz is only 0.32, satisfying the validity condition. Frequency band difference boundaries mark the dividing lines between frequency bands dominated by different physical sound sources or different acoustic mechanisms. The 1800Hz boundary of the transformer separates low-frequency mechanical vibration sound from mid-frequency electromagnetic noise, and the 8500Hz boundary separates mid-low-frequency noise from high-frequency partial discharge sound. The frequency positions and correlation strength comparisons on both sides of all frequency band difference boundaries are recorded.

[0047] Based on frequency band difference boundaries, segmented voiceprint feature clusters are differentially recombined to form a cross-frequency band feature chain. Through differential recombination, the segmented voiceprint feature clusters are divided into independent feature groups according to the frequency band difference boundaries. The frequency band difference boundaries of 1800Hz and 8500Hz divide the three clusters into three independent groups: [low-frequency cluster], [mid-frequency cluster], and [high-frequency cluster]. Voiceprint features within each independent group are enhanced, and the top 50 dimensions with the largest variance in the feature vector of that group are extracted. These dimensions correspond to the most discriminative voiceprint components in that frequency band. Features between different independent groups are differentially encoded, and the difference vector between the feature vectors of adjacent groups is calculated. The difference vector reflects the changes in voiceprint patterns when crossing frequency band difference boundaries. The difference vector D_low-mid between the low-frequency cluster and the mid-frequency cluster is obtained by subtracting F_low from F_mid. This vector highlights the incremental part of the mid-frequency features relative to the low-frequency features. The enhanced features of each independent group are concatenated with the differential codes of adjacent groups in ascending frequency order to form a cross-frequency band feature chain. The cross-frequency feature chain structure is [50-dimensional low-frequency enhancement features, 156-dimensional low-to-medium frequency difference features, 50-dimensional mid-frequency enhancement features, 156-dimensional mid-to-high frequency difference features, and 50-dimensional high-frequency enhancement features], with a total dimension of 462.

[0048] Vectorization reconstruction is performed on the cross-frequency band feature chain to generate voiceprint features. Through vectorization reconstruction, principal component analysis (PCA) is used to reduce the dimensionality of the 462-dimensional vector of the cross-frequency band feature chain, retaining principal components with a cumulative variance contribution rate of 95%. The covariance matrix of each dimension of the feature chain is calculated, and eigenvalue decomposition is performed on the covariance matrix. The top k principal components are selected according to the eigenvalues ​​in descending order, with the value of k such that the sum of the top k eigenvalues ​​accounts for more than 95% of the total sum of eigenvalues. For the sound signal of power equipment, 80-120 principal components are usually needed to achieve a 95% variance contribution rate. The cross-frequency band feature chain is projected onto the selected principal component space to obtain the dimensionality-reduced feature vector, which has a dimension of 80-120. The dimensionality-reduced feature vector is normalized, scaling the values ​​of each dimension to the range of 0-1. The normalized feature vector is the final voiceprint feature. The voiceprint feature has a dimension of 100 and includes comprehensive information on spectral energy distribution, inter-band correlation patterns, and cross-frequency band variation characteristics. The acoustic signature of a transformer in normal operation has a relatively large value on the low-frequency principal component, remaining in the range of 0.6-0.8. The acoustic signature of a loose core defect shows an abnormal peak value on the mid-frequency principal component.

[0049] Step S140: Modal response coupling monitoring is performed on image features and voiceprint features to identify differences in feature distribution consistency. Based on the differences in feature distribution consistency, a weighted modulation mechanism is triggered to perform weighted fusion of image features and voiceprint features to form an enhanced fusion feature set.

[0050] Specifically, modal response coupling monitoring is used to identify differences in the consistency of feature distribution between image features and acoustic fingerprint features. Through modal response coupling monitoring, image feature vectors and acoustic fingerprint feature vectors are extracted. The image feature vectors have a dimension of 36, including gradient statistical features, directional distribution features, and cross-channel correlation features; the acoustic fingerprint feature vectors have a dimension of 100, including spectral energy principal components and inter-band correlation principal components. Dimensional alignment is performed on the two feature vectors, expanding the 36-dimensional image feature vectors to 100 dimensions through zero-padding while preserving the original information. The statistical distribution characteristics of each dimension of the image and acoustic fingerprint feature vectors are calculated, including the mean and standard deviation. Under normal transformer operation, the image feature mean is 0.35 and the standard deviation is 0.18, while the acoustic fingerprint feature mean is 0.42 and the standard deviation is 0.22, indicating relatively similar statistical distributions. When the transformer experiences core loosening defects, the image feature value increases to the range of 0.65-0.75 in the gradient statistical dimension, and the acoustic fingerprint feature value increases to the range of 0.70-0.85 in the mid-frequency principal component dimension, showing a shift in the distribution characteristics of the two features. The method calculates a sequence of differences between two vectors, obtained by subtracting the image feature vector from the voiceprint feature vector. The mean, variance, and maximum value of the absolute differences are calculated. The mean of the absolute differences reflects the overall deviation between the two features, the variance reflects the stability of the deviation, and the maximum value indicates the location of the most significant modal difference. The statistical characteristics of the difference sequence are defined as the feature distribution consistency difference, which quantifies the degree of coordination between image features and voiceprint features when responding to the same device state.

[0051] In some embodiments, the step of triggering a weighted modulation mechanism based on the feature distribution consistency difference to perform weighted fusion of the image features and the voiceprint features to form an enhanced fusion feature set includes: evaluating the feature distribution consistency difference by performing a difference measurement to determine the difference level; triggering a weighted modulation strategy based on the difference level, enabling an enhanced modulation mode when the difference level is higher than a preset threshold, and enabling a balanced modulation mode when the difference level is lower than a preset threshold; generating weight coefficients according to the weighted modulation strategy; and using the weight coefficients to perform weighted fusion of the image features and the voiceprint features to form an enhanced fusion feature set.

[0052] The difference in feature distribution consistency is assessed using a difference measurement method to determine the difference level. Three key indicators are extracted from the difference in feature distribution consistency: mean absolute value of the difference, variance of the difference, and maximum difference. A difference measurement function is established, which comprehensively considers the impact of the three indicators on consistency. The calculation formula is: D = 0.5 × M_abs + 0.3 × √V + 0.2 × M_max, where D is the difference measurement value, M_abs is the mean absolute value of the difference, V is the variance of the difference, and M_max is the maximum difference. A larger value indicates a more significant difference between the distribution of image features and voiceprint features. For the normal operation of the transformer, the mean absolute value of the difference is 0.08, the variance of the difference is 0.015, the maximum difference is 0.22, and the measurement value is 0.117. When the transformer has a core loosening defect, the mean absolute value of the difference is 0.25, the variance of the difference is 0.082, the maximum difference is 0.58, and the measurement value is 0.327. A threshold is set to classify the difference levels, mapping the measurement values ​​to four difference levels: a measurement value less than 0.15 is classified as extremely low difference level 1, 0.15 to 0.25 as low difference level 2, 0.25 to 0.35 as medium difference level 3, and 0.35 to 0.50 as high difference level 4. The measurement value of 0.117 for normal transformer operation corresponds to level 1, the measurement value of 0.327 for loose iron core corresponds to level 3, and the measurement value of 0.48 for partial discharge in switchgear corresponds to level 4.

[0053] A weighted modulation strategy is triggered based on the difference level. When the difference level is higher than a preset threshold, an enhanced modulation mode is activated; when the difference level is lower than the preset threshold, a balanced modulation mode is activated. According to the weighted modulation strategy, the preset threshold for the difference level is set to level 3, which divides the difference level into low-consistency and high-difference regions. When the difference level is 1 or 2, it indicates that the distributions of image features and voiceprint features are highly consistent, triggering the balanced modulation mode. The balanced modulation mode assigns similar weights to the two modalities, with a weight ratio set to 0.45-0.55 for image and 0.45-0.55 for voiceprint, ensuring that the information from the two modalities is fused evenly. When the difference level is 3 or 4, it indicates that the distributions of the two features are significantly different, triggering the enhanced modulation mode. The enhanced modulation mode dynamically adjusts the weights according to the reliability of each modal feature, tilting the weights towards the more stable response modality. The internal consistency indices of image features and voiceprint features are analyzed. The internal consistency of image features is evaluated through the correlation between three types of features: gradient statistics, directional distribution, and cross-channel correlation. A correlation coefficient higher than 0.70 indicates good internal consistency. The internal consistency of acoustic signature features is evaluated using the orthogonality between different principal components. An orthogonality higher than 0.85 indicates strong principal component independence and high feature quality. Comparing the internal consistency indices of two modes, the mode with higher internal consistency is considered more reliable and receives a greater weight in the enhanced modulation mode. For the core loosening defect scenario, acoustic signature features are more reliable, and the enhanced modulation mode sets the acoustic signature weight to 0.65-0.75. For the partial discharge scenario, image features are more reliable, and the enhanced modulation mode sets the image weight to 0.70-0.80.

[0054] Weighting coefficients are generated based on a weighted modulation strategy. For the balanced modulation mode, precise weights are determined based on the effectiveness of the feature dimensions of image features and voiceprint features. The number of non-zero effective dimensions in the 36 original dimensions of image features and the number of effective dimensions with a cumulative variance contribution rate exceeding 1% in the 100 dimensionality-reduced dimensions of voiceprint features are counted. Weights are assigned based on the relative proportions of effective dimensions: when the difference in effective proportions between two modes is less than 0.10, equal weights of 0.50 and 0.50 are used; when the difference is in the range of 0.10-0.20, the mode with the higher effective proportion has a weight of 0.55, and the mode with the lower effective proportion has a weight of 0.45. For the enhanced modulation mode, dynamic weights are calculated based on the difference level and the modal reliability index. The weight of the more reliable mode increases with the level: when the difference level is 3, the reliable mode has a weight of 0.60, and the other mode has a weight of 0.40; when the level is 4, the reliable mode has a weight of 0.70, and the other mode has a weight of 0.30. For the core loosening scenario at level 3 and the voiceprint features are reliable, the generated weight coefficients are: image weight 0.40 and voiceprint weight 0.60. For partial discharge scenario level 4 and reliable image features, the generated weight coefficients are 0.70 for image weight and 0.30 for voiceprint weight. The generated weight coefficients are verified to meet the normalization condition, ensuring that the numerical range of the fused feature vector remains stable.

[0055] Weighted fusion of image features and voiceprint features is performed using weighting coefficients to form an enhanced fused feature set. The image feature vector and voiceprint feature vector are weighted and fused dimension-by-dimensionally. The fusion formula is: F_fused(i) = W_img × F_img(i) + W_snd × F_snd(i), where F_fused(i) is the value of the i-th dimension of the fused feature vector, W_img is the image weighting coefficient (dimensionless, ranging from 0 to 1), F_img(i) is the value of the i-th dimension of the image feature vector, W_snd is the voiceprint weighting coefficient (dimensionless, ranging from 0 to 1), and F_snd(i) is the value of the i-th dimension of the voiceprint feature vector, and W_img + W_snd = 1. Weighted fusion is performed one by one across 100 dimensions. Since only the first 36 dimensions of the image features have original values, and the last 64 dimensions are zero-padded, the first 36 dimensions of the fused result integrate information from both modalities, while the last 64 dimensions only contain voiceprint feature information. The fused feature vectors are normalized by rescaling the values ​​of each dimension to the range of 0-1. The normalized vector is the enhanced fused feature set. This enhanced fused feature set has 100 dimensions and integrates the image characteristics of spatial sound pressure distribution and the acoustic characteristics of spectral energy distribution. The fused feature vectors have higher information content, with the defective state exhibiting significantly higher information content than the normal state. The top 20 dimensions with the highest values ​​in the enhanced fused feature set are extracted as the key feature subset, which contains the most distinctive fused features. For core loosening, the key subset includes gradient statistical fused features and mid-frequency principal component fused features, with values ​​ranging from 0.65 to 0.82.

[0056] Step S150: Generate a defect candidate set based on the enhanced fusion feature set, perform competitive screening on the defect candidate set to identify the low-confidence candidate weight distribution, transfer and enhance the low-confidence candidate weight distribution to generate a high-confidence candidate set, and perform defect type determination based on the high-confidence candidate set to complete defect identification.

[0057] Specifically, a defect candidate set is generated based on an enhanced fusion feature set. This enhanced fusion feature set is input into a defect identification model, which is trained using a Support Vector Machine (SVM) algorithm. The training data includes eight typical defect types, such as transformer core loosening, winding deformation, switchgear partial discharge, and poor busbar contact. The model performs multi-class classification on the 100-dimensional feature vector of the enhanced fusion feature set, calculating the confidence score of each feature vector belonging to each defect type. The confidence score is represented in probabilistic form, ranging from 0 to 1, with the sum of the confidence scores for all defect types being 1. For a specific sound signal from a power equipment to be identified, the model outputs a sequence of confidence scores for eight defect types, such as core loosening (0.42), winding deformation (0.25), partial discharge (0.18), poor contact (0.08), and the remaining four defects each below 0.02. A confidence threshold of 0.10 is set, and defect types with confidence scores higher than 0.10 are marked as defect candidates. For the example above, three defect types—core loosening (0.42), winding deformation (0.25), and partial discharge (0.18)—were labeled as candidate items. The confidence score, characteristic response pattern, and corresponding defect type label were extracted for each candidate item. The characteristic response pattern recorded the top 10 dimensions and their values ​​that contributed most to the defect type in the enhanced fusion feature set. All candidate items, along with their confidence scores, response patterns, and labels, were integrated into a defect candidate set. This candidate set contains multiple possible defect interpretations, each corresponding to a potential equipment state.

[0058] A competitive screening process is implemented on the defect candidate set to identify the weight distribution of low-confidence candidates. Through competitive screening, each candidate in the defect candidate set is sorted from highest to lowest confidence score, resulting in the following candidate sequence: [Loose Core 0.42, Winding Deformation 0.25, Partial Discharge 0.18]. The confidence difference between candidates is calculated. The confidence difference between the first and second candidates is 0.17, and the confidence difference between the second and third candidates is 0.07. The confidence difference reflects the intensity of competition between candidates; a larger difference indicates a more significant advantage for the higher-ranked candidate. A competition threshold of 0.15 is set. When the confidence difference between adjacent candidates is less than 0.15, the two candidates are considered to be in competition. Through competitive screening, in the above example, the difference between the first and second candidates (0.17) is greater than the threshold, indicating that the first candidate, Loose Core, is in a dominant position; the difference between the second and third candidates (0.07) is less than the threshold, indicating that Winding Deformation and Partial Discharge are in competition. Candidates with relatively low confidence scores and in a competitive state are identified. These candidates, while exceeding the basic threshold of 0.10, have confidence scores significantly lower than the highest-scoring candidate and are defined as low-confidence candidates. The confidence score distribution features of low-confidence candidates are extracted, including the score value, the difference from the highest score, and the competitive intensity with neighboring candidates. The confidence scores and competitive relationship features of low-confidence candidates are integrated into a low-confidence candidate weight distribution, which describes the weight configuration of low-confidence candidates in the current feature space.

[0059] In some embodiments, the step of transferring and enhancing the low-confidence candidate weight distribution to generate a high-confidence candidate set includes: extracting confidence strength features and confidence weakness features based on the low-confidence candidate weight distribution; performing weight aggregation on the confidence strength features to form a weight transfer source; using the confidence weakness features to perform targeted allocation on the weight transfer source to generate enhanced weights; and generating a high-confidence candidate set based on the enhanced weights.

[0060] Confidence strength features and confidence weakness features are extracted based on the low-confidence candidate weight distribution. The feature response patterns of each low-confidence candidate in the low-confidence candidate weight distribution are analyzed to identify the feature dimensions with strong responses in the enhanced fusion feature set. For the winding deformation candidate, the feature response patterns show that the directional distribution fusion feature values ​​in dimensions 22-28 are in the range of 0.55-0.68, and the frequency band correlation principal component values ​​in dimensions 58-65 are in the range of 0.60-0.72. These dimensions contribute significantly to the identification of winding deformation. For the partial discharge candidate, the high-resolution channel feature values ​​in dimensions 8-11 are in the range of 0.58-0.65, and the high-frequency principal component values ​​in dimensions 75-82 are in the range of 0.62-0.78. The feature dimensions with strong responses and high values ​​for each low-confidence candidate are defined as confidence strength features; these feature dimensions are the main source of candidate confidence. We identified feature dimensions in the enhanced fusion feature set that showed weak or missing responses for each low-confidence candidate. These dimensions are valuable for identifying this defect type but are insufficiently represented in the current feature set. By comparing with typical feature templates for this defect type, we found that winding deformation responded weakly in the gradient-acoustic crossover feature of dimensions 35-40, with values ​​of only 0.28-0.35, while typical winding deformation samples should have values ​​of 0.50-0.65 in these dimensions. Partial discharge responded insufficiently in the mid-frequency transition feature of dimensions 48-52, with values ​​of 0.32-0.38, while typical values ​​should be 0.55-0.70. Feature dimensions with weak responses and values ​​below typical values ​​were defined as confidence weakness features. The insufficiency of these features leads to low confidence in the candidate.

[0061] A weight transfer source is formed by weighting and aggregating the confidence strength features. Through weight aggregation, the confidence strength feature dimensions and their values ​​for each low-confidence candidate are extracted. For winding deformation, the confidence strength features include 14 dimensions with values ​​ranging from 0.55 to 0.72. The weight contribution value corresponding to each confidence strength feature dimension is calculated; the weight contribution value is equal to the ratio of the dimension value to the total confidence of the candidate. The weight contribution values ​​of each confidence strength feature dimension are statistically analyzed, and high-contribution dimensions with weight contribution values ​​exceeding 2.0 are identified. The weights of the high-contribution dimensions are aggregated and accumulated; the total aggregated weight value represents the surplus weight resources currently possessed by the candidate. For winding deformation, 8 of the 14 confidence strength feature dimensions have weight contribution values ​​exceeding 2.0, with a total aggregated weight of 18.5. For partial discharge, 6 of the 12 dimensions exceed 2.0, with a total aggregated weight of 14.2. The total aggregated weight value of all low-confidence candidates is defined as the weight transfer source, representing the weight capacity that can be extracted from strong response features and transferred to weak response features. Record the weight transfer source values ​​and corresponding high contribution dimension information for each candidate.

[0062] For example, the step of using the confidence weakness feature to target and allocate weight transfer sources to generate enhanced weights includes: constructing a weight demand map based on the confidence weakness feature; performing intensity-level mapping on the weight demand map to form a hierarchical weight channel; monitoring the transferable weight capacity in the weight transfer source to determine an allocation threshold; and targeting and allocating weight transfer sources to generate enhanced weights based on the allocation threshold through the hierarchical weight channel.

[0063] A weighted demand map is constructed based on confidence weakness features. The confidence weakness feature dimensions of each low-confidence candidate are extracted, along with the difference between their current and typical values. For winding deformation, there are 6 confidence weakness feature dimensions (35-40), with current values ​​of 0.28-0.35 and typical values ​​of 0.50-0.65, resulting in a difference of 0.15-0.30. For partial discharge, there are 5 confidence weakness feature dimensions (48-52), with current values ​​of 0.32-0.38 and typical values ​​of 0.55-0.70, resulting in a difference of 0.17-0.38. The weight demand for each confidence weakness feature dimension is calculated. The weight demand equals the difference in value for that dimension multiplied by the importance coefficient of that dimension in defect identification. The importance coefficient is determined by statistically analyzing the contribution of that dimension to the classification accuracy of that defect type in the training samples. The weight demand for all confidence weakness feature dimensions of each candidate is statistically analyzed, and a weight demand distribution map is plotted. The horizontal axis of this distribution chart represents the feature dimension index, and the vertical axis represents the weight requirement. The chart labels the current value, typical value, and required amount for each dimension. This weight requirement distribution chart is defined as a weight requirement map, which visually shows which feature dimensions each low-confidence candidate lacks weight on and how much weight is needed to approximate the typical pattern.

[0064] A hierarchical weighted channel is formed by performing intensity-level mapping on the weighted demand map. Based on the magnitude of the weight demand in each dimension of the weighted demand map, the demand is divided into three intensity levels: high demand, medium demand, and low demand. Dimensions with a demand greater than 0.50 are marked as high demand, dimensions between 0.30 and 0.50 as medium demand, and dimensions below 0.30 as low demand. For the confidence weakness features of winding deformation, some dimensions belong to the high demand level, some to the medium demand level, and some to the low demand level. An independent weight transmission channel is established for each intensity level: high demand corresponds to a high-priority channel, medium demand to a medium-priority channel, and low demand to a low-priority channel. The weight allocation ratio for each channel is set according to priority: 50% for high priority, 30% for medium priority, and 20% for low priority. Each confidence weakness feature dimension is assigned to its corresponding weight channel according to its intensity level, establishing a hierarchical weighted channel structure through intensity-level mapping. The hierarchical weighted channel contains three parallel channels. Each channel records the list of feature dimensions assigned to that channel, the total demand, and the allocation ratio.

[0065] The allocation threshold is determined by monitoring the transferable weight capacity in the weight transfer sources. The actual transferable weight capacity in the weight transfer sources of each low-confidence candidate is evaluated. Not all weights in the weight transfer sources are transferable; some weights need to be retained to maintain the response level of the original confidence strength characteristics. For example, the confidence strength characteristics of the winding deformation candidate are prominent in the directional distribution and frequency band correlation dimensions. If all weights are transferred out, the identification ability of these strong response characteristics will be weakened. The transferable weight capacity W_available = W_total × (1-α) is calculated, where W_total is the total weight of the weight transfer sources, and α is a retention coefficient with a value of 0.40. The total demand D_total in the graded weight channels of each candidate is monitored, and the relationship between the transferable weight capacity and the total demand is compared to determine the allocation threshold. The allocation threshold is calculated using the formula T_alloc = min(W_available, D_total), where T_alloc is the allocation threshold, and the min function takes the smaller of the two values. When transferable capacity is sufficient, such as when the excess weight of winding deformation far exceeds the demand for weakness features, the allocation threshold is set to the total demand, which can fully satisfy the weight supplementation for weakness features. When transferable capacity is insufficient, such as when the excess weight of partial discharge cannot cover all demand, the allocation threshold is limited by available capacity, and limited resources must be allocated among various weakness features according to priority. This mechanism ensures that weight transfer can both improve low-confidence candidates and not weaken the identification ability of the original confidence strength features.

[0066] Enhanced weights are generated by targeted allocation of weight transfer sources through hierarchical weight channels based on an allocation threshold. The actual weight allocation received by each channel is calculated according to the allocation threshold and the allocation ratio of the hierarchical weight channels. The allocation threshold determines the total amount of weight available for allocation. The hierarchical weight channels allocate these weights to high, medium, and low priority channels, with allocation ratios of 50%, 30%, and 20% respectively, ensuring that high-demand weakness features receive priority weight supplementation. Taking winding deformation candidates as an example, their confidence weakness features are distributed across multiple dimensions of gradient-acoustic cross-features. Some dimensions with significant differences from typical values ​​are assigned to high-priority channels, while others with moderate differences are assigned to medium-priority channels. Through targeted allocation, the weights allocated to each channel are further subdivided into specific dimensions according to the demand ratio of each feature dimension within the channel. For a specific weakness dimension within a high-priority channel, a corresponding share of weight is obtained from the channel allocation based on its proportion of the channel's total demand. The allocated weight is then directly added to the current value of the corresponding confidence weakness feature dimension to obtain the enhanced value. For example, if the current value of winding deformation in a certain gradient-acoustic cross-feature dimension is low, this dimension is identified as a confidence weakness and belongs to a high-priority channel. A certain amount of weight is added through targeted allocation, and after these weights are added to the current value, the value of this dimension is significantly improved and approaches the performance level of typical winding deformation samples in this dimension. The same weight enhancement operation is performed on all confidence weakness feature dimensions, and the enhanced feature dimension values ​​approach or reach the typical value level. The enhanced feature dimension values ​​are defined as enhancement weights. These enhancement weights compensate for the weaknesses of the original features and improve the overall confidence of the candidates.

[0067] A high-confidence candidate set is generated based on enhanced weights. The confidence weakness feature dimension values ​​of each low-confidence candidate are replaced with the corresponding enhanced weight values, while the confidence strength feature dimension values ​​remain unchanged. For winding deformation candidates, the confidence weakness feature values ​​in dimensions 35-40 are updated from 0.28-0.35 to 0.67-0.80, while the confidence strength feature dimensions in dimensions 22-28 and 58-65 remain unchanged. The updated feature vectors are re-input into the defect identification model, and the enhanced confidence score is calculated. The confidence score of the winding deformation candidate increases from 0.25 to 0.58, and the partial discharge increases from 0.18 to 0.45. A high-confidence threshold of 0.50 is set, and candidates with enhanced confidence scores exceeding 0.50 are marked as high-confidence candidates. A winding deformation confidence score of 0.58 exceeds the threshold and is marked as a high-confidence candidate; a partial discharge confidence score of 0.45 does not exceed the threshold and remains a low-confidence candidate. Extract the defect type labels, enhanced confidence scores, and feature response patterns of all high-confidence candidates, and integrate these candidates into a high-confidence candidate set.

[0068] Defect identification is completed by performing defect type determination based on a high-confidence candidate set. Through defect type determination, candidates in the high-confidence candidate set are sorted according to their confidence scores, and the candidate with the highest confidence score is selected as the final defect determination result. In the example above, the winding deformation confidence score of 0.58 is the highest. Although the core loosening initially had a confidence score of 0.42, it did not enter the low-confidence transfer enhancement process and maintained its original score. After comparison, the winding deformation score of 0.58 is higher than the core loosening score of 0.42, and the final defect type is determined to be winding deformation. The defect type label, confidence score, and key feature dimensions supporting this determination are output. The key features of winding deformation include the enhanced gradient-acoustic cross feature in dimensions 35-40, the directional distribution fusion feature in dimensions 22-28, and the frequency band associated principal components in dimensions 58-65. The complete process information of defect identification is recorded, including the initial defect candidate set, the low-confidence candidate weight distribution, the weight transfer enhancement operation, the high-confidence candidate set, and the final determination result. By determining the type of defect, the defect identification process is completed, and a structured identification report is generated for equipment maintenance personnel to refer to.

[0069] To implement the power equipment defect identification method based on acoustic imaging and acoustic signature fusion corresponding to the above method embodiments, and to achieve the corresponding functions and technical effects. See also Figure 2 , Figure 2 This diagram illustrates the structural block diagram of a power equipment defect identification device 200 based on acoustic imaging and acoustic fingerprint fusion provided in an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The power equipment defect identification device 200 based on acoustic imaging and acoustic fingerprint fusion provided in this embodiment includes: The sound acquisition module 201 is used to acquire the sound signal of the power equipment, perform frequency domain integrity verification on the sound signal to identify the missing spectrum segment, and use the missing spectrum segment to perform acoustic feature repair analysis to generate a complete acoustic stream. The acoustic imaging processing module 202 is used to perform acoustic imaging conversion analysis on the complete acoustic flow to extract the spatial sound pressure distribution map, establish a multi-scale image feature channel using the spatial sound pressure distribution map, and perform spatial gradient analysis on the multi-scale image feature channel to generate image features. The voiceprint extraction module 203 is used to perform time-frequency decomposition processing on the complete acoustic flow to extract the spectral energy distribution matrix, identify the silent frequency band intervals in the spectral energy distribution matrix, use the silent frequency band intervals to segment the spectral energy to generate segmented voiceprint feature clusters, and perform cross-frequency band correlation analysis on the segmented voiceprint feature clusters to generate voiceprint features. Feature fusion module 204 is used to perform modal response coupling monitoring and identification of feature distribution consistency differences between the image features and the voiceprint features, and trigger a weight modulation mechanism based on the feature distribution consistency differences to perform weighted fusion of the image features and the voiceprint features to form an enhanced fusion feature set; The defect identification module 205 is used to generate a defect candidate set based on the enhanced fusion feature set, perform competitive screening on the defect candidate set to identify the low-confidence candidate weight distribution, transfer and enhance the low-confidence candidate weight distribution to generate a high-confidence candidate set, and perform defect type determination based on the high-confidence candidate set to complete defect identification.

[0070] The aforementioned power equipment defect identification device 200 based on acoustic imaging and voiceprint fusion can implement the power equipment defect identification method based on acoustic imaging and voiceprint fusion described in the above method embodiments. The options in the above method embodiments are also applicable to this embodiment, and will not be detailed here. The remaining content of this application embodiment can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.

[0071] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.

[0072] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.

Claims

1. A method for identifying defects in power equipment based on acoustic imaging and acoustic signature fusion, characterized in that, include: Collect the operating sound signal of the power equipment, perform frequency domain integrity verification on the operating sound signal to identify the missing spectrum segment, and use the missing spectrum segment to perform acoustic feature repair analysis to generate a complete acoustic flow; The complete acoustic flow is subjected to acoustic imaging conversion analysis to extract a spatial sound pressure distribution map. The spatial sound pressure distribution map is used to establish a multi-scale image feature channel. Spatial gradient analysis is performed on the multi-scale image feature channel to generate image features. The complete acoustic flow is subjected to time-frequency decomposition to extract the spectral energy distribution matrix. The silent frequency band intervals in the spectral energy distribution matrix are identified. The spectral energy is segmented and segmented using the silent frequency band intervals to generate segmented acoustic signature clusters. Cross-band correlation analysis is performed on the segmented acoustic signature clusters to generate acoustic signature features. Modal response coupling monitoring is performed on the image features and the voiceprint features to identify differences in feature distribution consistency. Based on the differences in feature distribution consistency, a weighted modulation mechanism is triggered to perform weighted fusion of the image features and the voiceprint features to form an enhanced fusion feature set. Based on the enhanced fusion feature set, a defect candidate set is generated. The defect candidate set is then subjected to competitive screening to identify the low-confidence candidate weight distribution. The low-confidence candidate weight distribution is then transferred and enhanced to generate a high-confidence candidate set. Based on the high-confidence candidate set, a defect type determination is performed to complete the defect identification.

2. The method according to claim 1, characterized in that, The process of generating a complete acoustic flow by performing acoustic feature restoration analysis using the missing spectral segments includes: Locate energy attenuation nodes in the spectrum missing segment and simultaneously obtain compensation energy references; Based on the energy attenuation node and the compensation energy reference, a missing-compensation coupling point is generated by frequency domain correlation. The missing-compensation coupling point is used to excite the spectral repair response to generate a compensated spectral signal; A complete acoustic flow is constructed based on the compensated spectral signal.

3. The method according to claim 1, characterized in that, The step of establishing multi-scale image feature channels using the spatial sound pressure distribution map includes: Obtain the peak sound pressure distribution characteristics of the spatial sound pressure distribution map; Locate the energy transition zone in the spatial sound pressure distribution map; Based on the energy transition band interval, the peak sound pressure distribution characteristics are scaled to generate a layered sound pressure signal; The layered sound pressure signal is processed by channelization to establish multi-scale image feature channels.

4. The method according to claim 1, characterized in that, The step of segmenting the spectral energy into segments to generate segmented acoustic signature clusters using the silent frequency band includes: Based on the silent frequency band interval, locate and identify segmentation points at the spectrum boundary; The aforementioned dividing points are used to partition the spectral energy region to form frequency band subsets; The frequency band subset is subjected to voiceprint feature extraction to generate feature segments; Clustering and integration are performed on the feature fragments to construct segmented voiceprint feature clusters.

5. The method according to claim 1, characterized in that, The step of performing cross-band correlation analysis on the segmented voiceprint feature clusters to generate voiceprint features includes: The frequency band similarity of the segmented voiceprint feature clusters is calculated to generate an association strength matrix; Weak correlation intervals are extracted from the correlation strength matrix as frequency band difference boundaries; Based on the frequency band difference boundary, the segmented voiceprint feature clusters are differentially recombined to form a cross-frequency band feature chain; The cross-frequency band feature chain is vectorized and reconstructed to generate voiceprint features.

6. The method according to claim 1, characterized in that, The weighted fusion of the image features and the voiceprint features based on the feature distribution consistency difference triggering weight modulation mechanism to form an enhanced fusion feature set includes: The differences in the consistency of the feature distribution are assessed by a difference measurement to determine the difference level; Based on the difference level triggering weighted modulation strategy, when the difference level is higher than a preset threshold, the enhanced modulation mode is activated, and when the difference level is lower than the preset threshold, the balanced modulation mode is activated. Weight coefficients are generated according to the weight modulation strategy; The image features and the voiceprint features are weighted and fused using the weighting coefficients to form an enhanced fusion feature set.

7. The method according to claim 1, characterized in that, The step of transferring and enhancing the low-confidence candidate weight distribution to generate a high-confidence candidate set includes: Confidence strength features and confidence weakness features are extracted based on the low-confidence candidate weight distribution; The confidence strength features are weighted and aggregated to form a weight transfer source; The confidence weakness features are used to target and allocate the weight transfer source to generate enhanced weights. A high-confidence candidate set is generated based on the enhanced weights.

8. The method according to claim 5, characterized in that, Extracting weakly correlated intervals from the correlation strength matrix as frequency band difference boundaries includes: The correlation strength matrix is ​​segmented using a threshold to identify the distribution of low-intensity elements; The low-intensity element distribution is mapped to a frequency band coordinate system to generate weakly correlated position markers; The weakly correlated position markers are merged into intervals to form continuous weakly correlated bands; The boundary nodes of the continuous weakly correlated bands are selected as the frequency band difference boundaries.

9. The method according to claim 7, characterized in that, The step of using the confidence weakness features to target and allocate weight transfer sources to generate enhanced weights includes: Construct a weighted demand graph based on the aforementioned confidence weakness features; The weight demand map is subjected to intensity-level mapping to form a hierarchical weight channel; The allocation threshold is determined by monitoring the transferable weight capacity in the weight transfer source. Based on the allocation threshold, the weight transfer source is targeted and enhanced weights are generated through the hierarchical weight channel.

10. A power equipment defect identification device based on acoustic imaging and acoustic fingerprint fusion, characterized in that, include: The sound acquisition module is used to acquire the sound signals of power equipment operation, perform frequency domain integrity verification on the sound signals to identify missing spectral segments, and use the missing spectral segments to perform acoustic feature repair analysis to generate a complete acoustic stream; The acoustic imaging processing module is used to perform acoustic imaging conversion analysis on the complete acoustic flow to extract a spatial sound pressure distribution map, establish a multi-scale image feature channel using the spatial sound pressure distribution map, and perform spatial gradient analysis on the multi-scale image feature channel to generate image features. The voiceprint extraction module is used to perform time-frequency decomposition processing on the complete acoustic stream to extract the spectral energy distribution matrix, identify the silent frequency band intervals in the spectral energy distribution matrix, use the silent frequency band intervals to perform spectral energy segmentation to generate segmented voiceprint feature clusters, and perform cross-frequency band correlation analysis on the segmented voiceprint feature clusters to generate voiceprint features. The feature fusion module is used to perform modal response coupling monitoring and identification of feature distribution consistency differences between the image features and the voiceprint features, and trigger a weight modulation mechanism based on the feature distribution consistency differences to perform weighted fusion of the image features and the voiceprint features to form an enhanced fusion feature set; The defect identification module is used to generate a defect candidate set based on the enhanced fusion feature set, perform competitive screening on the defect candidate set to identify low-confidence candidate weight distribution, transfer and enhance the low-confidence candidate weight distribution to generate a high-confidence candidate set, and perform defect type determination based on the high-confidence candidate set to complete defect identification.