Far-field sound source localization method applied to transformer substation
By converting the target frequency sound in the substation into a spherical harmonic domain multi-channel observation vector and generating a spherical harmonic domain mask matrix using a preset network model, and constructing an in-band smooth covariance, the problem of insufficient sound source localization accuracy in the substation is solved, and accurate localization is achieved.
Patent Information
- Application Number
- CN202511761255.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-27
- Publication Date
- 2026-03-06
AI Technical Summary
Existing sound source localization methods are difficult to effectively filter out direct target sound source signals in complex situations such as strong reflections, multiple high-frequency noise sources, and multiple sound sources in substations. This leads to the occurrence of false peaks during the localization process, causing sound source localization errors.
The target frequency sound in the substation is converted into a spherical harmonic domain multi-channel observation vector. A spherical harmonic domain mask matrix is generated using a preset network model. A spatial spectrum matrix is constructed by in-band smooth covariance to distinguish multiple sound sources, obtain a set of multi-source directions, and realize far-field sound source localization.
It effectively suppresses noise and interference in the complex environment of substations, reduces the generation of spurious peaks, achieves accurate localization of sound sources within substations, and solves the problems of noise interference and difficulty in distinguishing multiple sound sources.
Smart Images

Figure CN121613403A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound source localization technology, and in particular to a far-field sound source localization method applied to substations. Background Technology
[0002] Equipment in a substation (such as transformers, switches, and instrument transformers) generates stable sound signals during normal operation. However, when a fault occurs, it is accompanied by high-frequency sound source signals of a specific frequency. These fault-related high-frequency sound source signals are a direct reflection of the abnormal state of the equipment. Since the fault sound source in a substation is often located inside the equipment or in a distant area, far-field sound source localization technology is usually used to capture and locate the location of the distant fault source without having to come into close contact with the high-voltage equipment.
[0003] Traditional sound source localization methods typically employ various microphone arrays to acquire sound source signals. By extracting features such as signal time difference and sound level difference, and combining these with techniques like generalized cross-correlation phase transform and minimum variance distortionless response beamforming, a spatial spectrum or spatial covariance matrix is constructed. The direction of the sound source is then determined based on the peak value. However, the strong reflections, multiple high-frequency noise sources, and the simultaneous presence of multiple sound sources within substations make it difficult for existing methods to effectively filter out signals directly reaching the target sound source. This leads to the occurrence of false peaks during the localization process, causing sound source localization errors. Summary of the Invention
[0004] This invention provides a far-field sound source localization method for substations to solve the above-mentioned problems.
[0005] This invention provides a far-field sound source localization method applied to substations, comprising: Acquire the target frequency sound within the substation and transform the target frequency sound into a spherical harmonic domain multi-channel observation vector; The spherical harmonic domain multi-channel observation vector is input into a preset network model to obtain a spherical harmonic domain mask matrix, and based on the spherical harmonic domain mask matrix, the in-band smooth covariance is obtained; the preset network model is used to filter the spherical harmonic domain multi-channel observation vector. Based on the in-band smooth covariance, a spatial spectrum matrix is obtained, which includes spatial spectrum values in different candidate sound source directions. The spatial spectrum matrix is divided into multiple sound sources to obtain a set of multiple source directions, thereby realizing the far-field sound source localization of the substation. The set of multiple source directions includes the polar angle and azimuth coordinates corresponding to the direction of each target sound source, and the set of multiple source directions is used to reflect the direction of sound sources within the substation.
[0006] According to certain embodiments of the present invention, a method for acquiring target frequency sound within a substation and transforming the target frequency sound into a spherical harmonic domain multi-channel observation vector includes: Inside the substation, acquisition devices are arranged in a spherical array to acquire the target frequency sound. The target frequency sound is subjected to short-time Fourier transform and spherical harmonic transform to obtain the spherical harmonic domain multi-channel observation vector.
[0007] According to certain embodiments of the present invention, the preset network model has a built-in embedded beamforming network; the spherical harmonic domain multi-channel observation vector is input into the preset network model to obtain a spherical harmonic domain mask matrix, including: Obtain the spherical harmonic domain multi-channel observation vector, and perform a concatenation operation on the real and imaginary parts of the spherical harmonic domain multi-channel observation vector according to their channel dimensions to obtain the concatenated feature after concatenation of the real and imaginary parts; The splicing features are input into a preset embedding-beamforming network to obtain the spherical harmonic domain mask matrix.
[0008] According to certain embodiments of the present invention, based on the spherical harmonic domain mask matrix, the in-band smooth covariance is obtained, including: Based on element-wise multiplication, the spherical harmonic domain mask matrix and the spherical harmonic domain multi-channel observation vector are operated to obtain the spherical harmonic domain observation matrix. Perform time-frequency second-order statistics on the spherical harmonic domain observation matrix to obtain the spatial covariance matrix; Based on the spatial covariance matrix, the in-band smoothing covariance is obtained.
[0009] According to the technical solutions provided in certain embodiments of the present invention, the in-band smoothing covariance is obtained based on the spatial covariance matrix, including: Acquire the frequency acoustic emission bands of equipment in the substation during abnormal conditions; The spatial covariance matrix is processed based on the frequency acoustic emission band to obtain the in-band smooth covariance.
[0010] According to the technical solutions provided in certain embodiments of the present invention, the spatial spectral matrix is obtained based on the in-band smoothing covariance, including: Multiple candidate sound source directions are acquired on the spherical distributed array of the acquisition device; Based on the candidate sound source directions, calculate the column vector of spherical harmonic functions corresponding to each candidate sound source direction, and stack them into a scan matrix; The spatial spectrum matrix is calculated based on the scan matrix and the in-band smoothing covariance.
[0011] According to the technical solutions provided in certain embodiments of the present invention, the spatial spectrum matrix is subjected to multi-source differentiation to obtain a multi-source direction set, including: Using the spatial spectrum matrix as the sound source search range, the candidate sound source direction corresponding to the maximum value of the spatial spectrum value in the spatial spectrum matrix is taken as the first target sound source direction; Based on the preset angle threshold and the direction of the first target sound source, the search range of the next sound source is determined, and the directions of the remaining target sound sources are searched within the search range of the next sound source until all the directions of the target sound sources are determined to form the multi-source direction set.
[0012] According to the technical solutions provided in certain embodiments of the present invention, based on a preset angle threshold and the direction of the first target sound source, the search range of the next sound source is determined, and the directions of other target sound sources are searched within the search range of the next sound source, including: Calculate the angle between the candidate sound source direction corresponding to other spatial spectrum values in the spatial spectrum matrix and the first target sound source direction; The directions of all candidate sound sources that satisfy the preset angle threshold in the spatial spectrum matrix are obtained as the next sound source search range; Within the next sound source search range, the candidate sound source direction corresponding to the maximum value of the spatial spectrum value in the spatial spectrum matrix is taken as the next target sound source direction, and the remaining target sound source directions are searched with the preset angle threshold as a constraint.
[0013] In summary, this invention provides a far-field sound source localization method for substations, comprising: acquiring target frequency sound within the substation and transforming the target frequency sound into a spherical harmonic domain (SHDN) multi-channel observation vector; inputting the SHDN multi-channel observation vector into a preset network model to obtain a SHDN mask matrix, and obtaining an in-band smoothing covariance based on the SHDN mask matrix; the preset network model being used to filter the SHDN multi-channel observation vector; obtaining a spatial spectrum matrix based on the in-band smoothing covariance, the spatial spectrum matrix including spatial spectrum values in different candidate sound source directions; performing multi-source differentiation on the spatial spectrum matrix to obtain a multi-source direction set, thereby achieving far-field sound source localization in the substation; the multi-source direction set including the polar angle and azimuth coordinates corresponding to each target sound source direction, the multi-source direction set being used to reflect the sound source directions within the substation.
[0014] This invention converts target frequency sound within a substation into a spherical harmonic domain (SHDN) multi-channel observation vector. Then, a pre-defined network model (used for filtering the SHDN multi-channel observation vector) generates a SHDN mask matrix. This mask matrix is used to filter out effective target signals, reducing the impact of noise and interference in the complex substation environment and minimizing the generation of spurious peaks during localization. This results in an in-band smooth covariance that accurately characterizes the spatial features of the sound source. Based on this in-band smooth covariance, a spatial spectrum matrix containing spatial spectral values of different candidate sound source directions is constructed. This spatial spectrum matrix is then used to distinguish multiple sound sources, extracting a set of multi-source directions containing the polar and azimuth coordinates of each target sound source. This effectively solves the problems of noise interference and difficulty in distinguishing multiple sound sources within substations, achieving accurate localization of sound sources within substations.
[0015] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this invention do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 A flowchart illustrating a far-field sound source localization method applied to a substation, provided by an embodiment of the present invention; Figure 2 A schematic diagram showing the unfolded flow of step S100 of the far-field sound source localization method provided in the embodiment of the present invention; Figure 3 A schematic diagram showing the unfolded flow of step S200 of the far-field sound source localization method provided in the embodiment of the present invention; Figure 4 A flowchart illustrating step S205 of the far-field sound source localization method provided in this embodiment of the invention; Figure 5 A schematic diagram showing the unfolded flow of step S300 of the far-field sound source localization method provided in an embodiment of the present invention; Figure 6 A schematic diagram showing the unfolded flow of step S400 of the far-field sound source localization method provided in the embodiment of the present invention; Figure 7 A schematic diagram showing the unfolded flow of step S402 of the far-field sound source localization method provided in the embodiment of the present invention; Figure 8 This is a schematic diagram of the experimental results of Experiment 1 provided in an embodiment of the present invention; Figure 9 This is a schematic diagram of the experimental results of Experiment 2 provided in an embodiment of the present invention; Figure 10 This is a schematic diagram of the experimental results of Experiment 3 provided in an embodiment of the present invention. Detailed Implementation
[0018] To enable those skilled in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. This description is merely illustrative and explanatory, and should not be construed as limiting the scope of protection of the present invention in any way. Specifically, the described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort should fall within the scope of protection of the present invention.
[0019] It should be noted that similar reference numerals and letters in the following figures denote similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.
[0020] To make the technical solutions of the embodiments of the present invention clearer and easier to understand, the application background of the embodiments of the present invention will be introduced below.
[0021] Transformers, switches, and instrument transformers in substations emit stable sound signals under normal operating conditions. However, when these devices experience faults such as partial discharge, insulation aging, or mechanical wear, they often generate specific frequency sound signals (e.g., ultrasonic signals corresponding to partial discharge). These frequency sound signals, directly related to the fault, are a direct indication of the abnormal state of the equipment. By detecting and locating these signals, it is possible to accurately determine whether there are potential faults in the equipment, providing a reliable basis for equipment condition monitoring. Fault sound sources in substations (such as transformer partial discharge or abnormal switch noises) are often located inside the equipment or in distant areas. Therefore, far-field sound source localization technology is commonly used. This allows for the capture and location of distant fault sources without close contact with high-voltage equipment. This avoids the safety risks of manual inspections near high-voltage equipment and enables rapid identification of the equipment corresponding to abnormal sound sources in complex sound field environments, providing efficient support for remote monitoring and fault early warning.
[0022] Existing sound source localization technologies are mainly divided into two major directions: those based on traditional signal processing and those based on deep learning. Methods based on traditional signal processing often employ various microphone arrays, including planar arrays, three-element arrays, and some spherical arrays. They extract features such as signal time difference and sound level difference, and combine these with techniques like generalized cross-correlation phase transform and minimum variance distortionless response beamforming to construct a spatial spectrum or spatial covariance matrix. Peak detection is then used to determine the sound source direction. Some methods introduce frequency smoothing or spherical harmonic domain conversion to optimize spatial information representation. However, these localization methods struggle to effectively distinguish between direct and reflected sound, and between the target sound source and environmental noise, leading to spurious peaks in the constructed spatial spectrum and a significant decrease in localization accuracy. Even spherical harmonic domain correlation processing methods lack dynamic interference suppression mechanisms, resulting in insufficient spatial resolution in strongly reflective environments. Deep learning-based methods, on the other hand, rely on models such as convolutional neural networks and long short-term memory networks. They use short-time Fourier transform features or generalized cross-correlation phase transform features as input to directly map the sound source direction, or learn beamforming weights to achieve noise suppression and localization. However, such methods require large-scale labeled spatial acoustic data, which is costly to obtain and demands significant computing resources. Furthermore, they have limited adaptability to the complex environment within substations.
[0023] In view of this, please refer to Figure 1 This embodiment provides a far-field sound source localization method applied to substations, including: acquiring target frequency sound within the substation and transforming the target frequency sound into a spherical harmonic domain multi-channel observation vector; inputting the spherical harmonic domain multi-channel observation vector into a preset network model to obtain a spherical harmonic domain mask matrix, and obtaining an in-band smoothing covariance based on the spherical harmonic domain mask matrix; the preset network model is used to filter the spherical harmonic domain multi-channel observation vector; obtaining a spatial spectrum matrix based on the in-band smoothing covariance, the spatial spectrum matrix including spatial spectrum values in different candidate sound source directions; distinguishing multiple sound sources in the spatial spectrum matrix to obtain a multi-source direction set, thereby realizing far-field sound source localization in the substation; the multi-source direction set includes the polar angle and azimuth coordinates corresponding to each target sound source direction, and the multi-source direction set is used to reflect the sound source directions within the substation.
[0024] As can be seen, this invention captures spatial acoustic information completely by converting target frequency sound within a substation into a spherical harmonic domain multi-channel observation vector. A spherical harmonic domain mask matrix is generated through filtering using a pre-defined network model. This mask matrix can selectively filter out effective signal components in the target frequency sound, suppressing interference from reflected sound, equipment background noise, and other factors in the substation environment. The in-band smooth covariance is then calculated based on the signal after denoising using the mask matrix, thereby constructing a spatial spectrum matrix and distinguishing multiple sound sources. This method solves the problem of insufficient sound source localization accuracy in the complex environment of substations in existing technologies. It avoids spatial spectrum spurious peaks caused by the difficulty in distinguishing between direct sound and reflected sound, and between target sound sources and environmental noise, as in traditional signal processing methods, and the dependence of deep learning methods on large-scale labeled data and high computational resources. Ultimately, it achieves accurate localization of far-field sound sources in substations, providing reliable support for equipment fault monitoring.
[0025] Please refer to the following. Figure 1 This embodiment provides a schematic flowchart of a far-field sound source localization method applied to a substation. The execution subject of this embodiment can be a substation equipment status monitoring system. The following is a further explanation of each step of this invention. The sound source localization method includes the following steps: S100: Acquire the target frequency sound within the substation and transform the target frequency sound into a spherical harmonic domain multi-channel observation vector; The target frequency sound refers to all kinds of sound signals radiated during the operation of all equipment in the substation. It includes both the stable frequency sound generated during normal operation of the equipment and the specific frequency sound signals (such as the ultrasonic signal corresponding to partial discharge) that accompany abnormalities such as partial discharge, insulation aging, and mechanical wear of the equipment.
[0026] Transforming the target frequency sound into a spherical harmonic domain multi-channel observation vector is to obtain the basic observations shared by all subsequent steps. On the one hand, the spherical harmonic domain can completely preserve the three-dimensional spatial information of the target frequency sound, characterize the phase and amplitude relationship of sound sources in different directions, and adapt to the propagation path and reflection characteristics of high-frequency sound signals in substations. On the other hand, the spherical harmonic domain has natural signal separation properties, which can initially distinguish the reverberation, environmental noise and other interference signals from the target sound source signal in the frequency-spatial dimension, reducing the difficulty of subsequent processing. At the same time, the spherical harmonic domain can transform the target frequency sound signal into a standardized dimension observation vector, ensuring that all subsequent steps can be carried out based on a unified basic observation, ensuring the continuity and data consistency of the entire positioning process, and ultimately laying the foundation for improving the accuracy of far-field sound source positioning in complex substation environments.
[0027] Specifically, such as Figure 2 As shown, S100 includes the following steps: S101: Inside the substation, the acquisition devices are arranged in a spherical array to acquire the target frequency sound; S102: Perform short-time Fourier transform and spherical harmonic transform on the target frequency sound to obtain the multi-channel observation vector in the spherical harmonic domain.
[0028] The target frequency sound of equipment in a substation may originate from any three-dimensional direction, and the presence of numerous reflective surfaces such as equipment casings and walls leads to a complex sound field. A spherical distributed array, by arranging multiple acquisition units (i.e., multiple channels) on a spherical surface, can more comprehensively capture target frequency sound from different directions, avoiding the omission of crucial spatial information due to array layout limitations. Short-time Fourier transform decomposes the continuous time-domain target frequency sound of multiple channels into two-dimensional time-frequency domain signals, extracting the sound source signal features at different time frames (t) and different frequency points (f). Spherical harmonic transform, based on the time-frequency domain processing of short-time Fourier transform, maps the sound source signal features at different time frames (t) and different frequency points (f) to the spherical harmonic domain, obtaining a multi-channel observation vector in the spherical harmonic domain.
[0029] Specifically, the process of obtaining the multi-channel observation vector in the spherical harmonic domain by the target frequency sound through short-time Fourier transform and spherical harmonic transform is shown in the following formula (1): Formula (1); Where X(t,f) is the spherical harmonic domain multi-channel observation vector (dimension C×1); t is the time frame index; f is the frequency index; A is the spherical harmonic domain steering matrix; s(t,f) is the sound source signal vector; and v(t,f) is the equivalent noise vector.
[0030] S200: Input the spherical harmonic domain multi-channel observation vector into the preset network model to obtain the spherical harmonic domain mask matrix, and obtain the in-band smooth covariance based on the spherical harmonic domain mask matrix; the preset network model is used to filter the spherical harmonic domain multi-channel observation vector. The preset network model is used to filter the multi-channel observation vectors in the spherical harmonic domain. The filtering process is used to select target sound source signals and suppress reverberation noise. The preset network model can be an embedded beamforming network or other deep learning networks with dynamic feature learning capabilities. By learning signal features in different dimensions, it can distinguish between direct sound and interference signals, thereby achieving the effect of selecting target sound source signals and suppressing reverberation noise.
[0031] In this embodiment of the invention, by inputting the spherical harmonic domain multi-channel observation vector into a preset network model, a spherical harmonic domain mask matrix can be obtained. This spherical harmonic domain mask matrix is used to filter out the target sound source portion in the spherical harmonic domain multi-channel observation vector and suppress reverberation and noise. At the same time, based on the spherical harmonic domain mask matrix, the in-band smooth covariance is obtained, providing more robust input data for subsequent sound source localization.
[0032] Specifically, such as Figure 3As shown, the preset network model has a built-in embedded beamforming network, and S200 includes the following steps: S201: Obtain the multi-channel observation vector in the spherical harmonic domain, and perform a concatenation operation on the real and imaginary parts of the multi-channel observation vector in the spherical harmonic domain according to their channel dimensions to obtain the concatenated feature after concatenation of the real and imaginary parts; The spherical harmonic domain multi-channel observation vector is essentially a complex-valued signal, mathematically expressed as: real part + imaginary unit × imaginary part. However, the underlying computational logic of current preset network models (including the embedded-beamforming network in this scheme) only supports real-valued data input and cannot directly analyze the imaginary part information in complex-valued signals. Discarding the imaginary part directly results in the loss of the signal's phase characteristics; using only the real part as input, the preset network model cannot fully capture the phase correlation between multiple channels (and phase correlation is crucial for distinguishing between direct sound (phase-consistent) and reflected sound (phase-chaotic) in substation scenarios).
[0033] Therefore, by concatenating the real and imaginary parts along the channel dimension, the complex-valued vector can be transformed into a concatenated feature of real values (i.e., the real and imaginary elements of each channel are arranged sequentially). This satisfies the requirements of the preset network model for real-valued inputs and avoids the loss of complex-valued information. The concatenated feature is calculated using the following formula (2): Formula (2); Where Z(t,f) is the splicing feature; To take the real part; To extract the imaginary part; This is a channel-based splicing operation. S202: Input the splicing features into the preset embedding-beamforming network to obtain the spherical harmonic domain mask matrix.
[0034] The spliced features are input into a preset embedding-beamforming network to obtain a spherical harmonic domain mask matrix. The dimension of the spherical harmonic domain mask matrix is the same as that of X(t,f) (each channel and each time-frequency point corresponds one-to-one). It is used to filter out the part of the target sound source in the spherical harmonic domain multi-channel observation vector and suppress reverberation and noise.
[0035] Specifically, the spherical harmonic domain mask matrix is calculated using the following formula (3): Formula (3); in, The mask matrix is a spherical harmonic domain. Let Θ represent the embedded beamforming network; Θ is the set of network parameters, which is fixed after training.
[0036] S203: Based on element-wise multiplication, the spherical harmonic domain mask matrix and the spherical harmonic domain multi-channel observation vector are operated to obtain the spherical harmonic domain observation matrix; Based on element-wise multiplication, the spherical harmonic domain mask matrix and the spherical harmonic domain multi-channel observation vector are calculated according to the following formula (4) to obtain the spherical harmonic domain observation matrix. This process is the actual denoising step of the spherical harmonic domain multi-channel observation vector by using the spherical harmonic domain mask matrix learned by the embedded beamforming network. The calculation process of the spherical harmonic domain observation matrix is shown in the following formula (4): Formula (4); in, is the observation matrix of the spherical harmonic domain; X(t,f) is the multi-channel observation vector of the spherical harmonic domain; ⊙ represents element-wise multiplication; Let be the spherical harmonic mask matrix.
[0037] S204: Perform second-order time-frequency statistics on the spherical harmonic domain observation matrix to obtain the spatial covariance matrix; Second-order statistics is a general term for second-order moments in signal processing, used to describe second-order characteristics of signals such as energy distribution and correlation. By performing time-frequency second-order statistics on the spherical harmonic domain observation matrix, a spatial covariance matrix is obtained. This allows for the integration and refinement of the spatial correlation characteristics of target sound sources scattered across various channels in the spherical harmonic domain observation matrix. The dispersed spatial correlation characteristics are quantified into unified statistical characteristics, resulting in a spatial covariance matrix where the effective correlation of the target sound source is more prominent and interference pseudo-correlation is weakened.
[0038] Specifically, the spatial covariance matrix is obtained by performing second-order time-frequency statistics using the following formula (5): Formula (5); in, It is the spatial covariance matrix; For the observation matrix of the masked sphere harmonic domain; for The conjugate transpose of .
[0039] It should be added that, as can be seen from formulas (3), (4), and (5), the spatial covariance matrix... The accuracy is related to the network parameter set Θ. The optimization of the network parameter set Θ is achieved through offline training. Specifically, a spherical harmonic domain signal containing only the direct path of the target sound source is generated using early reflection simulation techniques. And construct the target covariance matrix reflecting the ideal state according to formula (6): Formula (6); in, Let the target covariance matrix be denoted as 'covariance'. This is a spherical harmonic domain signal containing only the direct path to the target sound source; for The conjugate transpose of .
[0040] Then, the loss function is minimized during the training phase. The loss function is calculated according to the following formula (7): Formula (7) in, The loss function; It is the spatial covariance matrix; The target covariance;
[0041] It is the Frobenius norm.
[0042] When the loss function When convergence to the minimum, the spatial covariance matrix corresponding to the network parameter set Θ The target covariance matrix has been approximated, and the set of network parameters Θ is now fixed.
[0043] S205: Based on the spatial covariance matrix, the in-band smooth covariance is obtained.
[0044] The purpose of obtaining the in-band smooth covariance based on the spatial covariance matrix is to reduce the random fluctuations of the spatial covariance matrix, thereby enhancing the dynamic resolution capability of the target sound source. The in-band smooth covariance can more accurately characterize the spatial features of the target sound source than the spatial covariance matrix, providing reliable support for subsequent localization.
[0045] Specifically, such as Figure 4 As shown, S205 includes the following steps: S2051: Obtain the frequency acoustic emission band of equipment in a substation during abnormal conditions; S2052: Based on the frequency acoustic emission band, the spatial covariance matrix is processed to obtain the in-band smooth covariance.
[0046] Obtain the frequency acoustic emission band of the equipment in the substation when it is abnormal. For example, based on experience, it is known that the equipment will generate a high-frequency ultrasonic signal of 20–40 kHz when it is partially discharged. This frequency band is the frequency acoustic emission band of the equipment when it is abnormal. However, it should be selected according to the actual situation of the equipment. No specific limitation is made here.
[0047] Processing the spatial covariance matrix based on the frequency acoustic emission band includes: selecting the corresponding spatial covariance matrix within the frequency acoustic emission band, and smoothing and integrating these selected matrices in the frequency dimension to obtain the intra-band smooth covariance. The intra-band smooth covariance is calculated using the following formula (8): Formula (8) in, For in-band smoothing covariance; f low , f high ΔF represents the minimum and maximum values of the frequency band emission band; ΔF is the number of frequency points to be summed. Let be the spatial covariance matrix.
[0048] S300: Based on the in-band smooth covariance, the spatial spectrum matrix is obtained. The spatial spectrum matrix includes spatial spectrum values in different candidate sound source directions. Based on the in-band smooth covariance, the spatial spectrum matrix is obtained. The abstract in-band smooth covariance can be transformed into an intuitive spatial energy distribution (i.e., the spatial spectrum matrix). By quantifying the energy intensity in different directions, the direction of the target sound source can be initially located. The candidate sound source directions are uniformly selected on the acquisition device, and the spatial spectrum value is the spatial energy corresponding to each candidate direction.
[0049] Specifically, such as Figure 5 As shown, S300 includes the following steps: S301: Acquire the directions of multiple candidate sound sources on the spherical distribution array of the acquisition device; When acquiring multiple candidate sound source directions on the spherical distribution array of the acquisition device, T-design can be used for uniform sampling on the spherical surface, or other uniform spherical sampling strategies can be used; no specific limitations are made here.
[0050] S302: Based on the candidate sound source directions, calculate the column vector of spherical harmonic functions corresponding to each candidate sound source direction, and stack them into a scan matrix; The column vector of spherical harmonic functions corresponding to the direction of each candidate sound source is calculated according to the following formula (9): Formula (9); in, In direction The column vector of spherical harmonic functions at the location; Let n be a spherical harmonic function of degree n and order m. It is a spherical harmonic truncated order.
[0051] The scan matrix is calculated according to the following formula (10): Formula (10); in, This is the scan matrix; This represents the number of candidate sound source directions.
[0052] S303: The spatial spectrum matrix is calculated based on the scan matrix and the in-band smoothing covariance.
[0053] After obtaining the scan matrix and the in-band smoothing covariance, the spatial spectral matrix is calculated according to the following formula (11): Formula (11); in, It is the spatial spectral matrix; This is the scan matrix; For in-band smoothing covariance; Find the inverse of a matrix; This is the conjugate transpose.
[0054] S400: The spatial spectrum matrix is divided into multiple sound sources to obtain a set of multiple source directions, which enables the far-field sound source localization of the substation. The set of multiple source directions includes the polar angle and azimuth coordinates corresponding to the direction of each target sound source. The set of multiple source directions is used to reflect the direction of sound sources within the substation.
[0055] There are various methods for distinguishing multiple sound sources from the spatial spectral matrix to obtain the set of multiple source directions. For example, in calculating the spatial covariance matrix... Previously, blind source separation technology was introduced. First, the mixed multi-source signals were separated into single-source signals. Then, the separated single-source signals were localized to solve the multi-source interference problem. Other multi-source differentiation methods can also be used, but no specific limitation is made here.
[0056] In this embodiment, an internal iterative peak selection method with angle threshold constraints is used to distinguish multiple sound sources. Specifically, as shown... Figure 6 As shown, S400 includes the following steps: S401: Using the spatial spectrum matrix as the sound source search range, the candidate sound source direction corresponding to the maximum value of the spatial spectrum value in the spatial spectrum matrix is taken as the first target sound source direction; In the spatial spectrum matrix, the magnitude of the spatial spectrum value reflects the energy intensity of the sound source in the corresponding candidate direction; that is, the higher the spectrum value, the greater the probability that a sound source exists in that direction. Therefore, taking the spatial spectrum matrix as the sound source search range, the candidate sound source direction corresponding to the maximum value of the spatial spectrum value in the spatial spectrum matrix is taken as the first target sound source direction. The first target sound source direction is calculated according to the following formula (12): Formula (12); in, The direction of the first target sound source; It is the spatial spectral matrix; To obtain the maximum value.
[0057] S402: Based on the preset angle threshold and the direction of the first target sound source, determine the search range of the next sound source, and search for the directions of the remaining target sound sources in the search range of the next sound source until all target sound source directions are determined to form a multi-source direction set.
[0058] Since sound sources or strongly correlated interference sources from the same device may be distributed in similar directions, in order to avoid repeated identification or misjudgment, it is necessary to define the interference range of the first target sound source direction by using a preset angle threshold. That is, the area of the first direction ± the preset angle threshold is excluded from the next round of search (for example, if the first sound source direction is 120° and the angle threshold is 5°, then the range of 115°~125° is excluded). The remaining area that is not excluded is the next sound source search range. In this way, it can be ensured that the subsequent search is for independent sound sources that are spatially separated from the first sound source.
[0059] Furthermore, within the next sound source search range, the candidate sound source direction with the largest spectral value within that range is identified as the second target sound source direction. Then, using this second direction as the center, its interference range is excluded according to a preset angle threshold, thus obtaining the next sound source search range for which sound sources need to be searched. This process continues, using the preset angle threshold as a constraint, to search for the next sound source direction until no more valid sound sources are found within the search range. Finally, all found target sound source directions are summarized to form a multi-source direction set.
[0060] Specifically, such as Figure 7 As shown, S402 includes the following steps: S4021: Calculate the angle between the candidate sound source direction corresponding to other spatial spectrum values in the spatial spectrum matrix and the first target sound source direction; The angle between the candidate sound source direction corresponding to other spatial spectrum values in the spatial spectrum matrix and the first target sound source direction is calculated according to the following formula (13): Formula (13); Where δ is the angle between the candidate sound source direction corresponding to other spatial spectrum values in the spatial spectrum matrix and the first target sound source direction; The polar angle of the first sound source direction; The azimuth angle of the first sound source direction; The polar angle of the candidate sound source direction corresponding to other spatial spectrum values in the spatial spectrum matrix; This represents the azimuth angle of the candidate sound source direction corresponding to other spatial spectrum values in the spatial spectrum matrix.
[0061] S4022: Obtain the directions of all candidate sound sources corresponding to the included angles that satisfy the preset angle threshold in the spatial spectrum matrix as the next sound source search range; The relationship between the preset angle threshold and the included angle is shown in the following formula (14): Formula (14); Where δ is the angle between the candidate sound source direction and the target sound source direction corresponding to other spatial spectrum values in the spatial spectrum matrix; The preset angle threshold is used; the target sound source direction here includes the first target sound source direction and the remaining target sound source directions that are subsequently confirmed.
[0062] S4023: Within the next sound source search range, the candidate sound source direction corresponding to the maximum value of the spatial spectrum value in the spatial spectrum matrix is taken as the next target sound source direction, and the remaining target sound source directions are searched with a preset angle threshold as a constraint.
[0063] After determining the direction of the first target sound source and defining its interference range (the area to be excluded) using a preset angle threshold, the remaining unexcluded area becomes the search range for the next sound source. Within this range, the candidate direction with the largest spatial spectral value is selected as the next target sound source direction (i.e., the sound source with the largest spatial spectral value (strongest energy) within this range) using the same judgment criteria as the first sound source. Subsequently, with the direction of the second target sound source as the center, its interference range is again excluded according to the preset angle threshold, resulting in a new search range. The search continues within this range, looking for the candidate direction with the largest spectral value as the third target sound source direction, and so on, iterating until all target sound source directions are confirmed, at which point the search stops. Through this progressive method, all independent target sound source directions can be filtered from the spatial spectral matrix, ultimately forming a complete set of multi-source directions.
[0064] It should be noted that the above steps only yield a set of multi-source directions within a single frame. However, the actual activity of the target sound source within the substation (such as abnormal equipment discharge) often exhibits continuity (e.g., pulse signals lasting from several seconds to several minutes). The results of a single frame may be affected by transient noise interference, leading to jumps or misjudgments. Therefore, speech or pulse activity detection can be used to achieve continuous tracking across frames. Specifically, this involves: performing correlation analysis on the multi-source direction sets of adjacent time frames (e.g., calculating the angular deviation of directions between frames; if it is less than the tracking threshold, it is determined to be continuous movement or stable existence of the same sound source); combining the temporal activity of the sound source signal (e.g., the periodicity of partial discharge pulses, the persistence of mechanical noise); filtering out isolated false directions in a single frame; and retaining the stable sound source trajectory across frames, ultimately obtaining the multi-source dynamic localization result in the continuous time dimension. Alternatively, a time-series tracking method based on sound source activity detection and Kalman filtering can be used to achieve continuous tracking and accurate localization of multiple far-field sound sources by judging the sound source activity state and predicting its location.
[0065] Furthermore, this embodiment also demonstrates the advantages of the present invention through a series of experiments, specifically including: Experiment 1: such as Figure 8As shown, this experiment aims to verify the performance differences between the method of this invention and existing mainstream sound source localization methods under different signal-to-noise ratio (SNR) environments. Experiment 1 uses the SNR level as a variable (SNR is an indicator that measures the relative strength of the effective signal and the interference noise in a signal; a higher SNR indicates that the energy of the effective signal is stronger relative to the noise energy, and the signal is less affected by noise interference). Traditional signal processing methods (SHD-MVDR), deep learning-based methods (1D-CNN (GCC), Cross3D, icoCNN), and different spherical harmonic order configurations of this invention (orders 1 and 4) are selected as comparison objects, with localization accuracy as the evaluation index. The results show that as the signal-to-noise ratio (SNR) increases, the localization accuracy of all methods increases, but the method of this invention maintains the best performance, especially under low SNR conditions, where its advantages over traditional methods and other deep learning methods are more significant. Furthermore, the higher spherical harmonic order configuration in this invention performs better, indicating that this invention can effectively suppress noise interference through dynamic masking and optimization processing in the spherical harmonic domain. Even in scenarios where the signal is severely affected by noise, it can still stably extract the target sound source information, verifying its robustness in the low SNR environment of substations.
[0066] Experiment 2: such as Figure 9 As shown, this experiment focuses on the impact of different load levels in substations on the performance of the model of this invention. Different load levels reflect the operating load of power equipment in the substation (such as high load conditions with multiple devices operating versus low load conditions with one device operating; or high load conditions with multiple lines of a single device operating versus low load conditions with one line operating, etc.). By synchronously monitoring two core indicators, positioning accuracy and model loss, the adaptability of the model to dynamic changes in substation operating conditions is verified. In the experiment, the load level gradually changes from extremely low to extremely high. The results show that in the extremely low to medium load stage, the model positioning accuracy remains at a high and stable level, and the model loss continues to decrease and remains at a low value. This indicates that the equipment operating status in the substation is stable and the environmental interference is relatively fixed in this stage, and the dynamic analysis and positioning mechanism of this invention can work stably. When the load increases to a high level or above, the positioning accuracy begins to decrease slowly, while the model loss gradually increases. This is due to the increase in noise sources and signal interference complexity in the substation equipment under high load. However, even under extremely high load, the model still maintains high accuracy and low loss. Overall, the model of this invention has good stability under different load scenarios in substations, can adapt to the dynamic changes in substation operating conditions, and meets the needs of long-term monitoring.
[0067] Experiment 3: such as Figure 10As shown, this experiment sets up four models: a basic model with only basic signal processing capabilities, a basic model superimposed with a spherical harmonic domain mask matrix, a basic model superimposed with an angle threshold iterative peak selection module, and the complete model of this embodiment. Fault identification accuracy is used as the evaluation index. The results show that the basic model has the lowest accuracy, indicating that traditional signal processing alone is insufficient for locating fault sources in the complex environment of substations. Superimposing a single core module significantly improves the accuracy of all models. The model with the angle threshold iterative peak selection module has a slightly higher accuracy than the model with the spherical harmonic domain mask, indicating that the two modules play roles in multi-source differentiation and noise suppression, respectively. The complete model of this embodiment has the highest accuracy, significantly better than the other three models. This demonstrates a synergistic effect between the spherical harmonic domain mask matrix and the angle threshold iterative peak selection module—the mask module filters the target signal cleanly, providing high-quality input for iterative peak selection, while the iterative peak selection module accurately distinguishes between multi-source and pseudo-peaks. The combination of the two greatly improves fault identification capabilities, verifying the rationality and innovation of the technical solution of this invention.
[0068] To facilitate understanding by those skilled in the art, the workflow of the far-field sound source localization method for substations provided by this invention is as follows: Target frequency sound within the substation is acquired using acquisition devices arranged in a spherical array. The spherical harmonic domain (SHV) multi-channel observation vector is obtained through short-time Fourier transform and spherical harmonic transform. The real and imaginary parts of the SHV multi-channel observation vector are concatenated along the channel dimension and input into an embedding-beamforming network to obtain the SHV mask matrix. The SHV observation matrix is obtained through element-wise multiplication with the SHV multi-channel observation vector. A spatial covariance matrix is generated by performing second-order time-frequency statistics on the SHV observation matrix. The spatial covariance matrix is then processed to obtain an in-band smooth covariance by considering the frequency sound emission bands during substation equipment anomalies. Multiple candidate sound source directions are preset on the spherical distribution array. The column vector of spherical harmonic functions corresponding to each direction is calculated and stacked into a scanning matrix. The scanning matrix and the in-band smooth covariance are combined to calculate a spatial spectrum matrix containing spatial spectrum values of different candidate directions. The spatial spectrum matrix is used as the initial search range. The direction with the largest spatial spectrum value is selected as the first target sound source direction. The next sound source search range is determined according to the preset angle threshold. The search is repeated in the new range until all target sound source directions are confirmed. Finally, a multi-source direction set containing the polar angle and azimuth coordinates of each target sound source is formed, so as to realize the accurate positioning of far-field sound sources in substations.
[0069] As can be seen, this invention converts target frequency sound within a substation into a spherical harmonic domain (SHDN) multi-channel observation vector. Then, a pre-defined network model (used for filtering the SHDN multi-channel observation vector) generates a SHDN mask matrix. This mask matrix is used to filter out effective target signals, reducing the impact of noise and interference in the complex substation environment and minimizing the generation of spurious peaks during localization. This results in an in-band smooth covariance that accurately characterizes the spatial features of the sound source. Based on the in-band smooth covariance, a spatial spectrum matrix containing spatial spectrum values of different candidate sound source directions is constructed. This spatial spectrum matrix is then used to distinguish multiple sound sources, extracting a set of multi-source directions containing the polar and azimuth coordinates of each target sound source. This effectively solves the problems of noise interference and difficulty in distinguishing multiple sound sources within a substation, achieving accurate localization of sound sources within the substation. The model parameters are trained offline, and online signal processing and localization are performed quickly, requiring no large computational resources or retraining to adapt to different substation operating conditions, thus balancing real-time performance and environmental adaptability. By employing speech or impulse activity detection to achieve continuous tracking across frames, filtering out spurious directions that appear isolated in a single frame, and preserving the stable sound source trajectories that exist across frames, dynamic localization of multiple sound sources in the continuous time dimension can be achieved.
[0070] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention is not limited to the specific combination of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in this invention.
Claims
1. A far-field sound source positioning method applied to a substation, characterized in that, The method comprises the following steps: obtaining target frequency sound in a substation and transforming the target frequency sound into a spherical harmonic domain multi-channel observation vector; inputting the spherical harmonic domain multi-channel observation vector into a preset network model to obtain a spherical harmonic domain mask matrix, and obtaining in-band smooth covariance based on the spherical harmonic domain mask matrix; the preset network model is used for filtering processing of the spherical harmonic domain multi-channel observation vector; obtaining a spatial spectrum matrix based on the in-band smooth covariance, the spatial spectrum matrix comprising spatial spectrum values in different candidate sound source directions; carrying out multi-source differentiation on the spatial spectrum matrix to obtain a multi-source direction set, and realizing far-field sound source positioning of the substation; the multi-source direction set comprises polar angle and azimuth angle coordinates corresponding to each target sound source direction, and the multi-source direction set is used for reflecting sound source directions in the substation.
2. The method for locating a far-field sound source applied to a transformer substation according to claim 1, characterized in that, Obtaining target frequency sound in a substation and transforming the target frequency sound into a spherical harmonic domain multi-channel observation vector comprises: arranging a collecting device according to a spherical distribution array in the substation to obtain the target frequency sound; carrying out short-time Fourier transform and spherical harmonic transform on the target frequency sound to obtain the spherical harmonic domain multi-channel observation vector.
3. The method for locating a far-field acoustic source applied to a substation according to claim 1, characterized in that, The preset network model is embedded with an embedded-beamforming network; inputting the spherical harmonic domain multi-channel observation vector into the preset network model to obtain the spherical harmonic domain mask matrix comprises: obtaining the spherical harmonic domain multi-channel observation vector, and carrying out splicing operation on real parts and imaginary parts of the spherical harmonic domain multi-channel observation vector according to channel dimensions to obtain spliced features after splicing of the real parts and the imaginary parts; inputting the spliced features into a preset embedded-beamforming network to obtain the spherical harmonic domain mask matrix.
4. The method for locating a far-field acoustic source applied to a substation according to claim 1, characterized in that, Obtaining in-band smooth covariance based on the spherical harmonic domain mask matrix comprises: carrying out operation on the spherical harmonic domain mask matrix and the spherical harmonic domain multi-channel observation vector based on element-by-element multiplication to obtain a spherical harmonic domain observation matrix; carrying out time-frequency second-order statistics on the spherical harmonic domain observation matrix to obtain a spatial covariance matrix; obtaining in-band smooth covariance based on the spatial covariance matrix.
5. The method for locating a far-field acoustic source applied to a substation according to claim 4, characterized in that, Obtaining in-band smooth covariance based on the spatial covariance matrix comprises: obtaining a frequency sound emission band of a device in the substation under abnormal condition; processing the spatial covariance matrix based on the frequency sound emission band to obtain the in-band smooth covariance.
6. The far-field sound source localization method applied to a substation of claim 2, characterized in that, Obtaining a spatial spectrum matrix based on the in-band smooth covariance comprises: obtaining a plurality of candidate sound source directions on the spherical distribution array of the collecting device; calculating a spherical harmonic function column vector corresponding to each candidate sound source direction based on the candidate sound source directions, and stacking the spherical harmonic function column vector into a scanning matrix; calculating the spatial spectrum matrix based on the scanning matrix and the in-band smooth covariance.
7. The far-field acoustic source localization method applied to a substation according to claim 1, characterized in that, Carrying out multi-source differentiation on the spatial spectrum matrix to obtain a multi-source direction set comprises: taking the spatial spectrum matrix as a sound source search range, and taking a candidate sound source direction corresponding to a maximum value of a spatial spectrum value in the spatial spectrum matrix as a first target sound source direction; confirming a next sound source search range according to a preset angle threshold and the first target sound source direction, and searching for remaining target sound source directions in the next sound source search range until all target sound source directions are confirmed to form the multi-source direction set.
8. The method for locating a far-field acoustic source applied to a substation according to claim 7, characterized in that, According to the preset angle threshold and the first target sound source direction, a next sound source search range is determined, and the remaining target sound source directions are searched in the next sound source search range, including: An angle between a candidate sound source direction corresponding to other spatial spectrum values in the spatial spectrum matrix and the first target sound source direction is calculated; Candidate sound source directions corresponding to angles satisfying the preset angle threshold in the spatial spectrum matrix are obtained as the next sound source search range; In the next sound source search range, a candidate sound source direction corresponding to a maximum value of the spatial spectrum values in the spatial spectrum matrix is taken as the next target sound source direction, and the remaining target sound source directions are continuously searched with the preset angle threshold as a constraint.
Citation Information
Cited By
Cable internal fault positioning system using sound wave signals
CN121955617A