Fast iteration deconvolution sound source imaging method and system
By optimizing the point spread function through a fast iterative shrinkage threshold algorithm and combining it with the FISTA iteration technology, the problems of insufficient computational efficiency and imaging resolution of the deconvolution algorithm are solved, and efficient and accurate sound source imaging and multi-sound source localization are achieved.
Patent Information
- Application Number
- CN202510702055.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-09
AI Technical Summary
Existing deconvolution algorithms have deficiencies in computational efficiency and imaging resolution, making it difficult to meet the needs of precise positioning of multiple sound sources in complex sound fields.
The fast iterative shrinkage threshold algorithm is used to optimize the point spread function. Combined with the FISTA iterative technology, the audio and video signals are acquired to construct a cross-spectral matrix, simulate the sound wave propagation path, generate the steering vector, and perform frequency domain calculation to optimize the sound source distribution map.
It significantly improves the speed and accuracy of sound source imaging, enables accurate multi-sound source localization and complex sound field imaging, reduces calculation time and suppresses background noise interference.
Smart Images

Figure CN120610119A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of acoustic imaging, and in particular to a fast iterative deconvolution sound source imaging method and system. Background Art
[0002] Acoustic imaging, a means of visualizing spatial sound fields, is widely used in aviation, rail transit, automotive noise control, and other fields, and is becoming a key tool for noise control and acoustic analysis. The core steps of current acoustic imaging technology are to collect sound field signals using acoustic arrays and then perform signal processing and sound source localization using beamforming algorithms.
[0003] Traditional beamforming methods are primarily based on the Delay and Sum (DAS) algorithm, which offers low computational complexity and strong robustness. However, this algorithm suffers from poor spatial resolution at low frequencies and insufficient sidelobe suppression, making it difficult to accurately localize multiple sound sources in complex sound fields. To improve the resolution and localization accuracy of sound source imaging, researchers at home and abroad have proposed an improved method based on deconvolution, the Deconvolution Approach for the Mapping of Acoustic Sources (DAMAS) algorithm. The DAMAS algorithm deconvolves the traditional beamforming results by constructing a point spread function (PSF), effectively suppressing sidelobe interference and improving spatial resolution. The significant performance of deconvolution beamforming has made it a mainstream acoustic imaging algorithm. However, due to the large number of variables involved in the PSF, the DAMAS algorithm is computationally complex and time-consuming, limiting its real-time applicability and scope.
[0004] To address these issues, researchers have proposed a series of improved algorithms based on the DAMAS algorithm. For example, the DAMAS2 algorithm uses FFT-NNLS (non-negative least squares combined with fast Fourier transform) to optimize the solution process, reducing computational time compared to DAMAS. However, its least squares estimation has a relatively slow update speed, requires many iterations, and requires long computational time, and there is still room for further resolution improvement.
[0005] Therefore, traditional beamforming methods cannot balance computational efficiency and imaging resolution. Based on this, the paper "Research on Sound Source Identification Methods Using Rapid Iterative Shrinkage Threshold Beamforming, a PhD dissertation from Chongqing University" proposes a periodic boundary-based FFT-FISTA deconvolution beamforming algorithm. However, this algorithm uses FISTA to optimize the Fourier transform method rather than the DAMAS2 algorithm. The paper "Research on Sparse Equivalent Source Sound Field Reconstruction Algorithms for Planar Arrays, a PhD dissertation from Chongqing University" proposes a rapid iterative shrinkage threshold sound field reconstruction algorithm. However, its rapid iterations use the norm for iterative optimization rather than using rapid iterations to optimize the deconvolution beamforming algorithm. Summary of the Invention
[0006] The present invention aims to solve the technical problems of low computational efficiency and low imaging resolution in existing deconvolution algorithms.
[0007] The present invention solves the above technical problems through the following technical means:
[0008] A fast iterative deconvolution sound source imaging method is proposed, including:
[0009] Acquiring audio signals and video signals generated during a partial discharge process, extracting frequency domain information of the audio signals and constructing a cross-spectrum matrix;
[0010] Simulate the propagation path of the sound wave within the scanning plane area, which is a key parameter initialized according to the sound source localization requirements, and generate a steering vector;
[0011] Modeling the acoustic field response characteristics of each point in the scanning plane area to generate a point spread function;
[0012] A fast iterative shrinkage threshold algorithm is used to optimize and solve the point spread function to obtain an optimal sound source distribution map, wherein the step size factor and the intermediate variable are updated during the iteration process based on the residual vector constructed based on the steering vector and the cross-spectral matrix.
[0013] Furthermore, the step of obtaining the audio signal and the video signal generated during the partial discharge process includes:
[0014] Use microphone array to collect multi-directional sound field signals of partial discharge sound sources;
[0015] Synchronously collecting and digitally processing the output signals of all microphones in the microphone array to obtain the audio signal, wherein the audio signal is a sound pressure signal;
[0016] A camera is used to detect video signals at the scene of partial discharge, and the video frames and audio signals in the video signals are aligned using time stamps.
[0017] Furthermore, the key parameters initialized according to the sound source localization requirements also include sound speed, sound source frequency, sound imaging frequency constraint and camera resolution;
[0018] The scanning plane area is calculated based on the measured distance between the microphone array and the sound source and the pixels of the camera, and the scanning plane area is divided into equally spaced grid points to determine the grid step size.
[0019] Furthermore, extracting the frequency domain information of the audio signal to construct a cross-spectral matrix includes:
[0020] Transferring the audio signal from the time domain to the frequency domain to obtain an audio frequency domain signal;
[0021] The cross spectrum matrix is obtained by multiplying the audio frequency domain signal and its conjugate transpose and dividing the result by the number of frequency frames.
[0022] Furthermore, simulating the propagation path of the acoustic wave within the scanning plane area to generate the steering vector includes:
[0023] Calculate the propagation path length from the grid points of the scan plane area division to the microphone array;
[0024] According to the propagation path length and the speed of sound, the phase delay of the sound wave is calculated to obtain the steering vector.
[0025] Furthermore, the step of modeling the acoustic field response characteristics of each point in the scanning plane region to generate a point spread function includes:
[0026] Calculate the point spread function in the time domain based on the propagation path length and sound speed from the grid points divided into the scanning plane area to the microphone array;
[0027] A point spread function in the time domain is converted into a point spread function in the frequency domain based on the steering vector, and a discrete integral constant is calculated based on the point spread function in the frequency domain.
[0028] Furthermore, the calculation formula of the discrete integral constant is:
[0029] a=∑|PSF(x,y)|
[0030] PSF(x,y)=|w(x,y,k) H g(x,y,k)| 2
[0031] Where a represents the discrete integral constant, PSF(x,y) represents the point spread function in the frequency domain, w(x,y,k) represents the weighted steering vector, g(x,y,k) represents the steering vector, where k represents the frequency in the frequency domain transform, and H represents the conjugate operator.
[0032] Furthermore, the method of optimizing the point spread function by using a fast iterative shrinkage threshold algorithm to obtain a sound source distribution map includes:
[0033] Initializing iteration parameters, including the number of iterations, iteration variables, step factors, and intermediate variables;
[0034] Based on the residual vector and the discrete integral constant and in combination with the non-negative constraint, the current solution of the nth iteration is updated, where the current solution represents the sound source distribution estimate, wherein the update formula of the current solution is:
[0035]
[0036] Where q(n+1) is the updated result of the current solution q(n) of the nth iteration, y(n) represents the intermediate variable, and DAS represents the beamforming result. w(x,y,k) represents the weighted steering vector, CSM represents the cross spectrum matrix, H represents the conjugate operator, N mic represents the number of microphones; PFS represents the point spread function in the time domain, DAS-PFS*q(n) represents the residual vector, a represents the discrete integration constant, and b(n) represents the result calculated using PFS;
[0037] Update the step size factor and the intermediate variables based on the current solution before proceeding to the next iteration;
[0038] When the number of iterations is reached, a global optimal current solution is obtained as the optimal sound source distribution map.
[0039] Furthermore, after optimizing and solving the point spread function by using a fast iterative shrinkage threshold algorithm to obtain an optimal sound source distribution map, the method further includes:
[0040] The optimal sound source distribution map is fused with the video signal to obtain visualization information of the sound source location and the surrounding environment.
[0041] In addition, the present invention also proposes a fast iterative deconvolution sound source imaging system, comprising:
[0042] A signal acquisition module, used to acquire audio signals and video signals generated during the partial discharge process;
[0043] A cross-spectrum matrix construction module, used to extract frequency domain information of the audio signal to construct a cross-spectrum matrix;
[0044] A steering vector calculation module is used to simulate the propagation path of sound waves in a scanning plane area and generate a steering vector. The scanning plane area is a key parameter obtained based on the sound source localization requirement.
[0045] A point spread function calculation module is used to model the acoustic field response characteristics of each point in the scanning plane area and generate a point spread function;
[0046] An iterative module is used to optimize and solve the point spread function using a fast iterative shrinkage threshold algorithm to obtain an optimal sound source distribution map, wherein during the iterative process, a step factor and an intermediate variable are updated based on a residual vector constructed based on the steering vector and the cross-spectral matrix.
[0047] The advantages of the present invention are:
[0048] (1) The present invention combines the accelerated iteration of FISTA in the calculation process of the point spread function, which not only reduces the calculation time but also effectively suppresses the influence of background noise and multi-source interference, so that the generated sound source distribution map has higher resolution, can accurately restore the spatial position and intensity distribution of the local discharge sound source, significantly improves the speed and accuracy of sound source imaging, and can accurately perform multi-sound source positioning and imaging analysis of complex sound fields, thus bringing new directions for localization of local discharge sound source imaging.
[0049] (2) The FISTA algorithm adopts a dynamic step-size adjustment mechanism to predict the update of the solution through the accelerated iterative formula, so that each update step is closer to the optimal solution, while avoiding the oscillation phenomenon in the traditional fixed step-size method.
[0050] (3) By calculating the residual and iterative update through frequency domain convolution, the high complexity of a large number of matrix operations in spatial domain operations is effectively avoided, greatly improving the computational efficiency.
[0051] Additional aspects and advantages of the present invention will be set forth in part in the description which follows and, in part, will be obvious from the description which follows, or may be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 1 is a flow chart of a fast iterative deconvolution sound source imaging method proposed in one embodiment of the present invention;
[0053] Figure 2 is a schematic diagram of the microphone array layout in one embodiment of the present invention;
[0054] Figure 3 1 is a schematic diagram of a complete process of rapid iterative deconvolution sound source imaging in one embodiment of the present invention;
[0055] Figure 4 2 is a schematic diagram of a single sound source simulation result according to an embodiment of the present invention;
[0056] Figure 5 1 is a schematic diagram of simulation results of dual sound sources in one embodiment of the present invention;
[0057] Figure 6 The figure is a schematic structural diagram of a fast iterative deconvolution sound source imaging system proposed in one embodiment of the present invention. DETAILED DESCRIPTION
[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0059] like Figure 1 As shown, the first embodiment of the present invention proposes a fast iterative deconvolution sound source imaging method, which includes the following steps:
[0060] S10, acquiring audio signals and video signals generated during a partial discharge process, extracting frequency domain information of the audio signals and constructing a cross-spectrum matrix;
[0061] S20, simulating the propagation path of the sound wave in the scanning plane area to generate a steering vector, wherein the scanning plane area is a key parameter initialized according to the sound source localization requirement;
[0062] S30, modeling the acoustic field response characteristics of each point in the scanning plane area to generate a point spread function;
[0063] S40. Using a fast iterative shrinkage threshold algorithm to optimize and solve the point spread function to obtain an optimal sound source distribution map, wherein during the iteration process, the step size factor and the intermediate variable are updated based on the residual vector constructed based on the steering vector and the cross-spectral matrix.
[0064] It should be noted that this embodiment digitizes the audio signal generated by the collected partial discharge process, extracts frequency domain information and constructs a cross-spectral matrix (CSM), and at the same time initializes key parameters such as the scanning plane area according to the sound source localization requirements, accurately simulates the propagation path of the sound wave in the scanning area, and generates a steering vector for subsequent sound field calculation, thereby improving the resolution capability under complex sound fields. By modeling the sound field response characteristics of each grid point in the scanning plane area, a point spread function of the system is generated. This function serves as an important basis for describing the sound field distribution and provides a reference for subsequent iterative optimization. Finally, the point spread function is solved using the FISTA iterative algorithm, and the sound source distribution is accelerated using fast iteration technology. By dynamically adjusting the intermediate variables and step size parameters, the solution is gradually approximated to the actual sound source distribution, while maintaining the efficiency and stability of the algorithm. Because the partial discharge time is relatively fast, fast imaging is required. This embodiment not only reduces the calculation time but also effectively suppresses the influence of background noise and multi-source interference, so that the generated sound source distribution map has a higher resolution, can accurately restore the spatial position and intensity distribution of the partial discharge sound source, and significantly improves the speed and accuracy of sound source imaging.
[0065] As a further preferred technical solution, in step S10: obtaining the audio signal and video signal generated during the partial discharge process specifically includes:
[0066] Use microphone array to collect multi-directional sound field signals of partial discharge sound sources;
[0067] Synchronously collecting and digitally processing the output signals of all microphones in the microphone array to obtain the audio signal, wherein the audio signal is a sound pressure signal;
[0068] A camera is used to detect video signals at the scene of partial discharge, and the video frames and audio signals in the video signals are aligned using time stamps.
[0069] It should be noted that if Figure 2 As shown, the microphone array used in this embodiment is a high-density microphone array, consisting of 64 high-sensitivity microphones, and adopts a spiral array layout design, with an array size of 0.15*0.15m 2 This ensures uniform coverage of the target sound source's sound field and accurate acquisition of multi-directional signals. High-quality sound pressure signals are then acquired from the microphone array and digitized, providing basic data support for subsequent sound field reconstruction and sound source localization. Preferably, the output signals of all microphones in the array are synchronously acquired and digitized through multi-channel data acquisition.
[0070] It should be noted that this embodiment is equipped with a high-definition camera (resolution of 1920×1080 or higher, frame rate of 30fps to 60fps), which can record dynamic images of the partial discharge experiment site in real time. By collecting real-time visual information of the partial discharge detection site, it provides auxiliary reference for sound field positioning and environmental monitoring.
[0071] Furthermore, this embodiment shares the same time reference when collecting audio and video signals, and uses timestamps to accurately align video frames and audio signals, thereby ensuring temporal consistency between sound field reconstruction and visual information.
[0072] As a further preferred technical solution, the key parameters initialized according to the sound source localization requirements include the scanning plane area, sound speed, sound source frequency, sound imaging frequency constraint and camera resolution, where:
[0073] (1) According to the measured distance between the microphone array and the sound source and the pixel of the camera, the scanning plane area is calculated and divided into 200*200 equally spaced grid points to ensure the imaging accuracy. The image plane grid step size is determined as scan_resolution=0.05m and the imaging scanning plane range is determined as scan_x=[-x c , x c ], scan_y=[-y c ,y c ]m.
[0074] (2) The speed of sound in air: c = 343 m / s.
[0075] (3) This embodiment requires detecting the frequency of sound imaging: scan_freq = [9500 10500] Hz, and the sound source frequency is selected as 10 kHz.
[0076] (4) The camera resolution is 1920×1080.
[0077] (5) The sampling rate of the audio signal is 192kHz, which can collect more data for imaging positioning in a short time.
[0078] This embodiment defines the three-dimensional coordinates (x, y, z) of each microphone based on the geometric layout mic_pos of the microphone array, and calculates the center coordinate mic_centre of the microphone array. The scanning plane grid is generated, and the grid coordinates X and Y within the scanning plane are constructed. A fixed distance z = 3m is defined from the scanning plane to the microphone array. The scanning plane area can be calculated based on the camera pixels. The main research is on single-point sound source and dual-point sound source scenarios. The scanning area is calculated as follows:
[0079]
[0080] Where: M is the camera internal parameter matrix, the camera resolution is 1920×1080, x c 、y c The physical horizontal and vertical coordinates of the image corresponding to the pixel point can be calculated by the above formula (x c ,y c ), so the coordinates of the four focal points in the sound source scanning area are (-x c ,y c ,z)、(x c ,y c ,z)、(-x c , -y c ,z)、(x c , -y c , z).
[0081] As a further preferred technical solution, in step S10, extracting the frequency domain information of the audio signal to construct a cross-spectrum matrix specifically includes the following steps:
[0082] Transferring the audio signal from the time domain to the frequency domain to obtain an audio frequency domain signal;
[0083] The cross spectrum matrix is obtained by multiplying the audio frequency domain signal and its conjugate transpose and dividing the result by the number of frequency frames.
[0084] It should be noted that, in this embodiment, when selecting the imaging frequency, for example, selecting a frequency within the range of 20 kHz-23 kHz, can also be understood as the number of frames within this frequency range.
[0085] Specifically, this embodiment reads the original data from the audio signal collected by the microphone array (mic_signal is used to represent the audio signal) and performs a fast Fourier transform (FFT) to generate a frequency domain audio signal X(k):
[0086] X(k)=FFT(mic_signal)
[0087] It is necessary to constrain the frequency of the desired imaging. Assuming that the sound source frequency is 10kHz, the frequency constraint is scan_freq = [9500 10500] Hz. The cross spectral matrix (CSM) calculated based on the frequency domain audio signal is:
[0088]
[0089] Where k represents the frequency in the frequency domain transform, H represents the conjugate transpose operation of the matrix, and N is the number of frequency points required by the composite.
[0090] It should be noted that the cross-spectral matrix calculated in this embodiment is the frequency domain conjugate transpose of the original signal multiplied by itself. Dividing it by N (the number of frequencies) can normalize the result to a fixed range for easy comparison. The diagonal removal technology is used to eliminate related components to avoid affecting the subsequent sound field reconstruction.
[0091] As a further preferred technical solution, the step S20 of simulating the propagation path of the sound wave in the scanning plane area to generate the steering vector specifically includes the following:
[0092] S21, calculating the propagation path length from the grid points of the scanning plane area division to the microphone array;
[0093] Specifically, this embodiment calculates the propagation path length r from the scanning plane grid point to the microphone:
[0094]
[0095] Where (x, y, z) is the three-dimensional coordinate of the microphone, x mic Indicates the horizontal coordinate of the microphone in the microphone array, y mic represents the vertical coordinate of the microphone in the microphone array.
[0096] S22. Calculate the phase delay of the sound wave based on the propagation path length and the speed of sound to obtain a steering vector.
[0097] Specifically, this embodiment calculates the phase delay of the sound wave based on the propagation path and the speed of sound, thereby calculating the steering vector as:
[0098]
[0099] Where g(x, y, k) represents the steering vector from the scan plane to each microphone in the microphone array, c is the speed of sound, k is the frequency, and j is the imaginary unit.
[0100] Furthermore, this embodiment calculates a weighted steering vector w(x, y, k) based on the steering vector to describe the directionality and energy distribution of the sound wave. The weighted steering vector w(x, y, k) is weighted and corrected according to the distance ratio to perform a normalization function, thereby improving the ability to resolve the target sound source.
[0101]
[0102] As a further preferred technical solution, the step S30: modeling the sound field response characteristics of each point in the scanning plane area to generate a point spread function, specifically includes the following steps:
[0103] S31, calculating the point spread function in the time domain based on the propagation path length and sound speed from the grid points divided into the scanning plane area to the microphone array;
[0104] Specifically, the point spread function (PSF) is used to describe the response characteristics of the system to a point source in the sound field. The calculation formula is:
[0105]
[0106] Where N mic represents the number of microphones, t is time, c is the speed of sound, δ is the unit impulse function, and PSF reflects the superposition effect of this propagation delay in the scanning plane.
[0107] S32 . Convert the point spread function in the time domain into a point spread function in the frequency domain based on the steering vector, and calculate a discrete integral constant based on the point spread function in the frequency domain.
[0108] Specifically, in order to facilitate calculation, this embodiment converts the time domain PSF into the frequency domain and calculates the discrete integral constant a:
[0109] a=∑|PSF(x,y)|
[0110] PSF(x,y)=|w(x,y,k) H g(x,y,k)| 2
[0111] Where a represents the discrete integral constant, PSF(x,y) represents the point spread function in the frequency domain, w(x,y,k) represents the weighted steering vector, g(x,y,k) represents the steering vector, where k represents the frequency in the frequency domain transform, and H represents the conjugate operator.
[0112] This embodiment converts the PSF into the frequency domain to accelerate calculations, suppresses interference components through regularization operations, and calculates residuals and iterative updates through frequency domain convolution, effectively avoiding the high complexity of a large number of matrix operations in spatial domain operations and greatly improving computational efficiency.
[0113] As a further preferred technical solution, Figure 3 As shown, the step S40: using a fast iterative shrinkage threshold algorithm to optimize the point spread function to obtain a sound source distribution map specifically includes the following steps:
[0114] S41, initializing iteration parameters, wherein the iteration parameters include the number of iterations, iteration variables, step factors, and intermediate variables;
[0115] It should be noted that, in this embodiment, the iteration parameters are initialized as follows: number of iterations: 200, initialization iteration variable: q(0)=0, acceleration step factor: t(0)=1, intermediate variable: y(0)=q(0)=0.
[0116] S42. Based on the residual vector and the discrete integral constant and in combination with the non-negative constraint, update the current solution of the nth iteration, where the current solution represents the sound source distribution estimate, wherein the update formula of the current solution is:
[0117]
[0118] Where q(n+1) is the updated result of the current solution q(n) of the nth iteration, y(n) represents the intermediate variable, and DAS represents the beamforming result. w(x,y,k) represents the weighted steering vector, CSM represents the cross spectrum matrix, H represents the conjugate operator, N mic represents the number of microphones; PFS represents the point spread function in the time domain, DAS-PFS*q(n) represents the residual vector, a represents the discrete integration constant, and b(n) represents the result calculated using PFS;
[0119] S43, updating the step size factor and the intermediate variables based on the current solution, and then performing the next iteration;
[0120] Specifically, this embodiment starts updating the acceleration step factor t(n) and the intermediate variable y(n) after updating the current solution q(n):
[0121]
[0122] It should be noted that the acceleration is performed according to the current and current iteration results to avoid the limitation of gradual linear convergence in the deconvolution algorithm, ensuring that the step size gradually increases, thereby accelerating the convergence speed of the solution in each iteration.
[0123] S44. When the number of iterations is reached, a global optimal current solution is obtained as an optimal sound source distribution map.
[0124] This embodiment uses FISTA to rapidly process data and optimize the sound field distribution. It also includes the introduction of an accelerated step-size factor t(n) and an intermediate variable y(n), which allows the update direction to be combined with the prediction of the current and current iteration results, adjusting the next update strategy, thereby more quickly approaching the global optimal solution in each iteration. Furthermore, a dynamic step-size adjustment mechanism is used in the FISTA algorithm to predict the solution update through an accelerated iterative formula, making each update step closer to the optimal solution while avoiding the oscillation phenomenon in traditional fixed-step-size methods.
[0125] As a further preferred technical solution, the CSM in this embodiment is generated by the distribution of point sound sources S(x,y), and we can get:
[0126] CSM∝PSF*S(x,y)
[0127] b(r)∝PSF(x,y)*S(x,y)
[0128] In this embodiment, S(x,y) is replaced by q(r), and the calculation result of DAS can be Converts to:
[0129] b(r)=q*PSF=FFT -1 [FFT[q]FFT[PSF]]
[0130] Further construct the residual vector:
[0131] d(n)=DAS-PSF*q(n)
[0132] Where d(n) represents the residual vector of the nth iteration, DAS is the beamforming result as input observation data, and q(n) represents the current solution (source distribution estimate) of the nth iteration.
[0133] As a further preferred technical solution, after the step S40 of optimizing and solving the point spread function using a fast iterative shrinkage threshold algorithm to obtain an optimal sound source distribution map, the method further comprises the following steps:
[0134] The optimal sound source distribution map is fused with the video signal to obtain visualization information of the sound source location and the surrounding environment.
[0135] It should be noted that this embodiment fuses the sound pressure cloud map with the video capture image and presents the sound field reconstruction result in a visual manner, thereby achieving an intuitive display of the sound source position and the surrounding environment.
[0136] Furthermore, to verify the imaging effect of the algorithm proposed in this embodiment, a single-point sound source simulation experiment was used; the sound source position was located at (0, 0) on the scanning plane, and the vertical distance from the microphone array was 3m; the sound frequency of the single-point sound source was 3000Hz. The imaging effects were compared using the algorithms DAS, DAMAS2, FFT_NNLS, and the algorithm proposed in this invention. The comparison results are shown in Figure 2. Figure 4 As shown, it can be seen that the method of this embodiment significantly improves the imaging resolution.
[0137] To test the imaging effect of the algorithm proposed in this embodiment on multiple sound sources, a dual-point sound source simulation experiment was used; the sound source positions were located at (-1, 0) and (1, 0) on the scanning plane, the vertical distance from the microphone array was 3m, and the sound source interval was 2m; the sound frequency of the single-point sound source was 3000Hz. The imaging effect comparison results using the algorithms DAS, DAMAS2, FFT_NNLS, and the algorithm proposed in this invention are shown in Figure 2. Figure 5 As shown, it can be seen that the method of this embodiment can identify multiple sound sources, and the imaging resolution is improved without spatial aliasing.
[0138] In addition, if Figure 6 As shown, another embodiment of the present invention further provides a fast iterative deconvolution sound source imaging system, the system comprising:
[0139] A signal acquisition module 10 is used to acquire audio signals and video signals generated during the partial discharge process;
[0140] A cross-spectrum matrix construction module 20 is used to extract frequency domain information of the audio signal to construct a cross-spectrum matrix;
[0141] A steering vector calculation module 30 is used to simulate the propagation path of the sound wave in the scanning plane area and generate a steering vector. The scanning plane area is a key parameter obtained according to the sound source localization requirement;
[0142] a point spread function calculation module 40 for modeling the acoustic field response characteristics of each point in the scanning plane area and generating a point spread function;
[0143] The iterative module 50 is used to optimize and solve the point spread function using a fast iterative shrinkage threshold algorithm to obtain an optimal sound source distribution map, wherein during the iteration process, the step factor and the intermediate variable are updated based on the residual vector constructed based on the steering vector and the cross-spectral matrix.
[0144] As a further preferred technical solution, the signal acquisition module 10 is used to acquire the sound pressure signal emitted by the partial discharge in real time, ensuring sampling accuracy and integrity to support subsequent efficient data processing and sound field reconstruction. The signal acquisition module 10 is mainly composed of a microphone array and a data acquisition module; the data acquisition module includes an audio signal acquisition unit and a video signal acquisition unit, respectively, for acquiring the sound pressure signal and real-time visual image information generated during the partial discharge process, thereby realizing the synchronous acquisition and processing of multimodal data during the sound source imaging and localization process.
[0145] The microphone array is a high-density microphone array consisting of 64 high-sensitivity microphones and adopts a spiral array layout design to ensure uniform coverage of the target sound source sound field and accurate collection of multi-directional signals.
[0146] The audio signal acquisition unit is responsible for acquiring high-quality sound pressure signals from the microphone array and performing digital processing to provide basic data support for subsequent sound field reconstruction and sound source localization. Specific functions include the following:
[0147] The multi-channel data acquisition system synchronously collects and digitizes the output signals of all microphones in the array. Supporting sampling rates up to 192kHz, the acquisition system avoids signal distortion and ensures spectral integrity. Each sampling channel is equipped with an analog-to-digital converter (ADC) with a resolution of 16 bits or higher to ensure signal quantization accuracy, and a hardware anti-aliasing filter suppresses high-frequency noise. The system also supports simultaneous acquisition and transmission, utilizing high-speed data transmission interfaces (such as Ethernet or optical fiber) to transmit microphone signals in real time to the data processing module for analysis, eliminating data delays.
[0148] The video signal acquisition unit is used to obtain real-time visual information of the partial discharge test site, providing auxiliary reference for sound field positioning and environmental monitoring. Specifically, it is equipped with a high-definition camera (resolution of 1920×1080 or higher, frame rate of 30fps to 60fps) to record dynamic images of the partial discharge test site in real time.
[0149] Furthermore, the video signal acquisition unit and the audio acquisition unit share the same time reference, and use timestamps to accurately align video frames and audio signals to ensure the temporal consistency of sound field reconstruction and visual information.
[0150] As a further preferred technical solution, the cross-spectrum matrix construction module 20 includes:
[0151] A time-frequency domain transfer unit, configured to transfer the audio signal from the time domain to the frequency domain to obtain an audio frequency domain signal;
[0152] The cross spectrum matrix calculation unit is used to multiply the audio frequency domain signal by its conjugate transpose and then divide the result by the number of frequency frames to obtain the cross spectrum matrix.
[0153] As a further preferred technical solution, the steering vector calculation module 30 includes:
[0154] A propagation path length calculation unit, used to calculate the propagation path length from the grid points of the scanning plane area division to the microphone array;
[0155] The steering vector calculation unit is used to calculate the phase delay of the sound wave according to the propagation path length and the sound speed to obtain the steering vector.
[0156] As a further preferred technical solution, the point spread function calculation module 40 includes:
[0157] A time domain point spread function calculation unit, configured to calculate a time domain point spread function based on a propagation path length and a sound velocity from a grid point divided into a scanning plane area to a microphone array;
[0158] The discrete integral constant calculation unit is configured to convert a point spread function in the time domain into a point spread function in the frequency domain based on the steering vector, and calculate a discrete integral constant based on the point spread function in the frequency domain.
[0159] As a further preferred technical solution, the iteration module 50 is specifically configured to:
[0160] Initializing iteration parameters, including the number of iterations, iteration variables, step factors, and intermediate variables;
[0161] Based on the residual vector and the discrete integral constant and in combination with the non-negative constraint, the current solution of the nth iteration is updated, where the current solution represents the sound source distribution estimate, wherein the update formula of the current solution is:
[0162]
[0163] Where q(n+1) is the updated result of the current solution q(n) of the nth iteration, y(n) represents the intermediate variable, and DAS represents the beamforming result. w(x,y,k) represents the weighted steering vector, CSM represents the cross spectrum matrix, H represents the conjugate operator, N mic represents the number of microphones; PFS represents the point spread function in the time domain, DAS-PFS*q(n) represents the residual vector, a represents the discrete integration constant, and b(n) represents the result calculated using PFS;
[0164] Update the step size factor and the intermediate variables based on the current solution before proceeding to the next iteration;
[0165] When the number of iterations is reached, a global optimal current solution is obtained as the optimal sound source distribution map.
[0166] As a further preferred technical solution, the system further includes:
[0167] The display module fuses the sound pressure cloud map with the video capture image to visualize the sound source location and surrounding environment.
[0168] It should be noted that other embodiments or specific implementation methods of the fast iterative deconvolution sound source imaging system of the present invention can refer to the above-mentioned method embodiments and will not be repeated here.
[0169] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0170] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0171] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are illustrative and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.
Claims
1. A fast iterative deconvolution sound source imaging method, characterized in that: include: Acquiring audio signals and video signals generated during a partial discharge process, extracting frequency domain information of the audio signals and constructing a cross-spectrum matrix; Simulate the propagation path of the sound wave within the scanning plane area, which is a key parameter initialized according to the sound source localization requirements, and generate a steering vector; Modeling the acoustic field response characteristics of each point in the scanning plane area to generate a point spread function; A fast iterative shrinkage threshold algorithm is used to optimize and solve the point spread function to obtain an optimal sound source distribution map, wherein the step size factor and the intermediate variable are updated during the iteration process based on the residual vector constructed based on the steering vector and the cross-spectral matrix.
2. The fast iterative deconvolution sound source imaging method according to claim 1, wherein: The step of obtaining the audio signal and the video signal generated during the partial discharge process includes: Use microphone array to collect multi-directional sound field signals of partial discharge sound sources; Synchronously collecting and digitally processing the output signals of all microphones in the microphone array to obtain the audio signal, wherein the audio signal is a sound pressure signal; A camera is used to detect video signals at the scene of partial discharge, and the video frames and audio signals in the video signals are aligned using time stamps.
3. The fast iterative deconvolution sound source imaging method according to claim 1, wherein: Key parameters initialized according to sound source localization requirements also include sound speed, sound source frequency, sound imaging frequency constraint, and camera resolution; The scanning plane area is calculated based on the measured distance between the microphone array and the sound source and the pixels of the camera, and the scanning plane area is divided into equally spaced grid points to determine the grid step size.
4. The fast iterative deconvolution sound source imaging method according to claim 1, wherein: Extracting frequency domain information of the audio signal to construct a cross-spectrum matrix includes: Transferring the audio signal from the time domain to the frequency domain to obtain an audio frequency domain signal; The cross spectrum matrix is obtained by multiplying the audio frequency domain signal and its conjugate transpose and dividing the result by the number of frequency frames.
5. The fast iterative deconvolution sound source imaging method according to claim 1, wherein: The simulating the propagation path of the sound wave in the scanning plane area to generate the steering vector includes: Calculate the propagation path length from the grid points of the scan plane area division to the microphone array; According to the propagation path length and the speed of sound, the phase delay of the sound wave is calculated to obtain the steering vector.
6. The fast iterative deconvolution sound source imaging method according to claim 1, wherein: The step of modeling the acoustic field response characteristics of each point in the scanning plane region to generate a point spread function includes: Calculate the point spread function in the time domain based on the propagation path length and sound speed from the grid points divided into the scanning plane area to the microphone array; A point spread function in the time domain is converted into a point spread function in the frequency domain based on the steering vector, and a discrete integral constant is calculated based on the point spread function in the frequency domain.
7. The fast iterative deconvolution sound source imaging method according to claim 6, wherein: The calculation formula of the discrete integral constant is: PSF(x,y)=|w(x,y,k) H g(x,y,k)| 2 Where a represents the discrete integral constant, PSF(x,y) represents the point spread function in the frequency domain, w(x,y,k) represents the weighted steering vector, g(x,y,k) represents the steering vector, where k represents the frequency in the frequency domain transform, and H represents the conjugate operator.
8. The fast iterative deconvolution sound source imaging method according to claim 1, wherein: The method of optimizing the point spread function using a fast iterative shrinkage threshold algorithm to obtain a sound source distribution map includes: Initializing iteration parameters, including the number of iterations, iteration variables, step factors, and intermediate variables; Based on the residual vector and the discrete integral constant and in combination with the non-negative constraint, the current solution of the nth iteration is updated, where the current solution represents the sound source distribution estimate, wherein the update formula of the current solution is: Where q(n+1) is the updated result of the current solution q(n) of the nth iteration, y(n) represents the intermediate variable, and DAS represents the beamforming result. w(x,y,k) represents the weighted steering vector, CSM represents the cross spectrum matrix, H represents the conjugate operator, N mic represents the number of microphones; PFS represents the point spread function in the time domain, DAS-PFS*q(n) represents the residual vector, a represents the discrete integration constant, and b(n) represents the result calculated using PFS; Update the step size factor and the intermediate variables based on the current solution before proceeding to the next iteration; When the number of iterations is reached, a global optimal current solution is obtained as the optimal sound source distribution map.
9. The fast iterative deconvolution sound source imaging method according to claims 1 to 8, characterized in that: After optimizing and solving the point spread function using a fast iterative shrinkage threshold algorithm to obtain an optimal sound source distribution map, the method further includes: The optimal sound source distribution map is fused with the video signal to obtain visualization information of the sound source location and the surrounding environment.
10. A fast iterative deconvolution sound source imaging system, characterized in that: include: A signal acquisition module, used to acquire audio signals and video signals generated during the partial discharge process; A cross-spectrum matrix construction module, used to extract frequency domain information of the audio signal to construct a cross-spectrum matrix; A steering vector calculation module is used to simulate the propagation path of sound waves in a scanning plane area and generate a steering vector. The scanning plane area is a key parameter obtained based on the sound source localization requirement. A point spread function calculation module is used to model the acoustic field response characteristics of each point in the scanning plane area and generate a point spread function; An iterative module is used to optimize and solve the point spread function using a fast iterative shrinkage threshold algorithm to obtain an optimal sound source distribution map, wherein during the iterative process, a step factor and an intermediate variable are updated based on a residual vector constructed based on the steering vector and the cross-spectral matrix.
Citation Information
Cited By
Fast deconvolution sound source localization method based on array sparse and broadband comprehensive processing
CN121148416A