Multi-source localization and imaging method based on sparse representation and variational bayesian inference
By combining sparse representation and variational Bayesian inference with beamforming, signal gradient optimization, and sparse dictionary learning, the problems of positioning error and long computation time in multi-sound source localization technology in traditional sound source localization technology are solved, and high-precision, real-time multi-sound source localization and imaging are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SOUTHEAST UNIV
- Filing Date
- 2025-01-13
- Publication Date
- 2026-06-02
AI Technical Summary
Traditional sound source localization and imaging technologies suffer from problems such as localization errors, long calculation times, and inability to display localization results in real time in multi-sound-source and complex sound field environments.
A method based on sparse representation and variational Bayesian inference is adopted. An initial sound intensity matrix is generated by conventional beamforming algorithm, and signal intensity gradient optimization is used. Combined with sparse dictionary learning and variational Bayesian inference, high-precision multi-source localization and imaging are achieved.
It achieves high-precision real-time positioning and imaging of multiple sound sources in complex sound fields, with high spatial resolution and anti-interference capability.
Smart Images

Figure CN120070659B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of acoustic imaging technology, specifically relating to a multi-sound source localization and imaging method based on sparse representation and variational Bayesian inference. Background Technology
[0002] Traditional sound source localization and imaging techniques face numerous challenges in signal processing. For example, while traditional beamforming algorithms can provide basic localization capabilities, they suffer from significant shortcomings in resolution, interference resistance, and computational efficiency, especially in multi-source and complex sound field environments, often failing to meet practical needs. To overcome these issues, improved beamforming algorithms, such as deconvolution beamforming, have emerged in recent years, significantly improving the accuracy of localization results. However, most current improved algorithms still suffer from localization bias in multi-source scenarios, long computation times leading to insufficient real-time localization, and the inability to comprehensively and intuitively represent the localization results. Summary of the Invention
[0003] To address the aforementioned issues, this invention discloses a multi-source localization and imaging method based on sparse representation and variational Bayesian inference, which can achieve high-precision, interference-resistant real-time sound source localization and imaging in multi-source scenarios.
[0004] To achieve the above objectives, the technical solution of the present invention is as follows:
[0005] A multi-source localization and imaging method based on sparse representation and variational Bayesian inference includes the following steps:
[0006] (1) Use conventional beamforming algorithms to generate an initial sound intensity matrix at low resolution and estimate the approximate location of the sound source;
[0007] (2) Calculate the signal intensity gradient based on the approximate location of the current sound source, and update the search position along the gradient direction to obtain the accurate sound source location and high-resolution sound intensity matrix.
[0008] (3) Use the sparse dictionary learning algorithm to process the high-resolution acoustic intensity matrix obtained above, and perform alternating optimization of sparse coding and dictionary updating on the matrix to reduce signal redundancy and enhance signal sparsity.
[0009] (4) Construct a variational Bayesian inference model based on the sparse coefficients, select the variational distribution, optimize the variational lower bound, and then perform posterior estimation of the sound source to obtain accurate localization results under multiple sound sources.
[0010] (5) The positioning results are fused with the camera images to obtain the visualized location of the sound source.
[0011] Furthermore, the method for generating the initial acoustic intensity matrix using the conventional beamforming algorithm in step (1) is as follows:
[0012] Assuming the sound field signal at a certain point r in space is determined by the sound source signal s(t) and the response of the array sensor, the received signal model is as follows:
[0013] ;
[0014] in, It is an array receiving signals. For the number of array sensors, It is the first The signal from the sound source It is the first Location of the sound source The corresponding weighted vector, It is noise.
[0015] Transforming the signal to the frequency domain yields:
[0016] ;
[0017] in, Indicates the first The frequency response of a sound source This represents the frequency response of the noise term.
[0018] The beamforming output can be obtained by using a delay-summation algorithm:
[0019] ;
[0020] in, It is a weighted vector.
[0021] The beamforming output energy is accumulated in the frequency domain and spatial grid to obtain the initial acoustic intensity matrix:
[0022] ;
[0023] in, These are the grid points of the search space. For the frequency range of interest, The frequency of the sound source at the current search location.
[0024] In the initial stage, the resolution of the grid points is relatively low. This is the initial sound intensity matrix, which provides the approximate location of the sound source for subsequent refined searches.
[0025] Furthermore, the method for searching the precise location of the sound source along the gradient direction and obtaining the high-resolution sound intensity matrix in step (2) is as follows:
[0026] Since the initial sound intensity matrix is a spatial position The function whose gradient Sound intensity Position The partial derivative is defined as:
[0027] ;
[0028] The finite difference method can be used to... The directional gradient is approximated as:
[0029] ;
[0030] in, for The unit vector of direction can be obtained similarly. and Approximate gradient of the direction.
[0031] Update the search position using gradient descent based on the gradient direction:
[0032] ;
[0033] in, This is the step size parameter, which can be set from 0.01 to 0.1, and is used to control the update magnitude.
[0034] The iteration stops when the gradient norm approaches zero.
[0035] ;
[0036] The minimum threshold value is set.
[0037] At this point, a high-resolution sound intensity matrix can be obtained. and accurate sound source location .
[0038] Furthermore, the method for learning the sparsity of the source signals in the intensity matrix using a sparse dictionary in step (3) is as follows:
[0039] The high-resolution acoustic intensity matrix obtained in the previous steps The sparse representation is:
[0040] ;
[0041] in, It is a dictionary matrix. It is a sparse coefficient matrix. This is a noise or error term.
[0042] The objective function for sparse optimization is set as follows:
[0043] ;
[0044] in, It is the Frobenius norm. For sparse regularization, For sparse regularization parameters.
[0045] In the sparse encoding process, the dictionary is first fixed. Solve for sparsity coefficients The goal of sparse coding is: ;
[0046] Solving for sparse coefficients using convex optimization:
[0047] ;
[0048] Then fix the sparsity coefficient. Update the dictionary The dictionary update target is:
[0049] ;
[0050] The dictionary is solved using a global update method:
[0051] ;
[0052] Then, the sparse coding and dictionary update steps are repeated to alternately optimize the dictionary and sparse coefficients until the iteration stopping condition is met: ;
[0053] The minimum threshold value is set.
[0054] The sparse-optimized matrix is obtained as follows:
[0055] ;
[0056] The sparse matrix at this time It is a compressed matrix that still retains the main sound source characteristics, and can more efficiently and accurately characterize the sound field distribution.
[0057] Furthermore, in step (4), the variational Bayesian inference framework is introduced to calculate the accurate location of multiple sound sources as follows:
[0058] The sparse-optimized sound intensity matrix was obtained earlier. By introducing the latent variable Z, a probability model is established, and the joint distribution of the sound intensity matrix is constructed as follows: ;
[0059] in, It follows a Gaussian distribution. For latent variables The prior distribution of .
[0060] Choose variational distribution To approximate the posterior Variational distributions are generally in factorized form:
[0061] ;
[0062] Approximating the log-marginal likelihood by optimizing the variational lower bound:
[0063] ;
[0064] The first term is the expectation of the joint distribution, and the second term is the entropy of the variational distribution.
[0065] Decompose the joint distribution into the conditional distribution and the prior:
[0066] ;
[0067] The optimization objective at this point is to maximize Let the variational distribution The update formula is:
[0068] ;
[0069] in, In addition to All potential variables outside of this.
[0070] After optimization Maximum a posteriori probability estimation of sound source location:
[0071] ;
[0072] The final set of sparse sound source locations is:
[0073] .
[0074] Furthermore, the method for visualizing the multi-source localization results in step (5) is as follows:
[0075] The basic formula for mapping camera images from three-dimensional space to a two-dimensional image plane is: ;
[0076] in, This represents the homogeneous coordinates of the student's location in the world coordinate system. This represents the homogeneous pixel coordinates of the sound source in the image coordinate system. This is the intrinsic parameter matrix of the camera. This is the camera extrinsic parameter matrix, which can be obtained through camera calibration.
[0077] Camera Intrinsic Matrix for:
[0078] ;
[0079] in, and For camera focal length, and These are the coordinates of the principal point of the image.
[0080] Regarding the previously obtained sound source location Perform the transformation:
[0081] ;
[0082] The pixel coordinates are:
[0083] ;
[0084] Here and This refers to the location of the sound source in the image.
[0085] Based on the sparse sound intensity matrix The corresponding values are displayed on the image to show the intensity of each sound source. The color gradient is used to represent the intensity of the sound source. OpenCV is used to overlay the labeled sound sources onto the camera image to form a visual output.
[0086] The beneficial effects of this invention are as follows:
[0087] The present invention describes a multi-source localization and imaging method based on sparse representation and variational Bayesian inference. By combining low-resolution preliminary localization, signal intensity gradient optimization, sparse dictionary learning and variational Bayesian inference, it achieves high-precision real-time localization and imaging of multiple sound sources in complex sound fields, and has high spatial resolution and good anti-interference ability. Attached Figure Description
[0088] Figure 1 This is a flowchart illustrating the implementation of the method of the present invention;
[0089] Figure 2 This is a flowchart of the gradient search process in the method of this invention;
[0090] Figure 3 This is a flowchart of sparse dictionary learning in the method of the present invention;
[0091] Figure 4 This is a flowchart of variational Bayesian inference in the method of the present invention. Detailed Implementation
[0092] The present invention will be further illustrated below with reference to the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.
[0093] like Figure 1 As shown, the multi-source localization and imaging method based on sparse representation and variational Bayesian inference of the present invention includes the following steps:
[0094] (1) Use conventional beamforming algorithms to generate an initial sound intensity matrix at low resolution and estimate the approximate location of the sound source;
[0095] The conventional beamforming algorithm generates the initial acoustic intensity matrix as follows:
[0096] Assuming the sound field signal at a certain point r in space is determined by the sound source signal s(t) and the response of the array sensor, the received signal model is as follows:
[0097] ;
[0098] in, It is an array receiving signals. For the number of array sensors, It is the first The signal from the sound source It is the first Location of the sound source The corresponding weighted vector, It is noise.
[0099] Transforming the signal to the frequency domain yields:
[0100] ;
[0101] in, Indicates the first The frequency response of a sound source This represents the frequency response of the noise term.
[0102] The beamforming output can be obtained by using a delay-summation algorithm:
[0103] ;
[0104] in, It is a weighted vector.
[0105] The beamforming output energy is accumulated in the frequency domain and spatial grid to obtain the initial acoustic intensity matrix:
[0106] ;
[0107] in, These are the grid points of the search space. For the frequency range of interest, The frequency of the sound source at the current search location.
[0108] In the initial stage, the resolution of the grid points is relatively low. This is the initial sound intensity matrix, which provides the approximate location of the sound source for subsequent refined searches.
[0109] (2) Calculate the signal intensity gradient based on the approximate location of the current sound source, and update the search position along the gradient direction to obtain the accurate sound source location and the high-resolution sound intensity matrix;
[0110] like Figure 2 As shown, the method for searching the precise location of the sound source along the gradient direction and obtaining the high-resolution sound intensity matrix is as follows:
[0111] Since the initial sound intensity matrix is a spatial position The function whose gradient Sound intensity Position The partial derivative is defined as:
[0112] ;
[0113] The finite difference method can be used to... The directional gradient is approximated as:
[0114] ;
[0115] in, for The unit vector of direction can be obtained similarly. and Approximate gradient of the direction.
[0116] Update the search position using gradient descent based on the gradient direction:
[0117] ;
[0118] in, This is the step size parameter, which can be set from 0.01 to 0.1, and is used to control the update magnitude.
[0119] The iteration stops when the gradient norm approaches zero.
[0120] ;
[0121] The minimum threshold value is set.
[0122] At this point, a high-resolution sound intensity matrix can be obtained. and accurate sound source location .
[0123] (3) Use the sparse dictionary learning algorithm to process the high-resolution acoustic intensity matrix obtained above, and perform alternating optimization of sparse coding and dictionary update on the matrix to reduce signal redundancy and enhance signal sparsity.
[0124] like Figure 3 As shown, the method for learning the sparsity of source signals in the enhanced acoustic intensity matrix using a sparse dictionary is as follows:
[0125] The high-resolution acoustic intensity matrix obtained in the previous steps The sparse representation is:
[0126] ;
[0127] in, It is a dictionary matrix. It is a sparse coefficient matrix. This is a noise or error term.
[0128] The objective function for sparse optimization is set as follows:
[0129] ;
[0130] in, It is the Frobenius norm. For sparse regularization, For sparse regularization parameters.
[0131] In the sparse encoding process, the dictionary is first fixed. Solve for sparsity coefficients The goal of sparse coding is: ;
[0132] Solving for sparse coefficients using convex optimization:
[0133] ;
[0134] Then fix the sparsity coefficient. Update the dictionary The dictionary update target is:
[0135] ;
[0136] The dictionary is solved using a global update method:
[0137] ;
[0138] Then, the sparse coding and dictionary update steps are repeated to alternately optimize the dictionary and sparse coefficients until the iteration stopping condition is met: ;
[0139] The minimum threshold value is set.
[0140] The sparse-optimized matrix is obtained as follows:
[0141] ;
[0142] The sparse matrix at this time It is a compressed matrix that still retains the main sound source characteristics, and can more efficiently and accurately characterize the sound field distribution.
[0143] (4) Construct a variational Bayesian inference model based on sparse coefficients, select the variational distribution, optimize the variational lower bound, and then perform posterior estimation of the sound source to obtain accurate localization results under multiple sound sources.
[0144] like Figure 4 As shown, the variational Bayesian inference framework calculates the accurate location of multiple sound sources using the following method:
[0145] The sparse-optimized sound intensity matrix was obtained earlier. By introducing the latent variable Z, a probability model is established, and the joint distribution of the sound intensity matrix is constructed as follows: ;
[0146] in, It follows a Gaussian distribution. For latent variables The prior distribution of .
[0147] Choose variational distribution To approximate the posterior Variational distributions are generally in factorized form:
[0148] ;
[0149] Approximating the log-marginal likelihood by optimizing the variational lower bound:
[0150] ;
[0151] The first term is the expectation of the joint distribution, and the second term is the entropy of the variational distribution.
[0152] Decompose the joint distribution into the conditional distribution and the prior:
[0153] ;
[0154] The optimization objective at this point is to maximize Let the variational distribution The update formula is:
[0155] ;
[0156] in, In addition to All potential variables outside of this.
[0157] After optimization Maximum a posteriori probability estimation of sound source location:
[0158] ;
[0159] The final set of sparse sound source locations is:
[0160] .
[0161] (5) The positioning results are fused with the camera images to obtain the visualized location of the sound source.
[0162] The visualization method for multi-source localization results is as follows:
[0163] The basic formula for mapping camera images from three-dimensional space to a two-dimensional image plane is: ;
[0164] in, This represents the homogeneous coordinates of the student's location in the world coordinate system. This represents the homogeneous pixel coordinates of the sound source in the image coordinate system. This is the intrinsic parameter matrix of the camera. This is the camera extrinsic parameter matrix, which can be obtained through camera calibration.
[0165] Camera Intrinsic Matrix for:
[0166] ;
[0167] in, and For camera focal length, and These are the coordinates of the principal point of the image.
[0168] Regarding the previously obtained sound source location Perform the transformation:
[0169] ;
[0170] The pixel coordinates are:
[0171] ;
[0172] Here and This refers to the location of the sound source in the image.
[0173] Based on the sparse sound intensity matrix The corresponding values are displayed on the image to show the intensity of each sound source. The color gradient is used to represent the intensity of the sound source. OpenCV is used to overlay the labeled sound sources onto the camera image to form a visual output.
[0174] The present invention describes a multi-source localization and imaging method based on sparse representation and variational Bayesian inference. By combining low-resolution preliminary localization, signal intensity gradient optimization, sparse dictionary learning and variational Bayesian inference, it achieves high-precision real-time localization and imaging of multiple sound sources in complex sound fields, and has high spatial resolution and good anti-interference ability.
[0175] It should be noted that the above content merely illustrates the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. For those skilled in the art, various improvements and modifications can be made without departing from the principle of the present invention, and all such improvements and modifications fall within the scope of protection of the claims of the present invention.
Claims
1. A multi-source localization and imaging method based on sparse representation and variational Bayesian inference, characterized in that, Includes the following steps: (1) Use conventional beamforming algorithms to generate an initial sound intensity matrix at low resolution and estimate the approximate location of the sound source; (2) Calculate the signal intensity gradient based on the approximate location of the current sound source, and update the search position along the gradient direction to obtain the accurate sound source location and high-resolution sound intensity matrix. (3) Use the sparse dictionary learning algorithm to process the high-resolution acoustic intensity matrix obtained in step (2), and perform alternating optimization of sparse coding and dictionary update on the matrix to reduce signal redundancy and enhance signal sparsity. Specifically as follows: High-resolution acoustic intensity matrix The sparse representation is: ; in, It is a dictionary matrix. It is a sparse coefficient matrix. This is a noise or error term; The objective function for sparse optimization is set as follows: ; in, It is the Frobenius norm. For sparse regularization, For sparse regularization parameters; In the sparse encoding process, the dictionary is first fixed. Solve for sparsity coefficients The goal of sparse coding is: ; Solving for sparse coefficients using convex optimization: ; Then fix the sparsity coefficient. Update the dictionary The dictionary update target is: ; The dictionary is solved using a global update method: ; Then, the sparse coding and dictionary update steps are repeated to alternately optimize the dictionary and sparse coefficients until the iteration stopping condition is met: ; The set minimum threshold; The sparse-optimized matrix is obtained as follows: ; The sparse matrix at this time It is a compressed matrix that still retains the main sound source characteristics, and can more efficiently and accurately characterize the sound field distribution; (4) Construct a variational Bayesian inference model based on sparse coefficients, select the variational distribution, optimize the variational lower bound, and then perform posterior estimation of the sound source to obtain accurate localization results under multiple sound sources. Specifically as follows: Sparse optimized sound intensity matrix By introducing the latent variable Z, a probability model is established, and the joint distribution of the sound intensity matrix is constructed as follows: ; in, It follows a Gaussian distribution. For latent variables The prior distribution; Choose variational distribution To approximate the posterior Variational distributions are generally in factorized form: ; Approximating the log-marginal likelihood by optimizing the variational lower bound: ; The first term is the expectation of the joint distribution, and the second term is the entropy of the variational distribution. Decompose the joint distribution into the conditional distribution and the prior: ; The optimization objective at this point is to maximize Let the variational distribution The update formula is: ; in, In addition to All potential variables outside; After optimization Maximum a posteriori probability estimation of sound source location: ; The final set of sparse sound source locations is: ; (5) The positioning results are fused with the camera images to obtain the visualized location of the sound source.
2. The multi-source localization and imaging method based on sparse representation and variational Bayesian inference according to claim 1, characterized in that, The method for generating the initial acoustic intensity matrix using the conventional beamforming algorithm in step (1) is as follows: Assuming the sound field signal at a certain point r in space is determined by the sound source signal s(t) and the response of the array sensor, the received signal model is as follows: ; in, It is an array receiving signals. For the number of array sensors, It is the first The signal from the sound source It is the first Location of the sound source The corresponding weighted vector, For noise; Transforming the signal to the frequency domain yields: ; in, Indicates the first The frequency response of a sound source The frequency response of the noise term; The beamforming output is obtained by using a delay-summation algorithm: ; in, It is a weighted vector; The beamforming output energy is accumulated in the frequency domain and spatial grid to obtain the initial acoustic intensity matrix: ; in, These are the grid points of the search space. For the frequency range of interest, The frequency of the sound source at the current search location; In the initial stage, the resolution of the grid points is relatively low. This is the initial sound intensity matrix, which provides the approximate location of the sound source for subsequent refined searches.
3. The multi-source localization and imaging method based on sparse representation and variational Bayesian inference according to claim 1, characterized in that, The method for searching the precise location of the sound source along the gradient direction and obtaining the high-resolution sound intensity matrix in step (2) is as follows: Since the initial sound intensity matrix is a spatial position The function whose gradient Sound intensity Position The partial derivative is defined as: ; Using the finite difference method The directional gradient is approximated as: ; in, for The unit vector of direction, similarly obtained and Approximate gradient of direction; Update the search position using gradient descent based on the gradient direction: ; in, The step size parameter, set to 0.01 to 0.1, is used to control the update magnitude; The iteration stops when the gradient norm approaches zero. ; The set minimum threshold; At this point, a high-resolution sound intensity matrix is obtained. and accurate sound source location .
4. The multi-source localization and imaging method based on sparse representation and variational Bayesian inference according to claim 1, characterized in that, The method for visualizing the sound source localization results in step (5) is as follows: The basic formula for mapping camera images from three-dimensional space to a two-dimensional image plane is: ; in, This represents the homogeneous coordinates of the student's location in the world coordinate system. This represents the homogeneous pixel coordinates of the sound source in the image coordinate system. This is the intrinsic parameter matrix of the camera. This is the camera extrinsic parameter matrix, obtained through camera calibration; Camera Intrinsic Matrix for: ; in, and For camera focal length, and The coordinates of the principal point in the image; The obtained sound source location Perform the transformation: ; The pixel coordinates are: ; and The location of the sound source in the image; Based on the sparse sound intensity matrix The corresponding values are displayed on the image to show the intensity of each sound source. The color gradient is used to represent the intensity of the sound source. OpenCV is used to overlay the labeled sound sources onto the camera image to form a visual output.