Multi-sound-source localization and imaging method based on sparse representation and variational Bayesian inference

By combining sparse characterization and variational Bayesian inference technology, the problem of positioning deviation and insufficient computing efficiency in the multi-sound source environment of traditional sound source positioning technology is solved, and a high-precision, anti-interference and multi-sound source positioning and imaging effect is achieved.

CN120070659AActive Publication Date: 2025-05-30SOUTHEAST UNIV

Patent Information

Application Number
CN202510046011.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-30
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Traditional sound source positioning and imaging technologies have problems such as positioning deviation, long calculation time and unintuitive positioning results in multi-sound sources and complex sound field environments.

Method used

Multi-sound source positioning and imaging methods based on sparse characterization and variational Bayes inference are adopted, and real-time sound source positioning and imaging with high precision, anti-interference and strong real-time sound source positioning and imaging through low-resolution preliminary positioning, signal intensity gradient optimization, sparse dictionary learning and variational Bayes inference are achieved.

Benefits of technology

High-precision real-time positioning and imaging of multiple sound sources is realized in complex sound fields, with high spatial resolution and good anti-interference ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070659A_ABST
    Figure CN120070659A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-sound-source localization and imaging method based on sparse representation and variational Bayesian inference, and the method comprises the following steps: (1), generating an initial sound intensity matrix at a low resolution through employing a conventional beam forming algorithm, and estimating the position of a sound source; (2) calculating a signal intensity gradient, updating a search position along the gradient direction, and obtaining an accurate sound source position and a high-resolution sound intensity matrix; (3) carrying out sparse coding and dictionary updating optimization on the high-resolution sound intensity matrix by using a sparse dictionary learning algorithm; (4) constructing a variational Bayesian inference model based on a sparse coefficient, optimizing a variational lower bound, and performing posterior positioning estimation of multiple sound sources; and (5) fusing a positioning result with a camera image to obtain a visual position of the sound source. According to the method, high-precision real-time positioning and imaging are realized in a complex sound field by combining low-resolution positioning, gradient optimization, sparse dictionary learning and Bayesian inference, and the method has relatively high spatial resolution and relatively strong anti-interference capability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of acoustic imaging, and particularly relates to a multi-source localization and imaging method based on sparse representation and variational Bayesian inference. Background Art

[0002] Traditional sound source localization and imaging technologies face many challenges in signal processing. For example, although the traditional beamforming algorithm can provide basic localization functions, it has obvious deficiencies in terms of resolution, anti-interference ability, and computational efficiency. Especially in the multi-source and complex sound field environment, it often fails to meet the actual needs. To overcome these problems, in recent years, algorithms such as the deconvolution beamforming method have emerged to improve the traditional beamforming. Such methods can relatively well improve the accuracy of the localization results. However, most current improved algorithms still have problems such as localization deviation in the case of multiple sound sources, insufficient real-time performance of localization due to long calculation time, and the inability to comprehensively and intuitively display the localization results. Summary of the Invention

[0003] To solve the above problems, the present invention discloses a multi-source localization and imaging method based on sparse representation and variational Bayesian inference, which can achieve high-precision and strong anti-interference real-time sound source localization and imaging in a multi-source scenario.

[0004] To achieve the above object, the technical solution of the present invention is as follows:

[0005] A multi-source localization and imaging method based on sparse representation and variational Bayesian inference includes the following steps:

[0006] (1) Use the conventional beamforming algorithm to generate an initial sound intensity matrix at low resolution and estimate the approximate position of the sound source;

[0007] (2) Calculate the signal intensity gradient according to the current approximate position of the sound source, and update the search position along the gradient direction to obtain the accurate sound source position and a high-resolution sound intensity matrix;

[0008] (3) Use the sparse dictionary learning algorithm to process the high-resolution sound intensity matrix obtained above, and alternately optimize the sparse coding and dictionary update of the matrix to reduce signal redundant information and strengthen signal sparsity;

[0009] (4) Construct a variational Bayesian inference model according to the sparse coefficients, select a variational distribution, optimize the variational lower bound, and then perform sound source posterior estimation to obtain the precise localization results under multiple sound sources;

[0010] (5) Fuse the localization results with the camera image to obtain the visual position results of the sound source.

[0011] Further, the method for generating the initial sound intensity matrix by the conventional beamforming algorithm in step (1) is as follows:

[0012] Assuming that the sound field signal at a certain point r in space is determined by the sound source signal s(t) and the response of the array sensor, the received signal model is:

[0013]

[0014] Where x(t)∈R M is the array receiving signal, M is the number of array sensors, s k (t) is the signal of the kth sound source, a(r k ) is the position r of the kth sound source k The corresponding weight vector, n(t) is the noise.

[0015] Transforming the signal to the frequency domain yields:

[0016]

[0017] Among them, S k (ω) represents the frequency response of the kth sound source, and N(ω) represents the frequency response of the noise term.

[0018] Using the delay-sum algorithm, the beamforming output can be obtained as:

[0019] I(r)=w H (r)X(ω)

[0020] in, is the weight vector.

[0021] The beamforming output energy is accumulated in the frequency domain and spatial grid to obtain the initial sound intensity matrix:

[0022]

[0023] Among them, r g is the grid point of the search space, is the frequency range of interest, and f is the sound source frequency at the current search position.

[0024] In the initial stage, the resolution of the grid points is low. init This is the initial sound intensity matrix, which provides the approximate location of the sound source for subsequent refined search.

[0025] Furthermore, the method for searching the precise location of the sound source along the gradient direction in step (2) and obtaining a high-resolution sound intensity matrix is ​​as follows:

[0026] Since the initial sound intensity matrix is ​​a function of the spatial position r = (x, y, z), its gradient is the partial derivative of the sound intensity I(r) with respect to the position r, defined as:

[0027]

[0028] The gradient in the x - direction can be approximated using the finite - difference method as follows:

[0029]

[0030] where \(e\) x is the unit vector in the x - direction. Similarly, the approximate gradients in the y - and z - directions can be obtained.

[0031] Update the search position using the gradient - descent method along the gradient direction:

[0032]

[0033] where \(\eta\) is the step - size parameter, which can be set between 0.01 and 0.1 and is used to control the update amplitude.

[0034] Set the iteration to stop when the gradient norm approaches zero:

[0035]

[0036] I is the set minimum - value threshold.

[0037] At this time, the high - resolution sound - intensity matrix \(I\) HR (r) and the accurate sound - source position \(r\) * can be obtained.

[0038] Furthermore, the method for enhancing the sparsity of the source signal in the sound - intensity matrix using sparse dictionary learning in step (3) is as follows:

[0039] The sparse representation of the high - resolution sound - intensity matrix \(I\) HR obtained in the previous steps is:

[0040] \(I\) HR = DS+E

[0041] where D is the dictionary matrix, S is the sparse - coefficient matrix, and E is the noise or error term.

[0042] Set the sparse - optimization objective function as:

[0043]

[0044] where is the Frobenius norm, \(\|S\|\) 1 is the sparse - regularization term, and \(\lambda\) is the sparse - regularization parameter.

[0045] During the sparse - coding process, first fix the dictionary D and solve for the sparse coefficient S. The sparse - coding objective is:

[0046]

[0047] Solving for the sparse coefficients using convex optimization:

[0048]

[0049] Subsequently, fix the sparse coefficients S and update the dictionary D. The dictionary update objective is:

[0050]

[0051] Adopt a global update method to solve for the dictionary:

[0052] D = I HR S T (SS T ) -1

[0053] Subsequently, repeat the sparse coding and dictionary update steps to achieve alternating optimization of the dictionary and sparse coefficients until the iteration stop condition is met:

[0054]

[0055] ∈ is the set minimum threshold value.

[0056] The matrix obtained after sparse optimization is:

[0057] I sparse = DS

[0058] The sparse matrix I at this time sparse is a compressed matrix that still retains the characteristics of the main sound sources and can more efficiently and accurately represent the sound field distribution.

[0059] Furthermore, the method for calculating the precise multi - sound source positions by introducing the variational Bayesian inference framework in step (4) is as follows:

[0060] Previously, the sound intensity matrix I after sparse optimization was obtained sparse . Introduce a latent variable Z to establish a probability model and construct the joint distribution of the sound intensity matrix as:

[0061] p(I sparse ,Z) = p(I sparse |Z)p(Z)

[0062] where p(I sparse ,Z) is a Gaussian distribution and p(Z) is the prior distribution of the latent variable Z.

[0063] Select a variational distribution q(Z) to approximate the posterior p(I sparse ,Z). The variational distribution is generally in a factorized form:

[0064]

[0065] Approximating the log marginal likelihood by optimizing the variational lower bound:

[0066]

[0067] Among them, the first term is the expectation of the joint distribution, and the second term is the entropy of the variational distribution.

[0068] Decompose the joint distribution into conditional distribution and prior:

[0069]

[0070] The optimization objective at this time is to maximize Let the update formula of the variational distribution q(Z i ) be:

[0071]

[0072] Among them, Z -i represents all latent variables except Z i .

[0073] Estimate the sound source position by the maximum a posteriori probability of q(Z) after optimization:

[0074] r i = arg max q(Z i )

[0075] The final set of sparse sound source positions obtained is:

[0076] R sparse = {r i | i ∈ non-zero sparse coefficient positions}.

[0077] Furthermore, the method for visualizing the multi-sound source localization result in step (5) is as follows:

[0078] Map the camera imaging from three-dimensional space to a two-dimensional image plane, and the basic formula is:

[0079] p = K[R|t]P

[0080] Among them, P = [X, Y, Z, 1] T represents the homogeneous coordinates of the sound source in the world coordinate system, and p = [u, v, 1] T represents the homogeneous pixel coordinates of the sound source in the image coordinate system. K is the internal parameter matrix of the camera, and [R|t] is the external parameter matrix of the camera, which can be obtained through camera calibration.

[0081] The internal parameter matrix K of the camera is:

[0082]

[0083] Among them, f x and f y is the focal length of the camera, and c x and c y are the coordinates of the principal point of the image.

[0084] Perform a transformation on the sound source position r i = [X i , Y i , Z i T ∈R sparse as follows:

[0085]

[0086] The obtained pixel coordinates are:

[0087]

[0088] Here, u i and v i are the positions of the sound source in the image.

[0089] According to the corresponding values of the sparse sound intensity matrix I sparse , display the intensity corresponding to each sound source on the image, use a color gradient to represent the magnitude of the sound source intensity, and use OpenCV to superimpose the marked sound sources on the camera image to form a visual output.

[0090] The beneficial effects of the present invention are:

[0091] A multi-source sound source localization and imaging method based on sparse representation and variational Bayesian inference according to the present invention realizes high-precision real-time localization and imaging of multiple sound sources in a complex sound field by combining technologies such as low-resolution preliminary localization, signal intensity gradient optimization, sparse dictionary learning, and variational Bayesian inference, and has high spatial resolution and good anti-interference ability. Description of the Drawings

[0092] Figure 1 is the implementation flowchart of the method of the present invention;

[0093] Figure 2 is the flowchart of gradient search in the method of the present invention;

[0094] Figure 3 is the flowchart of sparse dictionary learning in the method of the present invention;

[0095] Figure 4 is the flowchart of variational Bayesian inference in the method of the present invention. Detailed Embodiments

[0096] ​The present invention will be further illustrated below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the following specific embodiments are only used to illustrate the present invention and not to limit the scope of the present invention.

[0097] As Figure 1 shown, a multi-source localization and imaging method based on sparse representation and variational Bayesian inference according to the present invention includes the following steps:

[0098] (1) Use a conventional beamforming algorithm to generate an initial sound intensity matrix at low resolution and estimate the approximate positions of the sound sources;

[0099] The method for generating the initial sound intensity matrix by the conventional beamforming algorithm is as follows:

[0100] Assume that the sound field signal at a certain point r in space is determined by the sound source signal s(t) and the response of the array sensors. The received signal model is:

[0101]

[0102] where, x(t) ∈ R M is the array received signal, M is the number of array sensors, s k (t) is the signal of the k-th sound source, a(r k ) is the weighting vector corresponding to the position r k of the k-th sound source, and n(t) is the noise.

[0103] Transform the signal to the frequency domain to obtain:

[0104]

[0105] where, S k (ω) represents the frequency response of the k-th sound source, and N(ω) represents the frequency response of the noise term.

[0106] Adopt the delay-and-sum algorithm to calculate, and the beamforming output can be obtained as:

[0107] I(r) = w H (r)X(ω)

[0108] where, is the weighting vector.

[0109] Accumulate the beamforming output energy in the frequency domain and the spatial grid to obtain the initial sound intensity matrix:

[0110]

[0111] where, r g is the grid point of the search space, is the frequency range of interest, and f is the sound source frequency at the current search position.

[0112] In the initial stage, the resolution of the grid point division is relatively low, and at this time, I init is the initial sound intensity matrix, which provides the approximate position of the sound source for subsequent refined search.

[0113] (2) Calculate the signal intensity gradient according to the approximate position of the current sound source, and update the search position along the gradient direction to obtain the accurate sound source position and the high-resolution sound intensity matrix;

[0114] As Figure 2 shown, the method for searching for the accurate position of the sound source along the gradient direction and obtaining the high-resolution sound intensity matrix is as follows:

[0115] Since the initial sound intensity matrix is a function of the spatial position r = (x, y, z), its gradient is the partial derivative of the sound intensity I(r) with respect to the position r, and is defined as:

[0116]

[0117] Using the finite difference method, the gradient in the x direction can be approximated as:

[0118]

[0119] where, e x is the unit vector in the x direction. Similarly, the approximate gradients in the y and z directions can be obtained.

[0120] Use the gradient direction to update the search position using the gradient descent method:

[0121]

[0122] where, η is the step size parameter, which can be set between 0.01 and 0.1 to control the update amplitude.

[0123] Set the iteration to stop when the gradient norm approaches zero:

[0124]

[0125] ∈ is the set minimum threshold value.

[0126] At this time, the high-resolution sound intensity matrix I HR (r) and the accurate sound source position r * can be obtained.

[0127] (3) Use the sparse dictionary learning algorithm to process the high-resolution sound intensity matrix obtained above, and perform alternating optimization of sparse coding and dictionary update on this matrix to reduce signal redundant information and enhance signal sparsity;

[0128] As Figure 3As shown in the figure, the method for enhancing the sparsity of the source signal in the sound intensity matrix using sparse dictionary learning is as follows:

[0129] The high-resolution sound intensity matrix I obtained in the previous steps HR has a sparse representation as:

[0130] I HR = DS + E

[0131] where D is the dictionary matrix, S is the sparse coefficient matrix, and E is the noise or error term.

[0132] Set the sparse optimization objective function as:

[0133]

[0134] where is the Frobenius norm, ∥S∥ 1 is the sparse regularization term, and λ is the sparse regularization parameter.

[0135] Perform the sparse coding process. First, fix the dictionary D and solve for the sparse coefficient S. The sparse coding objective is:

[0136]

[0137] Use convex optimization to solve for the sparse coefficient:

[0138]

[0139] Subsequently, fix the sparse coefficient S and update the dictionary D. The dictionary update objective is:

[0140]

[0141] Adopt the global update method to solve for the dictionary:

[0142] D = I HR S T (SS T ) -1

[0143] Subsequently, repeat the sparse coding and dictionary update steps to achieve alternating optimization of the dictionary and the sparse coefficient until the iteration stop condition is met:

[0144]

[0145] ∈ is the set minimum threshold value.

[0146] The matrix after sparse optimization is:

[0147] I sparse = DS

[0148] The sparse matrix I at this time sparse is a compressed matrix that still retains the characteristics of the main sound sources and can more efficiently and accurately represent the sound field distribution.

[0149] (4) Construct a variational Bayesian inference model based on the sparse coefficients, select the variational distribution, optimize the variational lower bound, and then perform the posterior estimation of the sound sources to obtain the precise positioning results under multiple sound sources;

[0150] As Figure 4 shown, the method for calculating the precise positions of multiple sound sources using the variational Bayesian inference framework is as follows:

[0151] Previously, the sound intensity matrix I after sparse optimization was obtained sparse . Introduce the latent variable Z to establish a probability model, and construct the joint distribution of the sound intensity matrix as:

[0152] p(I sparse ,Z) = p(I sparse |Z)p(Z)

[0153] where p(I sparse ,Z) is a Gaussian distribution, and p(Z) is the prior distribution of the latent variable Z.

[0154] Select the variational distribution q(Z) to approximate the posterior p(I sparse ,Z). The variational distribution is generally in a factorized form:

[0155]

[0156] Approximate the log marginal likelihood by optimizing the variational lower bound:

[0157]

[0158] where the first term is the expectation of the joint distribution, and the second term is the entropy of the variational distribution.

[0159] Decompose the joint distribution into the conditional distribution and the prior:

[0160]

[0161] The optimization objective at this time is to maximize Let the update formula of the variational distribution q(Z i ) be:

[0162]

[0163] where Z -i represents all the latent variables except Z i .

[0164] Estimate the sound source positions through the maximum a posteriori probability of q(Z) after optimization:

[0165] r i = arg maxq(Z i )

[0166] The final set of sparse sound source positions obtained is as follows:

[0167] R sparse = {r i | i ∈ non - zero sparse coefficient positions}.

[0168] (5) Fuse the positioning result with the camera image to obtain the visual position result of the sound source.

[0169] The visualization method for multi - sound - source positioning results is as follows:

[0170] Map the camera imaging from three - dimensional space to a two - dimensional image plane. The basic formula is:

[0171] p = K[R|t]P

[0172] where P = [X, Y, Z, 1] T represents the homogeneous coordinates of the sound source in the world coordinate system, p = [u, v, 1] T represents the homogeneous pixel coordinates of the sound source in the image coordinate system, K is the internal parameter matrix of the camera, and [R|t] is the external parameter matrix of the camera, which can be obtained through camera calibration.

[0173] The internal parameter matrix K of the camera is:

[0174]

[0175] where f x and f y are the focal lengths of the camera, and c x and c y are the coordinates of the principal point of the image.

[0176] Perform a transformation on the previously obtained sound source position r i = [X i , Y i , Z i T ∈ R sparse as follows:

[0177]

[0178] The pixel coordinates obtained are:

[0179]

[0180] Here, u i and v i are the positions of the sound source in the image.​

[0181] According to the corresponding values of the sparse sound intensity matrix I sparse display the intensity corresponding to each sound source on the image, use a color gradient to represent the magnitude of the sound source intensity, and use OpenCV to superimpose the marked sound sources on the camera image to form a visual output.

[0182] A multi-source localization and imaging method based on sparse representation and variational Bayesian inference according to the present invention realizes high-precision real-time localization and imaging of multiple sound sources in a complex sound field by combining technologies such as low-resolution preliminary localization, signal intensity gradient optimization, sparse dictionary learning, and variational Bayesian inference, and has high spatial resolution and good anti-interference ability.

[0183] It should be noted that the above content only illustrates the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. For those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements all fall within the protection scope of the claims of the present invention.

Claims

1. A multi-sound source localization and imaging method based on sparse representation and variational Bayesian inference, characterized in that: The following steps are involved: (1) Generate an initial sound intensity matrix at low resolution using a conventional beamforming algorithm to estimate the approximate location of the sound source; (2) Calculate the signal intensity gradient according to the approximate location of the current sound source, and update the search position along the gradient direction to obtain the accurate sound source location and high-resolution sound intensity matrix; (3) using a sparse dictionary learning algorithm to process the high-resolution sound intensity matrix obtained in step (2), and performing alternating optimization of sparse coding and dictionary updating on the matrix to reduce signal redundant information and enhance signal sparsity; (4) A variational Bayesian inference model is constructed based on the sparse coefficients, the variational distribution is selected, and the sound source posterior estimation is performed after optimizing the variational lower bound to obtain accurate positioning results under multiple sound sources; (5) The positioning result is integrated with the camera image to obtain the visual position result of the sound source.

2. The method for multiple sound source localization and imaging based on sparse representation and variational Bayesian inference according to claim 1, characterized in that: The method for generating the initial sound intensity matrix using the conventional beamforming algorithm in step (1) is as follows: Assuming that the sound field signal at a certain point r in space is determined by the sound source signal s(t) and the response of the array sensor, the received signal model is: Where x(t)∈R M is the array receiving signal, M is the number of array sensors, s k (t) is the signal of the kth sound source, a(r k ) is the position r of the kth sound source k The corresponding weighted vector, n(t) is the noise; Transforming the signal to the frequency domain yields: Among them, S k (ω) represents the frequency response of the kth sound source, and N(ω) represents the frequency response of the noise term; Using the delay-sum algorithm, the beamforming output is: I(r)=w H (r)X(ω) in, is the weight vector; The beamforming output energy is accumulated in the frequency domain and spatial grid to obtain the initial sound intensity matrix: Among them, r g is the grid point of the search space, is the frequency range of interest, and f is the sound source frequency at the current search position; In the initial stage, the resolution of the grid points is low. init This is the initial sound intensity matrix, which provides the approximate location of the sound source for subsequent refined search.

3. The method for multiple sound source localization and imaging based on sparse representation and variational Bayesian inference according to claim 1, characterized in that: In step (2), the method for searching the exact location of the sound source along the gradient direction and obtaining a high-resolution sound intensity matrix is ​​as follows: Since the initial sound intensity matrix is ​​a function of the spatial position r = (x, y, z), its gradient is the partial derivative of the sound intensity I(r) with respect to the position r, defined as: The x-direction gradient is approximated using the finite difference method as: Among them, e x is the unit vector in the x direction, and the approximate gradients in the y and z directions are similarly obtained; Update the search position using gradient descent using the gradient direction: Where η is the step size parameter, which is set to 0.01 to 0.1 and is used to control the update amplitude; Set iterations to stop when the gradient norm approaches zero: ∈ is the set minimum threshold; At this time, the high-resolution sound intensity matrix I is obtained HR (r) and the exact sound source location r * .

4. The method for multiple sound source localization and imaging based on sparse representation and variational Bayesian inference according to claim 1, characterized in that: The method for using sparse dictionary learning to enhance the sparsity of the source signal in the sound intensity matrix in step (3) is as follows: High Resolution Sound Intensity Matrix I HR The sparse representation of is: I HR =DS+E Where D is the dictionary matrix, S is the sparse coefficient matrix, and E is the noise or error term; Set the sparse optimization objective function to: in, is the Frobenius norm, ∥S∥1 is the sparse regularization term, and λ is the sparse regularization parameter; In the sparse coding process, the dictionary D is fixed first, and the sparse coefficient S is solved. The sparse coding target is: Solve for sparse coefficients using convex optimization: Then fix the sparse coefficient S, update the dictionary D, and the dictionary update target is: Solve the dictionary using a global update method: D=I HR S T (SS T ) -1 Then repeat the sparse coding and dictionary update steps to alternately optimize the dictionary and sparse coefficients until the iteration stop condition is met: ∈ is the set minimum threshold; The sparse optimized matrix is: I sparse =DS At this time, the sparse matrix I sparse It is a matrix that is compressed but still retains the characteristics of the main sound sources, and can more efficiently and accurately characterize the sound field distribution.

5. The method for multiple sound source localization and imaging based on sparse representation and variational Bayesian inference according to claim 1, characterized in that: The variational Bayesian inference model introduced in step (4) is used to calculate the precise location of multiple sound sources as follows: The sound intensity matrix I after sparse optimization sparse , introduce the latent variable Z to establish a probability model, and construct the joint distribution of the sound intensity matrix as follows: p(I sparse ,Z)=p(I sparse |Z)p(Z) Among them, p(I sparse ,Z) is a Gaussian distribution, p(Z) is the prior distribution of the latent variable Z; Choose a variational distribution q(Z) to approximate the posterior p(I sparse ,Z), the variational distribution is generally in factorized form: Approximate the log-marginal likelihood by optimizing a variational lower bound: Among them, the first term is the expectation of the joint distribution, and the second term is the entropy of the variational distribution; Decompose the joint distribution into the conditional distribution and the prior: The optimization goal at this time is to maximize Let the variational distribution q(Z i ) is updated as: Among them, Z -i Indicates that except Z i All potential variables except The sound source position is estimated by optimizing the maximum a posteriori probability of q(Z): r i =argmaxq(Z i ) The final sparse sound source position set is: R sparse = {r i |i∈non-zero sparse coefficient position}.

6. The method for multiple sound source localization and imaging based on sparse representation and variational Bayesian inference according to claim 1, characterized in that: The method for visualizing the sound source localization result in step (5) is as follows: The basic formula for mapping the camera imaging from three-dimensional space to a two-dimensional image plane is: p=K[R|t]P in, represents the homogeneous coordinates of the source in the world coordinate system, represents the homogeneous pixel coordinates of the sound source in the image coordinate system, K is the intrinsic parameter matrix of the camera, and [R|t] is the extrinsic parameter matrix of the camera, which is obtained through camera calibration; The camera intrinsic parameter matrix K is: Among them, f x With f y is the focal length of the camera, c x With c y is the coordinate of the principal point of the image; The obtained sound source position To transform: The pixel coordinates are: u i With v i is the position of the sound source in the image; According to the sparse sound intensity matrix I sparse The corresponding value of the sound source is displayed on the image, and the color gradient is used to represent the intensity of the sound source. OpenCV is used to overlay the marked sound source on the camera image to form a visual output.

Citation Information

Patent Citations

  • Sound source localization method achieved in strong reverberation environment on basis of parameterization Bayes dictionary learning

    CN106842112A

  • Silk image denoising method and system based on KSVD algorithm

    CN117876245A

Cited By

  • Target imaging method and system fused with sound field gradient information

    CN120314961A

  • Gearbox fault diagnosis method based on lightweight variational Bayesian learning

    CN120524160A