A sound source positioning method based on laplacian norm regularization
By introducing Laplace norm regularization and iterative shrinkage threshold algorithm into sound source localization technology, the problems of insufficient sparsity and adaptive selection of regularization parameters are solved, and high-precision and high-speed sound source localization effect is achieved.
Patent Information
- Application Number
- CN202510094134.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-01-21
AI Technical Summary
In the existing compressed beamforming sound source localization technology, insufficient sparsity and uneven penalty make the positioning results susceptible to interference noise, and the adaptive selection of regularization parameters is difficult, which affects the positioning accuracy.
Laplace norm regularization and iterative shrinkage threshold algorithm are adopted to construct a Laplace norm sparse constraint model by dividing the sound source plane into grid points. The regularization parameter is adaptively selected in the iterative process to improve sparsity and robustness.
It achieves high-precision and high-speed sound source positioning, effectively suppresses pseudo-sound source interference, and improves positioning accuracy and spatial resolution.
Smart Images

Figure CN120009828B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of sound source localization and relates to a sound source localization method based on Laplace norm regularization. Background Art
[0002] With the rapid development of microphone array signal processing technology, sound source localization has been widely used in fields such as mechanical engineering, aerospace, and marine acoustics. High-precision, high-spatial-resolution, and clear, intuitive sound source localization is crucial for mechanical noise control, target identification, and fault diagnosis. Acoustic beamforming, a key branch of sound source localization technology, uses signal processing techniques such as beamforming to capture sound field data from a microphone array to pinpoint or locate the sound source.
[0003] The beamforming technology under the compressed sensing framework is called compressed beamforming. It can accurately locate the position of the sound source even when the number of array elements is less than the number of sound sources. It has been favored by researchers in recent years. The sound source localization technology based on compressed beamforming generally uses the l1 norm as a sparse constraint to obtain a unique solution. However, the l1 norm has problems such as insufficient sparsity and uneven penalty, which makes the positioning result susceptible to interference noise, and the solved sound source intensity has the problem of underestimation. These problems reduce the positioning performance of the algorithm. To alleviate these problems, there are currently many compressed beamforming models based on non-convex norms for sound source localization, such as l1 norm. p norm and generalized min-max concave penalty, etc. These methods have been shown to be able to induce stronger sparsity and impose more uniform penalties than l1 norm or weighted l1 norm. p Compressed beamforming models using the concave norm or generalized minimum-maximum penalty still face the challenge of adaptively selecting the regularization parameter. The regularization parameter is used to balance the fidelity term, which reflects the approximation of the solution to the real data, with the regularization term, which ensures the sparsity of the solution. The appropriate value of the regularization parameter is crucial to the quality of the optimal solution.
[0004] Therefore, in sound source localization technology based on compressed beamforming, how to further promote the sparsity of the solution, ensure the consistency of the penalty, and adaptively select the regularization parameter is one of the key issues to improve positioning accuracy and reduce the impact of interference noise. It is necessary to adopt regularization techniques that can induce stronger sparsity and design a regularization parameter adaptive selection strategy to overcome the shortcomings of existing technologies and solve or alleviate these problems. Summary of the Invention
[0005] In view of this, the object of the present invention is to provide a sound source localization method based on Laplace norm regularization, which can induce stronger sparsity, alleviate the problem that the sound source localization results are interfered by noise and have many pseudo sound sources, and can achieve fast and accurate sound source localization.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A sound source localization method based on Laplace norm regularization is proposed. The method first divides the plane where the sound source is located into a series of grid points, and assumes that the coordinates of the grid points are the potential locations of the sound source. Then, the Laplace norm is introduced as a regularization term to construct a compressed beamforming sound source localization model with Laplace norm sparsity constraint. Finally, based on the iterative shrinkage threshold algorithm (ISTA) framework, the non-convex Laplace norm sparsity constraint minimization problem is solved. In the solution process, an adaptive regularization parameter selection strategy is introduced to improve the quality of the optimal solution and enhance the robustness of the algorithm.
[0008] The method specifically comprises the following steps:
[0009] S1: Divide the plane where the sound source is located into a series of grid points, assuming that the coordinates of the grid points are the potential locations of the sound source;
[0010] S2: Based on the Green function, the transfer matrix between all grid points and the microphone array is obtained, and a minimization model with Laplace norm sparsity constraint is constructed;
[0011] S3: Iteratively solve the minimization model of non-convex Laplace norm sparsity constraints based on the iterative shrinkage threshold algorithm framework;
[0012] S4: During the solution process, the regularization parameter is adaptively determined based on the sound field data to obtain the strength of the potential sound source corresponding to all grid points. The greater the strength of the grid point, the greater the probability of it being the true location of the sound source;
[0013] S5: Based on the prior information that the sparsity of the solution is S, the grid coordinates corresponding to the first S largest sound intensities are taken as the estimated results of the sound source position.
[0014] Furthermore, step S1 specifically includes the following steps:
[0015] S11: There are S sound sources in the free field space, and the coordinates of the sound sources are r s =(x s ,y s ,z s ); The microphone array consists of M array elements, and the coordinates of the array elements are r m =(x m ,y m ,z m) ; Divide the plane where the sound source is located into N grid points according to the grid spacing d, and the coordinates of each grid point are r g =(x g ,y g ,z g ); The xoy plane of the Cartesian coordinate system overlaps with the microphone array, and the origin o is the center of the array; the microphone array is placed parallel to the sound source at a distance of z0 meters, and the distance between the mth array element and the sth sound source is defined as d ms :
[0016]
[0017] S12: The complex sound pressure measured by the mth array element is Represents a complex field:
[0018]
[0019] in, is the sound intensity of the sth sound source, The error of the data collected by the channel corresponding to the mth array element caused by the mathematical model and environmental noise; is the Green's function between the mth array element and the sth sound source:
[0020]
[0021] Where j is the imaginary unit, κ = 2πf / c is the wave number, f is the frequency, c is the speed of sound, and exp(·) represents the exponential function;
[0022] S13: Assuming that the coordinates of the grid points are the potential locations of the sound sources, the complex sound pressure measured by the mth array element can be re-expressed as:
[0023]
[0024] in, is the sound intensity of the potential sound source corresponding to the grid point, is the Green's function between the mth array element and the sound source corresponding to the gth grid.
[0025] Furthermore, step S2 specifically includes the following steps:
[0026] S21: Considering N grid points and M array elements, the sound pressure data P measured by the microphone array is expressed in matrix form:
[0027] P=Aq+E
[0028] in, is the sound pressure vector measured by the microphone array, is the sound intensity vector of the potential sound source corresponding to N grid points, is the error vector; [·] T represents transpose; is the transfer matrix between M array elements and N grid points:
[0029]
[0030] in, is the transfer vector between the M array elements and the i-th grid;
[0031] S22: Assume that S sound sources are sparsely distributed in the free field. Since S<<N, q is a sparse vector. Since M<N, this results in an underdetermined problem. The Laplace norm is introduced as a sparse constraint to obtain a unique solution. The sound source localization problem can be modeled as a minimization problem under the Laplace norm sparse constraint:
[0032]
[0033] Where λ>0 is the regularization parameter, ||·||2 represents the l2 norm; |||·||| γ represents the Laplace norm, which is defined as:
[0034]
[0035] Here, γ>0 is the order of the Laplace norm; when γ approaches zero, it can be regarded as an approximation of the l0 norm.
[0036] Furthermore, step S3 specifically includes the following steps:
[0037] S31: Introducing an additional l1 norm term, the minimization model of the Laplace norm sparse constraint is constructed into the following differential form:
[0038]
[0039] Where F(q) is the overall objective function, H(q) is the sub-objective function 1, G(q) is the sub-objective function 2, and α>0 is a constant;
[0040] S32: H(q) can be directly solved using the iterative shrinkage threshold algorithm, which is equivalent to solving the following expression:
[0041]
[0042] Where V(q) is the gradient of the squared term in H(q):
[0043] V(q)=q+A H (P-Aq)
[0044] in,[·] Hrepresents the conjugate transpose; then the iterative expression for solving H(q) is:
[0045] q k+1 =S λ (q k +A H (P-Aq k ))
[0046] Where k is the number of iterations; S λ (q)=[s λ (q1),s λ (q2),…,s λ (q N )] T For threshold operation:
[0047]
[0048] Among them, |·| is the modulus length of the complex number; sign(q i ) is a symbolic function:
[0049]
[0050] in, represents the field of real numbers;
[0051] S33: The solution for G(q) can be linearly approximated as:
[0052] G(q)≈G(q k )+ <G'(q k ),(qq k )>
[0053] Where <·> represents the inner product, G'(q)=[G'(q1),…,G'(q N )] T is the gradient of G(q):
[0054]
[0055] The iterative expression for solving the minimization problem of the Laplace norm sparsity constraint in the framework of the iterative shrinkage threshold algorithm is:
[0056]
[0057] The expression of D(q) is:
[0058]
[0059] Repeat the above iterative solution process and set the number of iterations until the objective function reaches the iteration stopping condition.
[0060] Further, step S4 specifically includes the following steps:
[0061] S41: When the regularization term is separable, solving the expression for H(q) is equivalent to the following univariate problem:
[0062]
[0063] Among them, v(q) is the single variable corresponding to the quadratic gradient V(q) in the objective function H(q); q k+1 The iterative expression is h v '(q)=0; however, h v '(q) is a transcendental equation. Considering that the physical meaning of q is sound intensity, the following expression can be obtained:
[0064]
[0065] Obviously, v(q) = αλ / 2γ can be used as a threshold for iteration. Therefore, the appropriate regularization parameter can be determined based on this threshold. Rearrange the elements in V(q) by size:
[0066] |[V(q k )]1|≥|[V(q k )]2|≥…≥|[V(q k )] N |
[0067] Among them, |[V(q k )] i | is the i-th largest element in V(q); after sorting, the first S elements in V(q) are larger than the threshold αλ / 2γ, so for the S-th element and the S+1-th element, the following inequality holds:
[0068]
[0069] S42: Without loss of generality, we no longer strictly distinguish between the parameters λ and αλ; at the kth iteration, the strategy for selecting the regularization parameter is:
[0070] λ k =2γ((1-μ)|[V(q k )] S+1 |+μ|[V(q k )] S |)
[0071] Among them, the value range of μ is μ∈[0,1); λ k The value of increases with the increase of μ, and λ k The larger the value, the stronger the sparsity; therefore, we can choose μ = 0, or more generally: k=min{λ k-1 ,2γ|[V(q k )] S+1 |}.
[0072] Further, step S5 specifically includes the following steps:
[0073] S51: Solve the non-convex Laplace norm constraint minimization problem through the above iterative shrinkage threshold algorithm. The solution is the sound intensity q of the potential sound source corresponding to each grid. Assume that the sparsity of the optimal solution is S, that is, the number of sound sources is regarded as prior information, and the first S largest elements in q are regarded as the estimated sound intensity of the sound source, which is recorded as q on ;
[0074] S52: By searching all grids, q on The corresponding grid coordinates are regarded as the estimated value of the sound source location, denoted as r on ; Therefore, the sound source information estimation result solved by a sound source localization method based on Laplace norm regularization in the framework of iterative shrinkage threshold algorithm is:
[0075]
[0076] in, It means taking the values corresponding to the first S largest elements.
[0077] The beneficial effects of the present invention are as follows: by introducing the Laplace norm as a regularization term and adopting an adaptive regularization parameter selection strategy, the present invention achieves greater sparsity in the sound source localization model and improves the quality of the optimal solution. This enables high-accuracy, high-spatial-resolution localization and effectively suppresses pseudo-sound sources caused by interfering noise. Furthermore, an iterative solution framework based on an iterative shrinkage threshold algorithm is employed to quickly and efficiently solve the non-convex Laplace norm sparsity-constrained minimization problem.
[0078] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:
[0080] Figure 1 This is a flow chart of the sound source localization method based on Laplace norm regularization according to the present invention;
[0081] Figure 2 The geometric figures of the regularization functions with different norms, where (a) is the geometric figure comparison of the four norm regularization functions, namely: l2 norm, l1 norm, l p Norm and Laplace norm, (b) is the geometric graph of Laplace norm under different γ;
[0082] Figure 3 56-channel multi-arm spiral microphone array used for simulation;
[0083] Figure 4 This is a comparison diagram of the sparsity of the method of the present invention and two traditional methods when the signal-to-noise ratio (SNR) is 20 dB and the frequency f is 20 kHz;
[0084] Figure 5 The figure shows the comparison of the simulation positioning effect between the method of the present invention and two traditional methods when the signal-to-noise ratio (SNR) is 20dB and the frequency f is 20kHz.
[0085] Figure 6 The figure shows the comparison of the simulation positioning effect between the method of the present invention and two traditional methods when the signal-to-noise ratio (SNR) is 0 dB and the frequency f is 20 kHz.
[0086] Figure 7 The Monte Carlo simulation results of the method of the present invention and two traditional methods are shown in FIG. 1 when the signal-to-noise ratio (SNR) range is 0 dB to 20 dB and the frequency f is 20 kHz.
[0087] Figure 8 The Monte Carlo simulation results of the method of the present invention and two traditional methods are shown when the frequency f ranges from 10kHz to 30kHz and the signal-to-noise ratio (SNR) is 20dB. DETAILED DESCRIPTION
[0088] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0089] See also Figures 1 to 8 The present invention provides a sound source localization method based on Laplace norm regularization, which includes the following steps:
[0090] Step S1: Divide the plane where the sound source is located into a series of grid points. Assume that the coordinates of the grid points are the potential locations of the sound source. Specifically, the following steps are included:
[0091] S11: There are S sound sources in the free field space, and the coordinates of the sound sources are r s =(x s ,y s ,z s ); The microphone array consists of M array elements, and the coordinates of the array elements are r m =(x m ,y m ,z m ) ; Divide the plane where the sound source is located into N grid points according to the grid spacing d, and the coordinates of each grid point are r g =(x g ,y g ,z g ); The xoy plane of the Cartesian coordinate system overlaps with the microphone array, and the origin o is the center of the array; the microphone array is placed parallel to the sound source at a distance of z0 meters, and the distance between the mth array element and the sth sound source is defined as d ms :
[0092]
[0093] S12: The complex sound pressure measured by the mth array element is Represents a complex field:
[0094]
[0095] in, is the sound intensity of the sth sound source, is the error caused by mathematical model, environmental noise, etc. in the data collected by the channel corresponding to the mth array element, is the Green's function between the mth array element and the sth sound source:
[0096]
[0097] Where j is the imaginary unit, κ = 2πf / c is the wave number, f is the frequency, c is the speed of sound, and exp(·) represents the exponential function;
[0098] S13: Assuming that the coordinates of the grid points are the potential locations of the sound sources, the complex sound pressure measured by the mth array element can be re-expressed as:
[0099]
[0100] in, is the sound intensity of the potential sound source corresponding to the grid point, is the Green's function between the mth array element and the sound source corresponding to the gth grid.
[0101] Step S2: Based on the Green's function, the transfer matrix between all grid points and the microphone array is obtained, and a minimization model of the Laplace norm sparse constraint is constructed, which specifically includes the following steps:
[0102] S21: Considering N grid points and M array elements, the sound pressure data P measured by the microphone array is expressed in matrix form:
[0103] P=Aq+E
[0104] in, is the sound pressure vector measured by the microphone array, is the sound intensity vector of the potential sound source corresponding to N grid points, is the error vector; [·] T represents transpose; is the transfer matrix between M array elements and N grid points:
[0105]
[0106] in, is the transfer vector between the M array elements and the i-th grid;
[0107] S22: Assume that S sound sources are sparsely distributed in the free field. Since S<<N, q is a sparse vector. Since M<N, this results in an underdetermined problem. The Laplace norm is introduced as a sparse constraint to obtain a unique solution. The sound source localization problem can be modeled as a minimization problem under the Laplace norm sparse constraint:
[0108]
[0109] Where λ>0 is the regularization parameter, ||·||2 represents the l2 norm, and |||·||| γ represents the Laplace norm, which is defined as:
[0110]
[0111] Here, γ>0 is the order of the Laplace norm; when γ approaches zero, it can be regarded as an approximation of the l0 norm.
[0112] Figure 2 The geometric figures of the regularization functions with different norms, where (a) is the geometric figure comparison of the four norm regularization functions, namely: l2 norm, l1 norm, l pnorm and Laplace norm, (b) is the geometric graph of the Laplace norm under different γ; from (a), it can be seen that compared with other norms, the Laplace norm has a stronger sparsity-promoting effect and is closer to the l0 norm; from (b), it can be seen that as γ decreases, the sparsity-promoting effect of the Laplace norm increases, and the stronger the sparsity-promoting effect, the better the sparsity of the solution.
[0113] Step S3: Iteratively solving the minimization model of the non-convex Laplace norm constraint based on the iterative shrinkage threshold algorithm framework, and adaptively determining the regularization parameter based on the sound field data during the solution process, specifically including the following steps:
[0114] S31: Introducing an additional l1 norm term, the minimization model of the Laplace norm sparse constraint is constructed into the following differential form:
[0115]
[0116] Where, F(q) is the overall objective function, H(q) is the sub-objective function 1, G(q) is the sub-objective function 2, and α>0 is a constant;
[0117] S32: H(q) can be directly solved using the iterative shrinkage threshold algorithm, which is equivalent to solving the following expression:
[0118]
[0119] Where V(q) is the gradient of the squared term in H(q):
[0120] V(q)=q+A H (P-Aq)
[0121] in,[·] H represents the conjugate transpose; then the iterative expression for solving H(q) is:
[0122] q k+1 =S λ (q k +A H (P-Aq k ))
[0123] Where k is the number of iterations; S λ (q)=[s λ (q1),s λ (q2),…,s λ (q N )] T For threshold operation:
[0124]
[0125] Among them, |·| is the modulus length of the complex number; sign(q i ) is a symbolic function:
[0126]
[0127] in, represents the field of real numbers;
[0128] S33: The solution for G(q) can be linearly approximated as:
[0129] G(q)≈G(q k )+ <G'(q k ),(qq k )>
[0130] Where <·> represents the inner product, G'(q)=[G'(q1),…,G'(q N )] T is the gradient of G(q):
[0131]
[0132] The iterative expression for solving the minimization problem of the Laplace norm sparsity constraint in the framework of the iterative shrinkage threshold algorithm is:
[0133]
[0134] The expression of D(q) is:
[0135]
[0136] Repeat the above iterative solution process and set the number of iterations until the objective function reaches the iteration stopping condition.
[0137] Step S4: During the solution process, a regularization parameter is adaptively determined based on the sound field data to obtain the strength of the potential sound source corresponding to all grid points. The greater the strength of the grid point, the greater the probability of it being the true location of the sound source. Specifically, the following steps are included:
[0138] S41: When the regularization term is separable, solving the expression for H(q) is equivalent to the following univariate problem:
[0139]
[0140] Among them, v(q) is the single variable corresponding to the quadratic gradient V(q) in the objective function H(q); q k+1 The iterative expression is h v '(q)=0; however, h v '(q) is a transcendental equation. Considering that the physical meaning of q is sound intensity, the following expression can be obtained:
[0141]
[0142] Obviously, v(q) = αλ / 2γ can be used as a threshold for iteration. Therefore, based on this threshold, a suitable regularization parameter can be determined. Rearrange the elements in V(q) by size:
[0143] |[V(q k )]1|≥|[V(q k )]2|≥…≥|[V(q k )] N |
[0144] Among them, |[V(q k )] i | is the i-th largest element in V(q); after sorting, the first S elements in V(q) are larger than the threshold αλ / 2γ, so for the S-th element and the S+1-th element, the following inequality holds:
[0145]
[0146] S42: Without loss of generality, we no longer strictly distinguish between the parameters λ and αλ; at the kth iteration, the strategy for selecting the regularization parameter is:
[0147] λ k =2γ((1-μ)|[V(q k )] S+1 |+μ|[V(q k )] S |)
[0148] Among them, the value range of μ is μ∈[0,1); λ k The value of increases with the increase of μ, and λ k The larger it is, the stronger the sparsity is; therefore, we can choose μ = 0, or more generally:
[0149] λ k =min{λ k-1 ,2γ|[V(q k )] S+1 |}
[0150] Step S5: Based on the prior information that the sparsity of the solution is S, the grid coordinates corresponding to the first S largest sound intensities are taken as the estimated result of the sound source position, which specifically includes the following steps:
[0151] S51: Solve the non-convex Laplace norm constraint minimization problem through the above iterative algorithm. The solution is the sound intensity q of the potential sound source corresponding to each grid. Assume that the sparsity of the optimal solution is S, that is, the number of sound sources is regarded as prior information, and the first S largest elements in q are regarded as the estimated sound intensity of the sound source, which is recorded as q on ;
[0152] S52: By searching all grids, q on The corresponding grid coordinates are regarded as the estimated value of the sound source location, denoted as r on ; Therefore, for the sound source localization method based on Laplace norm regularization described in claim 1, the sound source information estimation result solved under the framework of iterative shrinkage threshold algorithm is:
[0153]
[0154] in, It means taking the values corresponding to the first S largest elements.
[0155] Verification experiment:
[0156] The following simulation experiments demonstrate the beneficial effects of the method of the present invention:
[0157] Simulation test: A multi-arm spiral microphone array with 56 elements M is used for simulation test. Its diameter is about 0.20 meters. The positions of the 56 elements are distributed as follows: Figure 3 As shown in the figure, the sampling rate of the microphone array is 192kHz. The target sound source plane is parallel to the array plane and is located in a target area of 1.2m × 1.2m with a grid spacing of z0 = 1.5m in front of the array plane. The target area is divided into 13 × 13 grid coordinates with a grid spacing of d = 0.1m. In the simulation, four sound sources with a frequency of 20kHz and a sound pressure level of 93.98dB (1Pa) are set in the target area. The sound source coordinates are set to S1 (-0.16m, -0.51m), S2 (-0.11m, -0.48m), S3 (0.33m, 0.39m) and S4 (0.6m, -0.6m). The spatial distance between sound sources S1 and S2 is set to a very small value (0.27m) to demonstrate the spatial resolution performance of the algorithm. Sound source S4 falls on the grid point, and the other sound sources are all offset from the grid points. The sound source is located using the Laplace regularization-based sound source localization method (LapAIST), the Tychonov regularization-based localization method (Tik), and the fast iterative threshold shrinkage algorithm-based localization method (FISTA) of the present invention.
[0158] The simulation experiment was conducted with the signal-to-noise ratio (SNR) set to 20dB and the frequency f set to 20kHz. The sparsity comparison of the three methods is shown in the figure below. Figure 4As shown in the figure; the horizontal axis represents the grid index (one-to-one correspondence with the grid coordinates), and the vertical axis represents the modulus |q| of the sound source intensity estimated by the positioning method; the horizontal axis of the four black dotted lines in the figure is the grid index corresponding to the grid coordinate closest to the real sound source, and the theoretical modulus of the real sound source intensity is 1Pa. Figure 4 It can be seen that the localization method based on Tikhonov regularization (Tik) has the weakest sparsity, there are a large number of interference sound sources of comparable intensity near the real sound source, and the estimated sound intensity is quite different from the theoretical value; the localization method based on the fast iterative threshold shrinkage algorithm (FISTA) can effectively promote the sparsity of the solution and better estimate the sound intensity, but there is still interference from two pseudo sound sources; in contrast, the sound source localization method based on Laplace regularization (LapAIST) of the present invention can not only accurately estimate the position and intensity of the sound source, but also effectively suppress the pseudo sound sources caused by interference noise.
[0159] The positioning experiment was conducted in a simulation environment with a signal-to-noise ratio (SNR) of 20dB and a frequency of 20kHz. The positioning effects of the three methods are as follows: Figure 5 As shown in the figure, "*" and "o" represent the estimated sound source position and the actual sound source position respectively, and the dynamic display range is set to a 10dB difference from the maximum sound intensity of the sound source. Figure 5 It can be seen that the positioning method based on Tikhonov regularization (Tik) can only basically identify the locations of three sound sources, and there are many pseudo sound source interferences near the sound sources. It is impossible to accurately distinguish the positions of sound sources with adjacent spatial positions, and the spatial resolution is poor. Under high signal-to-noise ratio conditions, the positioning method based on the fast iterative threshold shrinkage algorithm (FISTA) has a sound source positioning effect comparable to the sound source localization method based on Laplace regularization (LapAIST) of the present invention. Both can accurately identify the grid coordinates closest to the real sound source and accurately distinguish two sound sources with adjacent spatial positions. However, the accuracy of sound intensity estimation is inferior to the positioning method of the present invention.
[0160] Furthermore, positioning experiments were conducted in a simulation environment with a signal-to-noise ratio (SNR) of 0 dB and a frequency of 20 kHz to verify the positioning effects of the three methods in a low signal-to-noise ratio environment. The results are shown in the figure. Figure 6 As shown. Figure 6It can be seen from the figure that the positioning method based on Tikhonov regularization (Tik) is affected by many interference noises and cannot accurately identify the sound source. This method completely fails in a low signal-to-noise ratio environment; the positioning method based on the fast iterative threshold shrinkage algorithm (FISTA) is interfered by two pseudo sound sources. Although the intensity of the pseudo sound source is weaker than that of the estimated sound source, it still interferes with the judgment of the operator in actual engineering. Further judgment is required for these two positions, which increases the detection cost; the sound source localization method based on Laplace regularization (LapAIST) of the present invention can still accurately estimate the position and intensity of the four sound sources under low signal-to-noise ratio, and accurately distinguish two sound sources in adjacent spatial positions. The reason is that the sparse promotion effect of the Laplace norm is stronger than the other two methods, which can effectively suppress interference noise.
[0161] Monte Carlo randomized simulations: Monte Carlo randomized trials of three methods at different signal-to-noise ratios and frequencies were simulated to further demonstrate the superiority of the present method. Four sound sources were randomly distributed within a 1.2m x 1.2m target area, divided into 13 x 13 grid coordinates using a grid spacing of d = 0.1m. Three hundred Monte Carlo randomized trials were conducted for each method at different signal-to-noise ratios and frequencies. To evaluate the performance of the positioning method, three evaluation metrics were defined: cumulative correct identification rate (CCIR), root mean square error (RMSEL) of sound source location, and root mean square error (RMSES) of sound intensity. A "correct identification" in the simulations was defined as the error between the estimated sound source location and intensity and the true value being less than a tolerance. In the present invention, the error tolerance for sound source location was set to 10cm, and the error tolerance for intensity was set to 5dB.
[0162] The signal-to-noise ratio (SNR) range is set to 0dB to 20dB, the SNR value is taken at intervals of 1dB, and the frequency f is 20kHz. The Monte Carlo simulation results of the three methods are as follows: Figure 7 As shown in the figure, the frequency f range is set to 10kHz~30kHz, the frequency value is taken at intervals of 1kHz, the signal-to-noise ratio SNR is 20dB, and the Monte Carlo simulation results of the three methods are shown in the figure. Figure 8 shown by Figure 7 、 Figure 8 It can be seen that the positioning performance of the sound source localization method based on Laplace regularization (LapAIST) of the present invention is far superior to the other two positioning methods based on traditional regularization functions; thanks to the excellent sparse promotion performance of the Laplace norm and the adaptive selection strategy of the regularization parameter, the positioning method of the present invention achieves a CCIR higher than 80% in various signal-to-noise ratios and frequency ranges, an RMSEL of around 5 cm, and an RMSES less than 3 dB, indicating that the method of the present invention can stably and accurately locate sound sources in a wide frequency band and low signal-to-noise ratio environment.
[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions, which should all be included in the scope of the claims of the present invention.
Claims
1. A sound source localization method based on Laplace norm regularization, characterized in that: The method specifically comprises the following steps: S1: Divide the plane where the sound source is located into a series of grid points, assuming that the coordinates of the grid points are the potential locations of the sound source; S2: Based on the Green function, the transfer matrix between all grid points and the microphone array is obtained, and a minimization model with Laplace norm sparsity constraint is constructed; S3: Iteratively solve the minimization model of non-convex Laplace norm sparsity constraints based on the iterative shrinkage threshold algorithm framework; S4: During the solution process, the regularization parameter is adaptively determined based on the sound field data to obtain the strength of the potential sound source corresponding to all grid points. The greater the strength of the grid point, the greater the probability of it being the true location of the sound source; S5: Based on the prior information that the sparsity of the solution is S, the grid coordinates corresponding to the first S largest sound intensities are taken as the estimated results of the sound source position.
2. The sound source localization method according to claim 1, wherein: Step S1 specifically includes the following steps: S11: There are S sound sources in the free field space, and the coordinates of the sound sources are r s =(x s ,y s ,z s ); The microphone array consists of M array elements, and the coordinates of the array elements are r m =(x m ,y m ,z m ) ; Divide the plane where the sound source is located into N grid points according to the grid spacing d, and the coordinates of each grid point are r g =(x g ,y g ,z g ); The xoy plane of the Cartesian coordinate system overlaps with the microphone array, and the origin o is the center of the array; the microphone array is placed parallel to the sound source at a distance of z0 meters, and the distance between the mth array element and the sth sound source is defined as d ms : S12: The complex sound pressure measured by the mth array element is Represents a complex field: in, is the sound intensity of the sth sound source, is the error caused by the mathematical model and environmental noise in the data collected by the channel corresponding to the mth array element; is the Green's function between the mth array element and the sth sound source: Where j is the imaginary unit, κ = 2πf / c is the wave number, f is the frequency, c is the speed of sound, and exp(·) represents the exponential function; S13: Assuming that the coordinates of the grid points are the potential locations of the sound sources, the complex sound pressure measured by the mth array element can be re-expressed as: in, is the sound intensity of the potential sound source corresponding to the grid point, is the Green's function between the mth array element and the sound source corresponding to the gth grid.
3. The sound source localization method according to claim 2, wherein: Step S2 specifically includes the following steps: S21: Considering N grid points and M array elements, the sound pressure data P measured by the microphone array is expressed in matrix form: P=Aq+E in, is the sound pressure vector measured by the microphone array, is the sound intensity vector of the potential sound source corresponding to N grid points, is the error vector; [·] T represents transpose; is the transfer matrix between M array elements and N grid points: in, is the transfer vector between the M array elements and the i-th grid; S22: Assume that S sound sources are sparsely distributed in the free field. Since S<<N, q is a sparse vector. Since M<N, this results in an underdetermined problem. The Laplace norm is introduced as a sparse constraint to obtain a unique solution. The sound source localization problem is then modeled as a minimization problem under the Laplace norm sparse constraint: Where λ>0 is the regularization parameter, ||·||2 represents the l2 norm; |||·||| γ represents the Laplace norm, which is defined as: Among them, γ>0 is the order of the Laplace norm; when γ tends to zero, it is regarded as an approximation of the l0 norm.
4. The sound source localization method according to claim 3, wherein: Step S3 specifically includes the following steps: S31: Introducing an additional l1 norm term, the minimization model of the Laplace norm sparse constraint is constructed into the following differential form: Where F(q) is the overall objective function, H(q) is the sub-objective function 1, G(q) is the sub-objective function 2, and α>0 is a constant; S32: H(q) is directly solved using the iterative shrinkage threshold algorithm, which is equivalent to solving the following expression: Where V(q) is the gradient of the squared term in H(q): V(q)=q+A H (P-Aq) in,[] H represents the conjugate transpose; then the iterative expression for solving H(q) is: q k+1 =S λ (q k +A H (P-Aq k )) Where k is the number of iterations; S λ (q)=[s λ (q1),s λ (q2),…,s λ (q N )] T For threshold operation: Among them, |·| is the modulus length of the complex number; sign(q i ) is a symbolic function: in, represents the field of real numbers; S33: For the solution of G(q), its linear approximation is: G(q)≈G(q k )+<G'(q k ),(q-q k )> Where <·> represents the inner product, G'(q)=[G'(q1),…,G'(q N )] T is the gradient of G(q): The iterative expression for solving the minimization problem of the Laplace norm sparsity constraint in the framework of the iterative shrinkage threshold algorithm is: The expression of D(q) is: Repeat the above iterative solution process and set the number of iterations until the objective function reaches the iteration stopping condition.
5. The sound source localization method according to claim 4, characterized in that: Step S4 specifically includes the following steps: S41: When the regularization term is separable, solving the expression for H(q) is equivalent to the following univariate problem: Among them, v(q) is the single variable corresponding to the quadratic gradient V(q) in the objective function H(q); q k+1 The iterative expression is h v '(q)=0; however, h v '(q) is a transcendental equation. Considering that the physical meaning of q is sound intensity, the following expression is obtained: Obviously, v(q) = αλ / 2γ can be used as a threshold for iteration. Therefore, the appropriate regularization parameter can be determined based on this threshold. Rearrange the elements in V(q) by size: |[V(q k )]1|≥|[V(q k )]2|≥…≥|[V(q k )] N | Among them, |[V(q k )] i | is the i-th largest element in V(q); after sorting, the first S elements in V(q) are larger than the threshold αλ / 2γ, so for the S-th element and the S+1-th element, the following inequality holds: S42: Without loss of generality, we no longer strictly distinguish between the parameters λ and αλ; at the kth iteration, the strategy for selecting the regularization parameter is: l k =2γ((1-μ)|[V(q k )] S+1 |+μ|[V(q k )] S |) Among them, the value range of μ is μ∈[0,1); λ k The value of increases with the increase of μ, and λ k The larger the value, the stronger the sparsity; therefore, select μ = 0, generally: λ k =min{λ k-1 ,2γ|[V(q k )] S+1 |}.
6. The sound source localization method according to claim 5, characterized in that: Step S5 specifically includes the following steps: S51: Solve the non-convex Laplace norm constrained minimization problem by iterative shrinkage threshold algorithm. The solution is the sound intensity q of the potential sound source corresponding to each grid. Assume that the sparsity of the optimal solution is S, that is, the number of sound sources is regarded as prior information, and the first S largest elements in q are regarded as the estimated sound intensity of the sound source, denoted by q on ; S52: By searching all grids, q on The corresponding grid coordinates are regarded as the estimated value of the sound source location, denoted as r on ; Therefore, the estimated result of the sound source information is: in, It means taking the values corresponding to the first S largest elements.
Citation Information
Patent Citations
Sound source identification method based on Laplace norm fast iteration shrinkage threshold
CN111257833A