Two-dimensional off-network compression beam forming sound source localization method based on Gauss-Newton method
By applying the Gaussian Newton method to correct the sound source coordinates and intensity in the off-grid compressed beamforming method, the problems of basis mismatch and high computing resource consumption in the prior art are solved, and sound source positioning with higher accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202510094135.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-23
AI Technical Summary
The existing off-grid compressed beamforming method has problems such as limited correction accuracy and high computing resource consumption in the problem of basis mismatch, especially in the case of low signal-to-noise ratio, which affects the positioning performance.
The two-dimensional off-grid compression beamforming method based on the Gaussian Newton method is used to conduct preliminary estimates through the rapid iterative shrinkage threshold algorithm, and the Gaussian Newton method is used to iteratively correct the sound source coordinates and intensity to improve positioning accuracy.
It effectively alleviates the fundamental mismatch problem caused by the sound source off-grid, improves positioning accuracy and calculation efficiency, and can more accurately correct the sound source information in the case of low signal-to-noise ratio.
Smart Images

Figure CN120028753A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of sound source localization, and relates to a two-dimensional off-grid compressed beamforming sound source localization method based on Gauss-Newton method. Background Art
[0002] In recent years, sound source localization technology based on microphone arrays has been widely used in the fields of noise control, target tracking and fault detection. In the field of fault diagnosis and maintenance, when complex mechanical systems or electrical equipment fail or have abnormalities, they are often accompanied by sound signals such as mechanical vibration or abnormal discharge. Therefore, sound source localization technology can timely detect and lock the location of the sound source, providing reference information for equipment fault diagnosis and maintenance, thereby reducing equipment downtime and extending service life.
[0003] Acoustic beamforming technology is an important method in the field of sound source localization. Since the introduction of compressed sensing theory into acoustic beamforming, the sound source localization method based on compressed beamforming has developed rapidly, which has high positioning accuracy and computational efficiency. The general process of compressed beamforming is to discretize the sound source plane into a set of grids, and assume that the coordinates of each grid may be the potential position of the sound source. A certain algorithm is used to solve the intensity of the potential sound source on the grid, and then the grid coordinates are searched according to the sound intensity to obtain the estimated value of the sound source information. This type of method is called an on-grid compressed beamforming model. However, in practical applications, the sound source position does not always fall precisely on the grid point, which leads to the basis mismatch problem and affects the positioning performance of the algorithm. Refining the grid can alleviate this problem to a certain extent, but excessive refinement increases the consumption of computing resources and increases the correlation of the measurement matrix columns, and may even cause positioning failure. In response to this, researchers proposed off-grid compressed beamforming. This method calculates the compensation correction amount of the sound source information based on the Taylor approximation of the grid coordinates closest to the off-grid sound source, which can effectively reduce the impact of basis mismatch on positioning performance. However, this method is based on the first-order Taylor expansion, and the correction accuracy is limited. In addition, all sound sources are corrected at the same time. There is a situation where the overall objective function is reduced but the correction of one or more sound sources is invalid, especially in the case of low signal-to-noise ratio, which may even further increase the positioning deviation.
[0004] Therefore, how to further estimate more accurate sound source information and ensure the effectiveness of off-grid compressed beamforming for sound source correction is one of the key issues in achieving accurate and efficient positioning of sound sources. Therefore, it is necessary to adopt a method that can achieve high-precision positioning without significantly increasing computational complexity and computing resource consumption to overcome the shortcomings of existing technologies and solve or alleviate the above problems. Summary of the invention
[0005] In view of this, the purpose of the present invention is to provide a two-dimensional off-grid compressed beamforming sound source localization method based on the Gauss-Newton method, which can effectively overcome the basis mismatch problem caused by the off-grid sound source and ensure the effectiveness of the off-grid model in correcting the sound source information, thereby achieving fast and accurate sound source localization.
[0006] In order to achieve the above object, the present invention provides the following technical solutions:
[0007] A two-dimensional off-grid compressed beamforming sound source localization method based on the Gauss-Newton method is proposed. First, the sound source plane is divided into a series of discrete grids to establish an on-grid compressed beamforming sound source localization model. Then, the model is solved using the fast iterative shrinking threshold algorithm (FISTA) to obtain an on-grid preliminary estimation result of the sound source information. Finally, in order to overcome the basis mismatch problem caused by the off-grid sound source, a two-dimensional off-grid compressed beamforming sound source localization model is constructed based on the Gauss-Newton method. The sound source coordinates and intensity of the initially obtained single sound source are iteratively corrected to obtain a more accurate sound source information estimation result.
[0008] The method specifically comprises the following steps:
[0009] S1: Divide the plane where the sound source is located into a series of discrete grid points, assuming that the grid coordinates are the potential locations of the sound source;
[0010] S2: Based on the Green function, a measurement matrix between the microphone array and the N potential sound sources corresponding to the grid is obtained, thereby constructing an on-grid compressed beamforming model;
[0011] S3: Solve l using fast iterative shrinkage threshold algorithm 1 In the norm-sparse constrained on-grid compressed beamforming model, the solution is the strength of N potential sound sources corresponding to the grid. Based on the prior information that the sparsity of the solution is S, the grid coordinates corresponding to the first S elements in terms of intensity are searched, and the preliminary estimation result of the sound source information is obtained.
[0012] S4: Single correction: A two-dimensional off-grid compressed beamforming model is constructed based on the Gauss-Newton method. Based on the preliminary estimation of the sound source coordinates and intensity, the single sound source information is corrected, and the S sound sources are corrected one by one;
[0013] S5: Global correction: After completing the correction of a single sound source, a global correction is performed on all sound source intensities to speed up the convergence. The single correction and global correction are repeated until the objective function reaches the iteration termination condition to obtain a more accurate sound source information estimation result.
[0014] Further, step S1 specifically includes the following steps:
[0015] S11: The microphone array consists of M array elements, and the coordinates of the array elements are rm =(x m ,y m ,z m ) There are S sound sources in the free field space, and the coordinates of the sound sources are r s =(x s ,y s ,z s ) ; Place the microphone array in parallel at h meters in front of the target sound source plane to collect sound field information. The distance between the mth array element and the sth sound source is defined as d ms :
[0016]
[0017] S12: The complex sound pressure output by the data channel corresponding to the mth array element is Represents a complex field:
[0018]
[0019] in, is the intensity of the sth sound source; is the error of the mth channel, including environmental noise, model error, and measurement error; is the Green function between the mth array element and the sth sound source:
[0020]
[0021] Where j is the imaginary unit, κ = 2πf / c is the wave number, f is the frequency, and c is the speed of sound;
[0022] S13: Divide the target sound source plane into a series of discrete grids with a spacing of δ to obtain N grid coordinates. Assuming that each grid coordinate may be a potential location of the sound source, the complex sound pressure output by the data channel corresponding to the mth array element can be re-expressed as:
[0023]
[0024] in, is the sound intensity of the potential sound source corresponding to the g-th grid, is the Green’s function between the mth array element and the corresponding sound source of the gth grid.
[0025] Further, step S2 specifically includes the following steps:
[0026] S21: Considering the N potential sound sources corresponding to the grid and the M array elements of the microphone array, the sound pressure Y output by all data acquisition channels of the microphone array can be written in the following matrix form:
[0027] Y=AX+N
[0028] in, is the sound pressure vector output by all data acquisition channels of the microphone array, is the sound intensity vector of the N potential sound sources corresponding to the grid, is the error vector; [·] T represents transpose; It is the measurement matrix between the N potential sound sources corresponding to the grid and the M array elements of the microphone array:
[0029] A=[a(r 1 ),…,a(r N )]
[0030] in, is the measurement vector between all array elements and the g-th grid:
[0031]
[0032] S22: Generally, the sound source is sparsely distributed in the free field, so the inequality S<<N holds, that is, X is a sparse vector, and in practice the number of microphone array elements M is usually less than the number of grids N, which makes solving the sound source intensity X an underdetermined inverse problem, and the sparse constraint is introduced to obtain a unique solution; l 0 The norm has natural sparsity and is the first choice as a sparse constraint, so the following l can be constructed 0 Norm-sparse constrained on-the-net compressed beamforming model:
[0033]
[0034] Solving this model directly is an NP-hard problem. 0 The norm convex relaxation is l 1 norm, and transforms into the following unconstrained l 1 Norm-sparse constrained on-the-net compressed beamforming model:
[0035]
[0036] where λ>0 is the regularization parameter, ||·|| 2 Indicates l 2 norm,||·|| 1 Indicates l 1 Norm.
[0037] Further, step S3 specifically includes the following steps:
[0038] S31: Solve l using fast iterative shrinkage threshold algorithm 1 The iterative expression of the in-network compressed beamforming model with norm sparsity constraint is as follows:
[0039]
[0040] Among them, t is the iteration step, k is the number of iterations, Z is the intermediate variable, S λ (Z)=[s λ (z 1 ),s λ (z 2 ),…,s λ (z N )] T For soft thresholding:
[0041] s λ (z i )=max{|z i |-λ,0}·sign(z i )
[0042] Among them, |·| represents the modulus length of the complex number, sign(z i ) is a symbolic function:
[0043]
[0044] in, represents the real number domain; in the above iterative expression, the iteration threshold is α k λ, α k The expression is:
[0045]
[0046] Among them, G(X) is the gradient of the objective function, expressed as follows:
[0047] G(X)=A H (Y-AX)
[0048] S32: Initialize k=1, t 0 =1,X 0 and X 1 is a zero vector, and the number of iterations is set to start the iteration until the objective function meets the iteration stop condition; 1 The solution X of the in-grid compressed beamforming model with norm sparsity constraint is the strength of the N potential sound sources corresponding to the grid; the number of sound sources S is regarded as prior information, that is, the sparsity of the solution is S; the first S elements with the largest intensity in X are taken as the estimated value of the sound source intensity, recorded as X on ; Search X on The corresponding grid coordinates are regarded as the estimated value of the sound source position, denoted as R on ; Search X on The S columns in the corresponding measurement matrix A are regarded as the measurement matrix between the microphone array and the estimated sound source, denoted by A on; Therefore, the preliminary estimation result of the sound source can be obtained by the following operation:
[0049]
[0050] in, It means taking the matching values corresponding to the first S largest elements.
[0051] Further, step S4 specifically includes the following steps:
[0052] S41: Traditional off-grid compressed beamforming is based on the first-order Taylor approximation to compensate for the coordinates of the off-grid sound source, and the off-grid sound source is corrected based on the preliminary estimation of the on-grid coordinates; the Newton method can obtain more accurate positioning results by using the second-order Taylor approximation; the Gauss-Newton method is an improvement on the Newton method, which realizes the linear approximation of the second-order derivative term of the Newton method through the Jacobian matrix, which can improve the calculation efficiency while maintaining the accuracy of the estimation result; the iterative update rules of the Gauss-Newton method are as follows:
[0053]
[0054] Among them, x is a variable, is the correction amount for updating the variable x in the next iteration, f(x k ) is the objective function, F(x)=1 / 2||f(x)|| 2 =1 / 2f(x) T f(x), J(x k ) is f(x k )’s Jacobian matrix; F’(x k ) is F(x k ), is F(x k ) is an approximation of the second-order derivative of; when {x|F(x k )≤F(x 0 )} is satisfied and the Jacobian matrix is full rank in all iterative steps, the convergence of this method can be guaranteed;
[0055] S42: In the sound source localization scenario, the objective function is to minimize the residual between the microphone array measurement signal and the signal estimated by the localization method; the more accurate the sound source coordinates and intensity are estimated, the smaller the residual; for the sth sound source, the objective function for a single sound source can be obtained as follows
[0056]
[0057] in, To remove the contribution of the identified s-1 sound sources, the residual is expressed as follows:
[0058]
[0059] a(r in the objective function s ) is the measurement matrix A on The sth column in X s For X on The Jacobian matrix of the objective function corresponding to the s-th sound source is defined as:
[0060]
[0061] Among them, a x '(r s ) and a y '(r s ) can be calculated by the following expression:
[0062]
[0063] Therefore, for the sth sound source, at the k+1th iteration, the coordinate correction formula of the sound source can be expressed as:
[0064]
[0065] Among them, Δr s-GN =[Δx s ,Δy s ] T is the off-grid coordinate correction of the sth sound source calculated based on the Gauss-Newton method, and The expression is defined as:
[0066]
[0067] Where Re(·) means taking the real part of a complex vector or matrix;
[0068] After obtaining the corrected coordinates of the sth sound source Then, update the Green's function corresponding to the corrected coordinates Then the intensity of the sth sound source can be modified according to the following expression:
[0069]
[0070] Further, step S5 specifically includes the following steps:
[0071] S51: After correcting the S sound sources one by one according to the coordinates and intensity correction formula of a single sound source, a global correction of the intensity according to the following formula is beneficial to speed up the convergence speed and obtain a more accurate intensity estimation result:
[0072]
[0073] Among them, A S are the S Green functions corresponding to the corrected coordinates The measurement vector composed of
[0074] S52: Repeat the above steps of single correction and global correction of the sound source until the objective function reaches the iteration termination condition, and obtain the l corrected based on the Gauss-Newton method. 1 Estimation results of off-grid sound source location and intensity using two-dimensional off-grid compressive beamforming with norm sparsity constraint.
[0075] The beneficial effects of the present invention are as follows: the present invention can effectively alleviate the basis mismatch problem caused by the off-grid sound source through the off-grid compressed beamforming model, and when correcting the off-grid sound source information, the Gauss-Newton method is used to calculate the correction compensation amount of the off-grid coordinates. Compared with the traditional method based on the first-order Taylor approximation, the method of the present invention can achieve more accurate and more effective correction while maintaining a high computational efficiency.
[0076] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:
[0078] Figure 1 It is a flow chart of a two-dimensional off-grid compressed beamforming sound source localization method based on Gauss-Newton method of the present invention;
[0079] Figure 2 This is a schematic diagram of the microphone array signal measurement scenario;
[0080] Figure 3 This is the 56-channel multi-arm spiral microphone array used in the simulation experiment;
[0081] Figure 4 The simulation positioning effect comparison between the method of the present invention and two traditional methods when the signal-to-noise ratio is 20dB and the frequency f is 20kHz;
[0082] Figure 5 The simulation positioning effect comparison between the method of the present invention and two traditional methods when the signal-to-noise ratio is 0dB and the frequency f is 20kHz;
[0083] Figure 6The Monte Carlo simulation results of the method of the present invention and two traditional methods when the signal-to-noise ratio range is 0dB to 20dB and the frequency f is 20kHz;
[0084] Figure 7 The Monte Carlo simulation results of the method of the present invention and two traditional methods are shown when the frequency f ranges from 10kHz to 30kHz and the signal-to-noise ratio is 20dB. DETAILED DESCRIPTION
[0085] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.
[0086] See also Figure 1 to Figure 7 The present invention provides a two-dimensional off-grid compressed beamforming sound source localization method based on Gauss-Newton method, which specifically includes the following steps:
[0087] Step S1: Divide the plane where the sound source is located into a series of discrete grid points, assuming that the grid coordinates are the potential positions of the sound source, specifically including the following steps:
[0088] S11: The microphone array consists of M array elements, and the coordinates of the array elements are r m =(x m ,y m ,z m ) There are S sound sources in the free field space, and the coordinates of the sound sources are r s =(x s ,y s ,z s ) ; Place the microphone array in parallel at h meters in front of the target sound source plane to collect sound field information. The distance between the mth array element and the sth sound source is defined as d ms :
[0089]
[0090] S12: The complex sound pressure output by the data channel corresponding to the mth array element is Represents a complex field:
[0091]
[0092] in, is the intensity of the sth sound source, is the error of the mth channel, including environmental noise, model error, measurement error, etc. is the Green function between the mth array element and the sth sound source:
[0093]
[0094] Where j is the imaginary unit, κ = 2πf / c is the wave number, f is the frequency, and c is the speed of sound;
[0095] S13: Divide the target sound source plane into a series of discrete grids with a spacing of δ to obtain N grid coordinates. Assuming that each grid coordinate may be a potential location of the sound source, the complex sound pressure output by the data channel corresponding to the mth array element can be re-expressed as:
[0096]
[0097] in, is the sound intensity of the potential sound source corresponding to the g-th grid, is the Green’s function between the mth array element and the corresponding sound source of the gth grid.
[0098] Step S2: Obtain a measurement matrix between the microphone array and the N potential sound sources corresponding to the grid from the sound field Green's function, thereby constructing an on-grid compression beamforming model, which specifically includes the following steps:
[0099] S21: Considering the N potential sound sources corresponding to the grid and the M array elements of the microphone array, the sound pressure Y output by all data acquisition channels of the microphone array can be written in the following matrix form:
[0100] Y=AX+N
[0101] in, is the sound pressure vector output by all data acquisition channels of the microphone array, is the sound intensity vector of the N potential sound sources corresponding to the grid, is the error vector; [·] T represents transpose; It is the measurement matrix between the N potential sound sources corresponding to the grid and the M array elements of the microphone array:
[0102] A=[a(r 1 ),…,a(r N )]
[0103] in, is the measurement vector between all array elements and the g-th grid:
[0104]
[0105] S22: Generally, the sound source is sparsely distributed in the free field, so the inequality S<<N holds, that is, X is a sparse vector, and in practice the number of microphone array elements M is usually less than the number of grids N, which makes solving the sound source intensity X an underdetermined inverse problem, and the sparse constraint is introduced to obtain a unique solution; l 0 The norm has natural sparsity and is the first choice as a sparse constraint, so the following l can be constructed 0 Norm-sparse constrained on-the-net compressed beamforming model:
[0106]
[0107] Solving this model directly is an NP-hard problem. 0 The norm convex relaxation is l 1 norm, and transforms into the following unconstrained l 1 Norm-sparse constrained on-the-net compressed beamforming model:
[0108]
[0109] where λ>0 is the regularization parameter, ||·|| 2 Indicates l 2 norm,||·|| 1 Indicates l 1 Norm.
[0110] Step S3: Solve l using the fast iterative shrinkage threshold algorithm 1 The in-grid compressed beamforming model with norm sparsity constraint has the solution of the intensity of N potential sound sources corresponding to the grid. Based on the prior information that the sparsity of the solution is S, the grid coordinates corresponding to the first S elements in terms of intensity are searched to obtain the preliminary estimation result of the sound source information. Specifically, the following steps are included:
[0111] S31: Solve l using fast iterative shrinkage threshold algorithm 1 The iterative expression of the in-network compressed beamforming model with norm sparsity constraint is as follows:
[0112]
[0113] Among them, t is the iteration step, k is the number of iterations, Z is the intermediate variable, S λ (Z)=[s λ (z 1 ),s λ (z 2 ),…,s λ (z N )] T For soft thresholding:
[0114] s λ (zi ) = max{|z i | - λ, 0}·sign(z i )
[0115] where |·| represents the modulus of a complex number, and sign(z i ) is the sign function:
[0116]
[0117] where represents the real number field; in the above iterative expression, the iteration threshold is α k λ, α k The expressions for are:
[0118]
[0119] where G(X) is the gradient of the objective function, and the expression is as follows:
[0120] G(X) = A H (Y - AX)
[0121] S32: Initialize k = 1, t 0 = 1, X 0 and X 1 as zero vectors, set the number of iterations to start the iteration until the objective function satisfies the iteration stop condition; l 1 The solution X of the in-network compression beamforming model with l1-norm sparse constraint is the intensity of N potential sound sources corresponding to the grid; regard the number of sound sources S as prior information, that is, the sparsity of the solution is S; take the first S elements with the largest intensity in X as the estimated values of the sound source intensities, denoted as X on ; Search for the grid coordinates corresponding to X on and regard them as the estimated values of the sound source positions, denoted as R on ; Search for S columns in the measurement matrix A corresponding to X on and regard them as the measurement matrix between the microphone array and the estimated sound sources, denoted as A on ; Therefore, the preliminary estimation results of the sound sources can be obtained through the following operations:
[0122]
[0123] where represents taking the matching values corresponding to the first S elements with the largest values.
[0124] Step S4: Construct a two-dimensional off-grid compression beamforming model based on the Gauss-Newton method. On the basis of the preliminary estimated values of the sound source coordinates and intensities, correct the information of a single sound source, and correct S sound sources one by one, specifically including the following steps:
[0125] S41: Traditional off-grid compressed beamforming is based on the first-order Taylor approximation to compensate for the coordinates of the off-grid sound source, and the off-grid sound source is corrected based on the preliminary estimation of the on-grid coordinates; the Newton method can obtain more accurate positioning results by using the second-order Taylor approximation; the Gauss-Newton method is an improvement on the Newton method, which realizes the linear approximation of the second-order derivative term of the Newton method through the Jacobian matrix, which can improve the calculation efficiency while maintaining the accuracy of the estimation result; the iterative update rules of the Gauss-Newton method are as follows:
[0126]
[0127] Among them, x is a variable, is the correction amount for updating the variable x in the next iteration, f(x k ) is the objective function, F(x)=1 / 2||f(x)|| 2 =1 / 2f(x) T f(x), J(x k ) is f(x k )’s Jacobian matrix; F’(x k ) is F(x k ), is F(x k ) is an approximation of the second-order derivative of; when {x|F(x k )≤F(x 0 )} is satisfied and the Jacobian matrix is full rank in all iterative steps, the convergence of this method can be guaranteed;
[0128] S42: In the sound source localization scenario, the objective function is to minimize the residual between the microphone array measurement signal and the signal estimated by the localization method; the more accurate the sound source coordinates and intensity are estimated, the smaller the residual; for the sth sound source, the objective function for a single sound source can be obtained as follows
[0129]
[0130] in, To remove the contribution of the identified s-1 sound sources, the residual is expressed as follows:
[0131]
[0132] a(r in the objective function s ) is the measurement matrix A on The sth column in X s For X on The Jacobian matrix of the objective function corresponding to the s-th sound source is defined as:
[0133]
[0134] Among them, a x '(r s ) and a y '(r s ) can be calculated by the following expression:
[0135]
[0136] Therefore, for the sth sound source, at the k+1th iteration, the coordinate correction formula of the sound source can be expressed as:
[0137]
[0138] Among them, Δr s-GN =[Δx s ,Δy s ] T is the off-grid coordinate correction of the sth sound source calculated based on the Gauss-Newton method, and The expression is defined as:
[0139]
[0140] Where Re(·) means taking the real part of a complex vector or matrix;
[0141] After obtaining the corrected coordinates of the sth sound source Then, update the Green's function corresponding to the corrected coordinates Then the intensity of the sth sound source can be modified according to the following expression:
[0142]
[0143] Step S5: After completing the correction of a single sound source, a global correction is performed on all sound source intensities to speed up the convergence speed, and the single correction and global correction are repeated until the objective function reaches the iteration termination condition to obtain a more accurate sound source information estimation result, which specifically includes the following steps:
[0144] S51: After correcting the S sound sources one by one according to the coordinates and intensity correction formula of a single sound source, a global correction of the intensity according to the following formula is beneficial to speed up the convergence speed and obtain a more accurate intensity estimation result:
[0145]
[0146] Among them, A S are the S Green functions corresponding to the corrected coordinates The measurement vector composed of
[0147] S52: Repeat the above steps of single correction and global correction of the sound source until the objective function reaches the iteration termination condition, and obtain the l corrected based on the Gauss-Newton method. 1 Estimation results of off-grid sound source location and intensity using two-dimensional off-grid compressive beamforming with norm sparsity constraint.
[0148] Verification experiment:
[0149] The beneficial effects of the method of the present invention are demonstrated through simulation experiments:
[0150] Simulation experiment: The schematic diagram of the microphone array signal measurement scenario is as follows Figure 2 As shown in the figure, the microphone array is placed parallel to the sound source plane, and the distance between the two is set to 1.5 meters. Four sound sources are set to be distributed in the range of 1.2 meters × 1.2 meters. The target area is divided with a grid spacing of 0.1 meters, and a total of 169 grid points are obtained. That is, it is considered that the coordinates of these 169 grids may become the potential locations of the four sound sources. The simulation uses a 56-channel multi-arm spiral microphone array with an array diameter of about 0.2 meters. The position distribution of the array elements is shown in the figure. Figure 3 As shown in the figure, the sampling rate of the microphone array is set to 192kHz. The frequency and sound pressure level of the four sound sources are set to 20kHz and 93.98dB (1Pa); the sound source coordinates are set to S 1 (-0.55m,-0.37m), S 2 (0.05m,0.58m),S 3 (0.28m,0.30m) and S 4 (0.38m, 0.23m), the four sound sources are all deviated from the grid point, among which the sound source S 3 and S 4 The sound sources are set to be spatially close to each other to show the spatial resolution performance of the algorithm, with a spacing of 0.12 meters.
[0151] In the simulation experiment, the two-dimensional off-grid compressed beamforming sound source localization method based on the Gauss-Newton method of the present invention is compared with the on-grid compressed beamforming sound source localization method based on the fast iterative shrinkage threshold algorithm and the traditional off-grid compressed beamforming sound source localization method based on the first-order Taylor approximation. The three methods are represented as FISTA-GN, FISTA and FISTA-FT, respectively. Figure 4 The comparison of the simulation positioning effects of the three methods when the signal-to-noise ratio (SNR) is 20dB and the frequency f is 20kHz. In the figure, "*" and "o" represent the estimated sound source position and the real sound source position respectively. The dynamic display range is set to a difference of 10dB from the maximum sound intensity of the sound source. Figure 4It can be seen that the three methods can locate the positions of the four sound sources and distinguish the two sound sources that are adjacent in space. However, FISTA and FISTA-FT are interfered by pseudo sound sources and one or more sound sources cannot accurately identify the position of the sound source, and the spatial resolution is reduced. In contrast, the FISTA-GN of the present invention can achieve high-precision positioning, effectively suppress the interference of noise, and accurately estimate the intensity of the estimated sound source. In order to verify the positioning performance of the method of the present invention in a low signal-to-noise ratio environment, the signal-to-noise ratio (SNR) is set to 0dB and the frequency f is set to 20kHz for simulation positioning experiments. The results are as follows: Figure 5 shown by Figure 5 It can be seen that the FISTA method cannot accurately distinguish between sound sources and noise, and there are many pseudo sound source interferences in the positioning results. Therefore, this method fails in a low signal-to-noise ratio environment. In comparison, the two off-grid compressed beamforming methods can still correctly identify the four sound sources, but a pseudo sound source with the same intensity as the real sound source appears in the positioning results of the FISTA-FT method, which interferes with the positioning results. In actual engineering applications, the location needs to be checked, which increases the detection cost. At the same time, compared with FISTA-GN, the positioning accuracy of FISTA-FT is significantly reduced, which verifies the analysis that the traditional off-grid cultivation method based on first-order Taylor approximation may lead to worse positioning results under low signal-to-noise ratio.
[0152] Monte Carlo random simulation test: In order to further illustrate the effectiveness of the method of the present invention, Monte Carlo random tests of three methods under different signal-to-noise ratios and different frequencies were simulated, and 300 positioning experiments were carried out in each environment. In the simulation, four sound sources were randomly distributed in a target area of 1.2 m × 1.2 m, and the target area was divided with a grid spacing of 0.1 m. The root mean square error RMSEL of the sound source position and the root mean square error RMSES of the sound intensity were defined as evaluation indicators to evaluate the performance of the positioning method.
[0153] Set the frequency f range to 15kHz~35kHz, take the frequency value at intervals of 1kHz, and the signal-to-noise ratio SNR to 20dB. The Monte Carlo simulation results of the three methods are as follows Figure 6 As shown; the signal-to-noise ratio SNR range is set to 0dB~20dB, the SNR value is taken at 1dB intervals, the frequency f is 20kHz, and the Monte Carlo simulation results of the three methods are as follows Figure 7 shown by Figure 6 It can be seen that with the improvement of signal-to-noise ratio, the estimation errors of the three methods for the sound source position and intensity are reduced. When the SNR is greater than 5dB, the positioning error of the FISTA-GN method of the present invention is about 4cm, and the intensity error is less than 2dB. Figure 7It can be seen that within the simulation frequency band, the estimation errors of the FISTA and FISTA-FT methods increase with the increase of frequency, especially the estimation error of the sound intensity increases significantly, but the FISTA-GN method of the present invention can more accurately estimate the sound source information within the simulated signal-to-noise ratio range and frequency range; in general, compared with the on-grid compressed beamforming method FISTA, the positioning error and intensity error of the two methods based on the off-grid compressed beamforming method are significantly reduced; compared with the two off-grid compressed beamforming methods, the positioning accuracy and intensity estimation error of the method based on the Gauss-Newton method of the present invention are smaller, and the robustness is stronger.
[0154] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.
Claims
1. A two-dimensional off-grid compressed beamforming sound source localization method based on Gauss-Newton method, characterized in that: The method specifically comprises the following steps: S1: Divide the plane where the sound source is located into a series of discrete grid points, assuming that the grid coordinates are the potential locations of the sound source; S2: Based on the Green function, the measurement matrix between the microphone array and the N potential sound sources corresponding to the grid is obtained, so as to construct the grid compression beamforming model; S3: Using a fast iterative shrinkage threshold algorithm to solve the l1-norm sparse constraint on-grid compressed beamforming model, the solution is the intensity of N potential sound sources corresponding to the grid. Based on the prior information that the sparsity of the solution is S, the grid coordinates corresponding to the first S elements in terms of intensity are searched, and the preliminary estimation result of the sound source information is obtained; S4: Single correction: A two-dimensional off-grid compressed beamforming model is constructed based on the Gauss-Newton method. Based on the preliminary estimation of the sound source coordinates and intensity, the single sound source information is corrected, and the S sound sources are corrected one by one; S5: Global correction: After completing the correction of a single sound source, a global correction is performed on the intensities of all sound sources. The single correction and global correction are repeated until the objective function reaches the iteration termination condition to obtain an accurate sound source information estimation result.
2. The two-dimensional off-grid compressed beamforming sound source localization method according to claim 1, characterized in that: Step S1 specifically includes the following steps: S11: The microphone array consists of M array elements, and the coordinates of the array elements are r m =(x m ,y m ,z m ) There are S sound sources in the free field space, and the coordinates of the sound sources are r s =(x s ,y s ,z s ) ; Place the microphone array in parallel at h meters in front of the target sound source plane to collect sound field information. The distance between the mth array element and the sth sound source is defined as d ms : S12: The complex sound pressure output by the data channel corresponding to the mth array element is Represents a complex field: in, is the intensity of the sth sound source; is the error of the mth channel; is the Green function between the mth array element and the sth sound source: Where j is the imaginary unit, κ = 2πf / c is the wave number, f is the frequency, and c is the speed of sound; S13: Divide the target sound source plane into a series of discrete grids with a spacing of δ to obtain N grid coordinates. Assuming that each grid coordinate may be a potential location of the sound source, the complex sound pressure output by the data channel corresponding to the mth array element is re-expressed as: in, is the sound intensity of the potential sound source corresponding to the g-th grid, is the Green’s function between the mth array element and the corresponding sound source of the gth grid.
3. The two-dimensional off-grid compressed beamforming sound source localization method according to claim 2, characterized in that: Step S2 specifically includes the following steps: S21: Considering the N potential sound sources corresponding to the grid and the M array elements of the microphone array, the sound pressure Y output by all data acquisition channels of the microphone array is written in the following matrix form: Y=AX+N in, is the sound pressure vector output by all data acquisition channels of the microphone array, is the sound intensity vector of the N potential sound sources corresponding to the grid, is the error vector; [·] T represents transpose; It is the measurement matrix between the N potential sound sources corresponding to the grid and the M array elements of the microphone array: A=[a(r1),…,a(r N )] in, is the measurement vector between all array elements and the g-th grid: S22: Construct the following l0-norm sparse constrained on-line compressed beamforming model: Solving this model directly is an NP-hard problem. The l0 norm is convexly relaxed to the l1 norm and transformed into the following unconstrained l1 norm sparse constrained on-net compressed beamforming model: Here, λ>0 is the regularization parameter, ||·||2 represents the l2 norm, and ||·||1 represents the l1 norm.
4. The two-dimensional off-grid compressed beamforming sound source localization method according to claim 3, characterized in that: Step S3 specifically includes the following steps: S31: The iterative expression of the on-line compressed beamforming model for solving the l1-norm sparse constraint using the fast iterative shrinkage threshold algorithm is as follows: Among them, t is the iteration step, k is the number of iterations, Z is the intermediate variable, S λ (Z)=[s λ (z1),s λ (z2),…,s λ (z N )] T For soft thresholding: s λ (With i )=max{|z i |-λ,0} sign(z i ) Among them, |·| represents the modulus length of the complex number, sign(z i ) is a symbolic function: in, represents the real number domain; in the above iterative expression, the iteration threshold is α k λ, α k The expression is: Among them, G(X) is the gradient of the objective function, expressed as follows: G(X)=A H (Y-AX) S32: Initialize k=1, t0=1, X0 and X1 as zero vectors, set the number of iterations and start iterating until the objective function meets the iteration stop condition; the solution X of the e1-norm sparse constrained on-grid compressed beamforming model is the strength of the N potential sound sources corresponding to the grid; the number of sound sources S is regarded as prior information, that is, the sparsity of the solution is S; take the first S elements with the largest intensity in X as the estimated value of the sound source intensity, recorded as X on ; Search X on The corresponding grid coordinates are regarded as the estimated value of the sound source position, denoted as R on ; Search X on The S columns in the corresponding measurement matrix A are regarded as the measurement matrix between the microphone array and the estimated sound source, denoted by A on ; Therefore, the preliminary estimation result of the sound source can be obtained by the following operation: in, It means taking the matching values corresponding to the first S largest elements.
5. The two-dimensional off-grid compressed beamforming sound source localization method according to claim 4, characterized in that: Step S4 specifically includes the following steps: S41: The iterative update rule of the Gauss-Newton method is as follows: Among them, x is a variable, is the correction amount for updating the variable x in the next iteration, f(x k ) is the objective function, F(x)=1 / 2||f(x)|| 2 =1 / 2f(x) T f(x), J(x k ) is f(x k )’s Jacobian matrix; F’(x k ) is F(x k ), is F(x k ) is an approximation of the second-order derivative of; when {x|F(x k )≤F(x 0 )} is satisfied and the Jacobian matrix is full rank in all iterative steps, the convergence of this method can be guaranteed; S42: In the sound source localization scenario, the objective function is to minimize the residual between the microphone array measurement signal and the signal estimated by the localization method; the more accurate the sound source coordinates and intensity are estimated, the smaller the residual; for the sth sound source, the objective function for a single sound source is obtained as follows in, To remove the contribution of the identified s-1 sound sources, the residual is expressed as follows: a(r in the objective function s ) is the measurement matrix A on The sth column in X s For X on The Jacobian matrix of the objective function corresponding to the s-th sound source is defined as: Among them, a x '(r s ) and a y '(r s ) is calculated by the following expression: Therefore, for the sth sound source, at the k+1th iteration, the coordinate correction formula of the sound source is expressed as: Among them, Δr s-GN =[Δx s ,Δy s ] T is the off-grid coordinate correction of the sth sound source calculated based on the Gauss-Newton method, and The expression is defined as: Where Re(·) means taking the real part of a complex vector or matrix; After obtaining the corrected coordinates of the sth sound source Then, update the Green's function corresponding to the corrected coordinates Then the intensity of the sth sound source is modified according to the following expression:
6. The two-dimensional off-grid compressed beamforming sound source localization method according to claim 5, characterized in that: Step S5 specifically includes the following steps: S51: After completing the correction of the S sound sources one by one according to the coordinates and intensity correction formula of a single sound source, a global correction is performed on the intensity according to the following formula: Among them, A S are the S Green functions corresponding to the corrected coordinates The measurement vector composed of S52: Repeat the above single correction and global correction steps for the sound source until the objective function reaches the iteration termination condition, and obtain the estimation results of the off-grid sound source position and intensity based on the l1 norm sparse constraint two-dimensional off-grid compressed beamforming modified by Gauss-Newton method.