A deconvolution high-resolution sound source localization method, device, equipment and storage medium based on non-convex L1-L2 sparse regularization

By using a deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization, utilizing microphone arrays and sound source propagation models, combined with L1-L2 regularization and gradient descent algorithms, the problems of low sound source localization accuracy and computational efficiency in existing technologies are solved, achieving higher resolution and more efficient sound source localization.

CN119104983BActive Publication Date: 2025-09-09NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411156766.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2025-09-09
Estimated Expiration
2044-08-22

AI Technical Summary

Technical Problem

Existing deconvolution high-resolution sound source localization methods suffer from sparsity in spatially distributed grid points, which affects the sound source intensity and positioning accuracy, and has low computational efficiency.

Method used

A high-resolution sound source localization method based on deconvolution of non-convex L1-L2 sparse regularization is adopted. The sound pressure signal is obtained through a microphone array, and the low-resolution beam pattern is calculated using a preset delay-sum algorithm. Combined with the sound source propagation model and L1-L2 regularization, a non-convex L1-L2 sparse regularized deconvolution model is constructed. The high-resolution beam pattern is obtained by iterative solution through the proximal algorithm and gradient descent algorithm of L1-L2 regularization.

Benefits of technology

The spatial resolution and robustness of sound source localization are improved, more efficient calculation and higher energy concentration are achieved, and the accuracy and computational efficiency of sound source localization are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119104983B_ABST
    Figure CN119104983B_ABST
Patent Text Reader

Abstract

The present invention discloses a deconvolution high-resolution sound source localization method, device, equipment and storage medium based on non-convex L1-L2 sparse regularization. Based on the sound source propagation model, the low-resolution beam pattern is rewritten, and the generation mechanism of the rewritten low-resolution beam pattern and the influence of the point spread function (PSF) are analyzed to construct a deconvolution model; L1-L2 regularization is added to the deconvolution model as a sparse constraint to construct a non-convex L1-L2 sparse regularized deconvolution high-resolution localization model; based on the L1-L2 regularized proximal algorithm and gradient descent algorithm, the non-convex L1-L2 sparse regularized deconvolution high-resolution localization model is iteratively solved to obtain a high-resolution beam pattern to determine the sound source localization information. In this way, the geometric structure advantage of the present invention, which is closer to the L0 norm, effectively improves the spatial resolution, robustness and positioning accuracy of sound source localization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of spatial sound source localization, and in particular to a deconvolution high-resolution sound source localization method, apparatus, device and storage medium based on non-convex L1-L2 sparse regularization. Background Art

[0002] Beamforming is a signal processing technique used to locate sound sources using microphone arrays. Beamforming overcomes the inability of far-field sound source reconstruction techniques to locate sound sources, making it an important tool for sound source localization. Deconvolution sound source imaging, a high-resolution noise source identification and localization technique based on traditional beamforming technology, achieves significantly higher spatial resolution than conventional beamforming techniques. Therefore, high-resolution deconvolution beamforming methods are crucial.

[0003] The classic deconvolution beamforming method DAMAS algorithm has higher resolution than the traditional beamforming delayed-sum DAS method, but it is not easy to be effectively used in actual engineering due to the large dimension of the point spread function, low computational efficiency, and long time consumption. Subsequently, many methods were proposed to improve computational efficiency without affecting resolution, such as DAMAS2, DAMAS3, and FFT-NNLS algorithms. However, none of the above-mentioned improved deconvolution beamforming methods take into account the sparsity of sound sources in the spatially distributed grid points. In other words, the existing deconvolution high-resolution sound source localization methods usually have the disadvantage of sparsity in the spatially distributed grid points, which affects the sound source intensity and the accuracy of sound source localization.

[0004] In related technologies, the ideal sparsity constraint is the L0 norm, which can impose the same penalty on non-zero elements. The L0 norm is essentially a non-convex function, and directly solving the L0 norm minimization problem is a non-deterministic polynomial hard (NP-hard) problem. The usual method is to use the optimal convex approximation of the L0 norm. Therefore, the L0 norm is always relaxed to the convex L1 norm and is widely used, such as SC-DAMAS, FFT-FISTA, and SALSA. The L1 norm cannot optimally approximate the L0 norm, and other regularization terms have been introduced, such as TVNCD for the convex norm class and L1-L2DAMAS for the L1+L2 norm, and L1 with the 1 / 2 norm for the non-convex norm class. 1 / 2 -DAMAS.

[0005] However, the regularization terms used to describe sparsity in these sparse deconvolution beamforming methods are insufficient to describe the ideal L0 norm performance, which limits the accuracy of sound source localization. Summary of the Invention

[0006] In order to solve the above-mentioned technical problems existing in the prior art, the present invention provides a deconvolution high-resolution sound source localization method, device, equipment and storage medium based on non-convex L1-L2 sparse regularization to solve the disadvantage that the deconvolution high-resolution sound source localization method in the prior art usually has sparsity in the spatially distributed grid points, which affects the sound source intensity and the sound source localization accuracy.

[0007] In order to achieve the above-mentioned purpose, the technical solution of the embodiment of the present invention is:

[0008] A first aspect of the present invention provides a deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization, the method comprising:

[0009] S1: Use a microphone array with M elements to obtain the sound pressure signal of the sound source and calculate a low-resolution beam pattern based on a preset delay-sum algorithm;

[0010] S2: Based on a sound source propagation model, rewrite the low-resolution beam pattern, analyze the generation mechanism of the rewritten low-resolution beam pattern and the influence of the point spread function (PSF), and construct a deconvolution model; the sound source propagation model refers to a relationship model between the sound source signal and the sound pressure signal;

[0011] S3: Add L1-L2 regularization as a sparse constraint to the deconvolution model to construct a non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model;

[0012] S4: Based on the L1-L2 regularized proximal algorithm and the gradient descent algorithm, the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model is iteratively solved to obtain a high-resolution beam map;

[0013] S5: Determine sound source localization information based on the high-resolution beam pattern.

[0014] In some embodiments, the low-resolution beam pattern b(r) in S1 is:

[0015]

[0016] Where p is the sound pressure signal, M is the number of array elements, r is the distance from the focus point to the center of the microphone array, and p H is the conjugate transpose of p, represents the energy of the sound pressure signal, C is the cross-spectral matrix, v(r)=[v1(r),v2(r),…,v M (r)] represents the steering vector of the focus point r on the focusing surface, and the mth element in the steering vector is v(r) H is the conjugate transpose of v(r), (f) Hrepresents the conjugate transpose, Indicates summed average.

[0017] In some embodiments, the rewritten low-resolution beam pattern in S2 is:

[0018]

[0019] Where r is the distance from the focus point to the center of the microphone array, r s is the distance from the sth focal point to the center of the microphone array, S is the number of sound sources, k is the wave number, M is the number of array elements, and the sound source intensity is q = [q1,q2,…,q s ] T ,q s is the sound source intensity q at the sth gathering point s =(jωρQ s / 4π|r s |), v(r) H is the conjugate transpose of v(r), Green's function is the Green's function corresponding to the s-th focusing point, for The conjugate transpose of .

[0020] In some embodiments, the point spread function (PSF) in S2 is defined as the response of the beamformer to a single sound source and is expressed as:

[0021]

[0022] In the case of a translation-invariant PSF,

[0023] Based on discrete spatial Fourier transform, the frequency domain convolution is converted into beam domain product for acceleration, and the deconvolution model is constructed. The deconvolution model is:

[0024] Where r is the distance from the focus point to the center of the microphone array, M is the number of array elements, and v(r) H is the conjugate transpose of v(r), g rs is the Green's function corresponding to the s-th focusing point, g rs The conjugate transpose of * represents the convolution operation, X represents the “clean” beamforming pattern, B represents the matrix composed of the traditional beamforming pattern, P represents the PSF of the focal plane, o represents the Hadamard product, and represent two-dimensional Fourier transform and inverse Fourier transform, respectively.

[0025] In some embodiments, the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model in S3 is:

[0026]

[0027] subject to X≥0.

[0028] Where X represents the “clean” beamforming pattern, B represents the matrix of the conventional beamforming pattern, P represents the PSF of the focal plane, and o represents the Hadamard product. and

[0029] denote the two-dimensional Fourier transform and inverse Fourier transform, respectively, ||·|| F is the Frobenius norm, ||·||1 is the 1-norm, ||·||2 is the 2-norm, λ is the regularization parameter, and α is the sparse weight parameter.

[0030] In some embodiments, the first iteration process in S4 includes:

[0031]

[0032] in, is the proximal operator of L1-αL2, X represents the “clean” beamforming map, and the gradient Expressed as The Lipschitz constant L satisfies The Lipschitz constant is the largest eigenvalue λ of the Hessian matrix max ; Based on the power method effective estimation, the expression is:

[0033]

[0034] Where X1 is an arbitrary matrix with the same shape as the matrix P, the Lipschitz constant L is obtained after 10 iterations, P represents the PSF of the focal plane, and o represents the Hadamard product. and denote two-dimensional Fourier transform and inverse Fourier transform respectively;

[0035] The proximal operator of the backward L1-αL2 in the iterative process is defined as follows:

[0036]

[0037] The closed-form solution of the L1-αL2 proximal operator in the iterative process is given (positive parameter λ>0):

[0038] Given λ>0, and α≥0, the optimal solution x in the optimization problem of the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model is * There are the following statements:

[0039] When ||y|| ∞ >λ,x * =z(||z||2+αλ) / ||z||2,

[0040] When ||y|| ∞ =λ, if and only if |y i |<λ,||x * ||2=αλ and for all i, When x * There is a unique solution that satisfies When multiple components of y have the largest absolute value, the optimal solution is not unique; that is, there are infinitely many optimal solutions;

[0041] When (1-α)λ<||y|| ∞ <λ, if and only if |y i |<||y|| ∞ ,||x * ||2=||y|| ∞ +(α-1)λ and for all i, When x * There is a unique solution, which is a single sparse vector, satisfying The number of optimal solutions and the number of solutions with the largest absolute value ||y|| ∞ The number of components is the same;

[0042] When ||y|| ∞ ≤(1-α)λ, The proximal operator for L1 is called soft contraction and is defined as:

[0043]

[0044] in, is the optimal solution, y is the input of the proximal operator, y i is the i-th element of the y vector, z is the soft threshold operator, i is the number of elements, α is the sparse weight parameter, and λ is the regularization parameter.

[0045] In some embodiments, the overall iterative process in S4 includes:

[0046] Step 1: Update X (k)

[0047]

[0048] Step 2: (k+1)

[0049]

[0050] Step 3: Y (k+1)

[0051]

[0052] Repeat the iteration until ||X (k+1) -X (k) || F ≤εmax(||X (k) || F ,1), obtain the high-resolution beam pattern;

[0053] Where ε is 10 -5 , t (1) is 1, Y (1) is a zero matrix, X (k) is the X value of the kth iteration. The X obtained when the iteration stops is the high-resolution beam pattern. k is the value of t at the kth iteration, t (k+1) is the value of t at the k+1th iteration, where t is the momentum parameter used to accelerate the algorithm calculation time, and Y (k+1) is the value of Y at the k+1th iteration, ||X (k+1) -X (k) || F is the F norm of the difference between the k+1th and kth iterations, ||X (k) || F is the F norm of the k-th iteration X.

[0054] An embodiment of the present invention provides a deconvolution high-resolution sound source localization device based on non-convex L1-L2 sparse regularization, the device comprising:

[0055] A calculation module is used to obtain a sound pressure signal of a sound source using a microphone array having M elements, and calculate a low-resolution beam pattern based on a preset delay-sum algorithm;

[0056] A construction module is configured to rewrite the low-resolution beam pattern based on a sound source propagation model, analyze the generation mechanism of the rewritten low-resolution beam pattern and the influence of the point spread function (PSF), and construct a deconvolution model; the sound source propagation model is a relationship model between the sound source signal and the sound pressure signal;

[0057] The construction module is further used to add L1-L2 regularization as a sparse constraint to the deconvolution model to construct a non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model;

[0058] A solution module, configured to iteratively solve the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model based on the L1-L2 regularized proximal gradient descent algorithm to obtain a high-resolution beam map;

[0059] A determination module is configured to determine sound source localization information based on the high-resolution beam pattern.

[0060] An embodiment of the present invention provides an electronic device, comprising: a memory for storing executable instructions; and a processor for implementing the above-mentioned deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization when executing the executable instructions stored in the memory.

[0061] An embodiment of the present invention provides a computer-readable storage medium storing executable instructions for causing a processor to execute the executable instructions to implement the above-mentioned deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization.

[0062] The present invention provides a deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization, which obtains the sound pressure signal of the sound source through a microphone array and calculates a low-resolution beam pattern according to a preset delay-sum algorithm; based on the sound source propagation model, the low-resolution beam pattern is rewritten, and the generation mechanism of the rewritten low-resolution beam pattern and the influence of the point spread function PSF are analyzed to construct a deconvolution model; L1-L2 regularization is added to the deconvolution model as a sparse constraint to construct a non-convex L1-L2 sparse regularized deconvolution high-resolution localization model; based on the proximal algorithm and gradient descent algorithm of L1-L2 regularization, the non-convex L1-L2 sparse regularized deconvolution high-resolution localization model is iteratively solved to obtain a high-resolution beam pattern to determine the sound source localization information. In this way, the present invention can, on the one hand, more effectively approximate the description of the L0 norm by constraining the non-convex regularization of the 1-norm and the 2-norm difference near the zero point, which is conducive to the spatial sparsity expression of the sound source distribution and can achieve higher-resolution sound source localization. Secondly, considering the non-negative nature of the sound source's acoustic power, a non-negative L1-L2 regularization was chosen. Thirdly, utilizing the analytical solution of the L1-L2 proximal operator to iteratively estimate the location and energy information of the sound source enables efficient and reliable calculations, improving computational efficiency and energy concentration. Fourthly, the present invention utilizes the advantages of a geometric structure closer to the L0 norm, effectively improving the spatial resolution, robustness, and positioning accuracy of sound source localization. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 Schematic diagram of a flow chart of a deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization provided by the present invention;

[0064] Figure 2 This is a schematic diagram of an anechoic chamber experimental device and layout provided by an embodiment of the present invention;

[0065] Figure 3 is a conventional beamforming diagram provided by the present invention;

[0066] Figure 4 is a regularization parameter selection curve diagram under different signal-to-noise ratios provided by the present invention;

[0067] Figure 5 It is the high-resolution beam pattern provided by the present invention;

[0068] Figure 6 Schematic diagram of the structure of a deconvolution high-resolution sound source localization device based on non-convex L1-L2 sparse regularization provided by the present invention;

[0069] Figure 7 It is a schematic diagram of the composition structure of the deconvolution high-resolution sound source localization device based on non-convex L1-L2 sparse regularization provided by the present invention. DETAILED DESCRIPTION

[0070] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limiting the present invention. All other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0071] In the following description, references to "some embodiments" describe a subset of all possible embodiments. However, it is understood that "some embodiments" may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict. Unless otherwise defined, all technical and scientific terms used in the embodiments of the present invention have the same meaning as commonly understood by those skilled in the art to which the embodiments of the present invention pertain. The terms used in the embodiments of the present invention are for the purpose of describing the embodiments of the present invention only and are not intended to limit the present invention.

[0072] This paper proposes a high-resolution sound source localization method based on deconvolution of a non-convex L1-L2 sparse regularization. By leveraging its geometric structure, which is closer to the L0 norm, it effectively improves the spatial resolution and robustness of sound source localization. This method is effective and superior in terms of positioning accuracy, energy conservation, and computational cost, and therefore has great potential for achieving higher-resolution sound source localization.

[0073] The embodiment of the present invention provides a deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization, see Figure 1 , Figure 1 This is a flowchart of a deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization provided by an embodiment of the present invention, which is combined with Figure 1 The steps shown are explained.

[0074] Step S1: Acquire a sound pressure signal of a sound source using a microphone array having M elements, and calculate a low-resolution beam pattern based on a preset delay-sum algorithm.

[0075] In this embodiment, in step S1, Figure 2 This is a schematic diagram of the anechoic chamber experimental setup and layout. The experimental system mainly consists of two sound sources and a microphone array. The location of the sound source is restored by the obtained microphone signal. The microphone array consists of 16 microphones, forming a rectangular array with an array element spacing of 0.06m. The microphone model used is 46AM, an analog sound transducer from CETC. Two speakers (HUAWEI AI Speaker 2) are placed 3m away from the center of the microphone array. The sampling frequency is 51200Hz. The two sound sources play the music "Hundred Birds Paying Homage to the Phoenix" with a main frequency of 1820Hz and a sound source spacing of 1.6m.

[0076] In this embodiment, it is assumed that there are unknown sound sources in the free field. On a grid plane with n focal points (n=N×N), the coordinates of the focal point s are r s =[x s ,y s ,z s There are M microphones arranged in the sound field, located on the xoy plane, parallel to the grid plane, and the coordinate of microphone m is r m =[x m ,y m ,z m ].

[0077] In some embodiments, the preset delay summation algorithm refers to receiving a sound signal through a microphone array and calculating the time delay of each microphone receiving the signal, thereby determining a low-resolution beam pattern of the sound source.

[0078] In some embodiments, the sound source localization information includes position information and intensity information of the sound source.

[0079] In step S1 of the present invention, the signal received by the microphone array is collected by the right-angle sensor array. The sound source plays the music "Hundred Birds Paying Homage to the Phoenix". The distance from the center of the microphone array is 3m, and the distance between the sound sources is 1.6m. The result of traditional beamforming is obtained as shown in formula (1): Figure 3 This is the low-resolution beamforming diagram obtained:

[0080]

[0081] Where p is the sound pressure signal, M is the number of array elements, r is the distance from the focus point to the center of the microphone array, and p H is the conjugate transpose of p, represents the energy of the sound pressure signal, C is the cross-spectral matrix, v(r)=[v1(r),v2(r),…,v M (r)] represents the steering vector of the focus point r on the focusing surface, and the mth element in the steering vector is v(r) H is the conjugate transpose of v(r), (f) H represents the conjugate transpose, Indicates summed average.

[0082] Step S2: rewrite the low-resolution beam pattern based on the sound source propagation model, analyze the generation mechanism of the rewritten low-resolution beam pattern and the influence of the point spread function (PSF), and construct a deconvolution model.

[0083] The sound source propagation model refers to a relationship model between the sound source signal and the sound pressure signal.

[0084] In some embodiments, the sound source propagation model means that the sound pressure signal received by the microphone array is equal to the product of the transfer matrix G between the focusing plane and the measurement plane and the potential sound source intensity q of the focusing plane, as shown in formula (2):

[0085] p=Gq (2),

[0086] Where, source signal q=[q1,q2,...,q s ] T , the sound source intensity of the sth focal point is q s =(jωρQ s / 4π|r s |), the output signals of M microphones p=[p1,p2,...,p M ] T .

[0087] The expression of the transfer matrix G is shown in formula (3):

[0088]

[0089] The low-resolution beam pattern is rewritten using the sound pressure signal of the measurement surface obtained by formula (1) to obtain the rewritten low-resolution beam pattern, as shown in formula (4):

[0090]

[0091] Where r is the distance from the focus point to the center of the microphone array, r s is the distance from the sth focal point to the center of the microphone array, M is the number of array elements, S is the number of sound sources, k is the wave number, and the sound source intensity is q = [q1,q2,...,q s ] T ,qs is the sound source intensity q at the sth gathering point s =(jωρQ s / 4π|r s |), v(r) H is the conjugate transpose of v(r), Green's function is the Green's function corresponding to the s-th focusing point, g rs The conjugate transpose of .

[0092] According to the propagation model of array sound pressure, the generation mechanism of the beam pattern is analyzed. The rewritten low-resolution beam pattern model is decomposed into two parts. The first part is the PSF, which is defined as the response of the beamformer to a single sound source. The expression is shown in formula (5):

[0093]

[0094] The second part is the sound source signal. In the case of the translation-invariant PSF, the rewritten low-resolution beam pattern can be rewritten as:

[0095]

[0096] The frequency domain convolution is converted into beam domain product through discrete spatial Fourier transform for acceleration, and a deconvolution model is constructed. The deconvolution model expression is shown in formula (7):

[0097]

[0098] Where r is the distance from the focus point to the center of the microphone array, M is the number of array elements, and v(r) H is the conjugate transpose of v(r), is the Green's function corresponding to the s-th focusing point, for The conjugate transpose of * represents the convolution operation, X represents the “clean” beamforming pattern, B represents the matrix composed of the traditional beamforming pattern, P represents the PSF of the focal plane, o represents the Hadamard product, and represent two-dimensional Fourier transform and inverse Fourier transform, respectively.

[0099] In the present invention, in order to avoid errors in calculation, zero padding can be used before Fourier transform. Therefore, here, the N×N grid is expanded to at least (2N-1)×(2N-1) size by zero padding, where the number of grids N=300.

[0100] Step S3: Add L1-L2 regularization as a sparse constraint to the deconvolution model to construct a non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model.

[0101] In step S3 of the present invention, the deconvolution model is rewritten as an optimization solver, and L1-L2 regularization is added to the deconvolution model as a sparse constraint. Considering that the beamforming power should be non-negative, a non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model is obtained. The model can be expressed as formula (8):

[0102]

[0103] subject to X≥0.

[0104] Where X represents the “clean” beamforming pattern, B represents the matrix of the conventional beamforming pattern, P represents the PSF of the focal plane, and o represents the Hadamard product. and denote the two-dimensional Fourier transform and inverse Fourier transform, respectively, ||·|| F is the Frobenius norm, ||·||1 is the 1-norm, ||·||2 is the 2-norm, λ is the regularization parameter, and α is the sparse weight parameter.

[0105] In the present invention, the values ​​of λ and α are determined according to Figure 4 The optimal parameter selection curve shown in FIG. 1 selects the value when SNR = 10 dB, that is, λ = 60, α = 18. At this time, the performance of the method of the present invention reaches the best.

[0106] Step S4: Based on the L1-L2 regularized proximal algorithm and the gradient descent algorithm, the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model is iteratively solved to obtain a high-resolution beam map.

[0107] In some preferred embodiments, in step S4, in order to solve the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model, the L1-L2 proximal operator and gradient descent step are used to iteratively solve the problem, eliminate the influence of the point spread function PSF from the low-resolution beam map, obtain a high-resolution beam map, and determine the position and intensity of the sound source.

[0108] The first iterative step is as follows, which includes the forward gradient descent and the backward proximal operator:

[0109]

[0110]

[0111] in, is the proximal operator of L1-αL2, X represents the “clean” beamforming map, and the gradient It can be expressed as Need to choose appropriately The Lipschitz constant L satisfies for The Lipschitz constant is the largest eigenvalue λ of the Hessian matrix max , so it can be estimated efficiently using the power method, and the expression for the method is:

[0112]

[0113] Where X1 is an arbitrary matrix with the same shape as the matrix P, the Lipschitz constant L is obtained after 10 iterations, P represents the PSF of the focal plane, and o represents the Hadamard product. and represent two-dimensional Fourier transform and inverse Fourier transform, respectively.

[0114] The proximal operator of the backward L1-αL2 is defined as follows:

[0115]

[0116] The closed-form solution of the L1-αL2 proximal operator in the iterative process is given (positive parameter λ>0):

[0117] Given λ>0, and α≥0, the optimal solution x in the optimization problem of the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model is * There are the following statements:

[0118] When ||y|| ∞ >λ,x * =z(||z||2+αλ) / ||z||2,

[0119] When ||y|| ∞ =λ, if and only if |y i |<λ,||x * ||2=αλ and for all i, When x * There is a unique solution that satisfies When multiple components of y have the largest absolute value, the optimal solution is not unique; that is, there are infinitely many optimal solutions;

[0120] When (1-α)λ<||y|| ∞ <λ, if and only if |y i |<||y|| ∞ ,||x*||2=||y|| ∞ +(α-1)λ and for all i, When x * There is a unique solution, which is a single sparse vector, satisfying The number of optimal solutions and the number of solutions with the largest absolute value ||y|| ∞ The number of components is the same;

[0121] When ||y|| ∞ ≤(1-α)λ, The proximal operator for L1 is called soft contraction and is defined as:

[0122]

[0123] in, is the optimal solution, y is the input of the proximal operator, y i is the i-th element of the y vector, z is the soft threshold operator, i is the number of elements, α is the sparse weight parameter, and λ is the regularization parameter.

[0124] In some embodiments, the overall iterative process in S4 includes:

[0125] Step 1: Update X (k)

[0126]

[0127] Step 2: (k+1)

[0128]

[0129] Step 3: Y (k+1)

[0130]

[0131] Repeat the iteration until ||X (k+1) -X (k) || F ≤εmax(||X (k) || F ,1), obtain the high-resolution beam pattern;

[0132] Where ε is 10 -5 , t (1) is 1, Y (1) is a zero matrix, X (k) is the X value of the kth iteration. The X obtained when the iteration stops is the high-resolution beam pattern. k is the value of t at the kth iteration, t (k+1) is the value of t at the k+1th iteration, where t is the momentum parameter used to accelerate the algorithm calculation time, and Y (k+1) is the value of Y at the k+1th iteration, ||X (k+1) -X (k) || Fis the F norm of the difference between the k+1th and kth iterations, ||X (k) || F is the F norm of the k-th iteration X.

[0133] Step S5: Determine sound source localization information based on the high-resolution beam pattern.

[0134] The present invention provides a deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization, which obtains the sound pressure signal of the sound source through a microphone array and calculates a low-resolution beam pattern according to a preset delay-sum algorithm; based on the sound source propagation model, the low-resolution beam pattern is rewritten, and the generation mechanism of the rewritten low-resolution beam pattern and the influence of the point spread function PSF are analyzed to construct a deconvolution model; L1-L2 regularization is added to the deconvolution model as a sparse constraint to construct a non-convex L1-L2 sparse regularized deconvolution high-resolution localization model; based on the proximal algorithm and gradient descent algorithm of L1-L2 regularization, the non-convex L1-L2 sparse regularized deconvolution high-resolution localization model is iteratively solved to obtain a high-resolution beam pattern to determine the sound source localization information. In this way, the present invention can, on the one hand, more effectively approximate the description of the L0 norm by constraining the non-convex regularization of the 1-norm and the 2-norm difference near the zero point, which is conducive to the spatial sparsity expression of the sound source distribution and can achieve higher-resolution sound source localization. Secondly, considering the non-negative nature of the sound source's acoustic power, a non-negative L1-L2 regularization was chosen. Thirdly, utilizing the analytical solution of the L1-L2 proximal operator to iteratively estimate the location and energy information of the sound source enables efficient and reliable calculations, improving computational efficiency and energy concentration. Fourthly, the present invention utilizes the advantages of a geometric structure closer to the L0 norm, effectively improving the spatial resolution, robustness, and positioning accuracy of sound source localization.

[0135] In the present invention, the position and intensity of the sound source are determined based on the energy distribution of the high-resolution beam pattern X. Figure 5 This is beam pattern X of the present invention (the black asterisk indicates the actual sound source location). Based on the energy distribution, two clear sound sources were located at their actual locations. The distance between the two sound sources was 1.6 meters, and the vertical distance between the plane where the sound sources were located and the plane where the microphone array was located was 3 meters. The error between the actual sound source location and the estimated sound source location was 0.1 meters, and the calculation time of the present invention was less than 0.1 seconds.

[0136] Figure 6 : is a schematic diagram of the composition structure of a deconvolution high-resolution sound source localization device based on non-convex L1-L2 sparse regularization provided by an embodiment of the present invention, such as Figure 6As shown, the deconvolution high-resolution sound source localization device 600 based on non-convex L1-L2 sparse regularization includes: a calculation module 601, which is used to obtain the sound pressure signal of the sound source using a microphone array with an array element number of M, and calculate the low-resolution beam pattern through a preset delay-sum algorithm; a construction module 602, which is used to rewrite the low-resolution beam pattern based on the sound source propagation model, and analyze the generation mechanism of the beam pattern and the influence of the point spread function PSF to construct a deconvolution model; the sound source propagation model refers to the relationship between the sound source signal and the sound pressure signal. system model; the construction module 602 is further used to add L1-L2 regularization as a sparse constraint to the deconvolution model to construct a non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model; the solution module 603 is used to iteratively solve the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model based on the proximal algorithm and gradient descent algorithm of the L1-L2 regularization to obtain a high-resolution beam map; the determination module 604 is used to determine the sound source localization information based on the high-resolution beam map.

[0137] In some embodiments, it should be noted that the description of the device embodiment of the present invention is similar to the description of the method embodiment described above, and has similar beneficial effects as the same method embodiment, so it is not repeated here. For technical details not disclosed in the present device embodiment, please refer to the description of the method embodiment of the present invention for understanding.

[0138] It should be noted that, in the embodiment of the present invention, if the above-mentioned deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization is implemented in the form of a software function module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present invention, or the part that contributes to the relevant technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a terminal to execute all or part of the methods described in each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM, Read Only Memory), a magnetic disk or an optical disk. In this way, the embodiment of the present invention is not limited to any specific combination of hardware and software.

[0139] Correspondingly, the embodiment of the present invention provides a deconvolution high-resolution sound source localization device based on non-convex L1-L2 sparse regularization. Figure 7 : is a schematic diagram of the composition structure of a deconvolution high-resolution sound source localization device based on non-convex L1-L2 sparse regularization provided by an embodiment of the present invention, such as Figure 7As shown, the high-resolution sound source localization device 700 based on deconvolution of non-convex L1-L2 sparse regularization includes at least: a processor 701 and a computer-readable storage medium 702 configured to store executable instructions, wherein the processor 701 generally controls the overall operation of the high-resolution sound source localization device based on deconvolution of non-convex L1-L2 sparse regularization. The computer-readable storage medium 702 is configured to store instructions and applications executable by the processor 701, and can also cache data to be processed or processed by each module of the processor 701 and the high-resolution sound source localization device 700 based on non-convex L1-L2 sparse regularization. This can be implemented using flash memory (FLASH) or random access memory (RAM).

[0140] An embodiment of the present invention provides a storage medium storing executable instructions, wherein the executable instructions are stored. When the executable instructions are executed by a processor, the processor will be caused to execute the method provided by the embodiment of the present invention, for example, Figure 2 The method shown.

[0141] In some embodiments, the storage medium can be a computer-readable storage medium, such as a ferroelectric random access memory (FRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); it can also be various devices including one or any combination of the above memories.

[0142] In some embodiments, executable instructions may be in the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.

[0143] As examples, executable instructions may, but need not necessarily, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as one or more scripts in a Hypertext Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinating files (e.g., files storing one or more modules, subroutines, or code portions). As examples, executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected by a communication network.

[0144] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and scope of the present invention are included in the scope of protection of the present invention.

[0145] It should be understood that "one embodiment" or "an embodiment" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present invention. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention. The serial numbers of the above-mentioned embodiments of the present invention are for description only and do not represent the advantages and disadvantages of the embodiments.

[0146] It should be noted that, in this article, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method or device. In the absence of further restrictions, an element defined by the statement "comprises a..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely schematic. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored, or not executed.

[0147] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A high-resolution sound source localization method based on deconvolution and non-convex L1-L2 sparse regularization, characterized in that: The method comprises: S1: Use a microphone array with M elements to obtain the sound pressure signal of the sound source and calculate a low-resolution beam pattern based on a preset delay-sum algorithm; S2: Based on a sound source propagation model, rewrite the low-resolution beam pattern, analyze the generation mechanism of the rewritten low-resolution beam pattern and the influence of the point spread function (PSF), and construct a deconvolution model; the sound source propagation model refers to a relationship model between the sound source signal and the sound pressure signal; S3: Add L1-L2 regularization as a sparse constraint to the deconvolution model to construct a non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model; S4: Based on the L1-L2 regularized proximal algorithm and the gradient descent algorithm, the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model is iteratively solved to obtain a high-resolution beam map; S5: Determine sound source localization information based on the high-resolution beam pattern.

2. The method according to claim 1, characterized in that The low-resolution beam pattern b(r) in S1 is: ; Where p is the sound pressure signal, M is the number of array elements, r is the distance from the focus point to the center of the microphone array, and p H is the conjugate transpose of p, represents the energy of the sound pressure signal, C is the cross-spectral matrix, Indicates the focus plane The steering vector corresponding to the focus point, the mth element in the steering vector is , k is the wave number, for The conjugate transpose of Indicates summed average.

3. The method according to claim 1, characterized in that The rewritten low-resolution beam pattern in S2 is: ; Where r is the distance from the focus point to the center of the microphone array, is the distance from the sth focal point to the center of the microphone array, Indicates the focus plane The steering vector corresponding to the focus point, for The conjugate transpose of is the Green's function corresponding to the s-th focusing point, for The conjugate transpose of , S is the number of sound sources, M is the number of array elements, and the sound source intensity is , is the sound source intensity at the sth focal point , ρ is the density of the propagation medium, Q s is the sound source intensity at the center of the microphone array, ω is the angular frequency, is a column vector The elements in represent the Green's function corresponding to the s-th focal point of the m-th microphone. is the distance between the sth focal point and the mth microphone, and k is the wave number.

4. The method according to claim 3, characterized in that The point spread function PSF in S2 is defined as the response of the beamformer to a single sound source and is expressed as: ; In the case of a translation-invariant PSF, ; Based on discrete spatial Fourier transform, the frequency domain convolution is converted into beam domain product for acceleration, and the deconvolution model is constructed. The deconvolution model is: , Where r is the distance from the focus point to the center of the microphone array, M is the number of array elements, for The conjugate transpose of is the Green's function corresponding to the s-th focusing point, for The conjugate transpose of represents the convolution operation, represents the "clean" beamforming pattern, represents the matrix of the traditional beamforming pattern, Indicates the focal plane , represents the Hadamard product, and represent two-dimensional Fourier transform and inverse Fourier transform, respectively.

5. The method according to claim 4, characterized in that The non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model in S3 is: ; in, represents the "clean" beamforming pattern, represents the matrix of the traditional beamforming pattern, Represents the focal plane , represents the Hadamard product, and denote the two-dimensional Fourier transform and inverse Fourier transform, respectively. is the Frobenius norm, is the 1 norm, is the 2-norm, represents the regularization parameter, Represents the sparse weight parameter.

6. The method according to claim 5, characterized in that The first iteration process in S4 includes: ; ; in, for The proximal operator of represents the "clean" beamforming map, the gradient Expressed as ; In the formula, you need to choose appropriately The Lipschitz constant L satisfies ; The iterative process of calculating the Lipschitz constant is as follows: ; ; in, is an arbitrary matrix with the same shape as the matrix P, and the Lipschitz constant L is obtained after 10 iterations. Represents the focal plane , represents the Hadamard product, and denote two-dimensional Fourier transform and inverse Fourier transform respectively; Backward in the iterative process The proximal operator of is defined as follows: ; Given the iterative process The closed-form solution of the proximal operator, where the regularization parameter : Given , ,and , the optimal solution to the optimization problem in the non-convex L1-L2 sparse regularized deconvolution high-resolution localization model There are the following statements: when , , ; when , if and only if , And for all i, hour, There is a unique solution that satisfies ; When multiple components of y have the largest absolute value, the optimal solution is not unique; that is, there are infinitely many optimal solutions; when , if and only if , And for all i, hour, There is a unique solution, which is a single sparse vector, satisfying ; The number of optimal solutions and the maximum absolute value The number of components is the same; when , , The proximal operator for L1 is called soft contraction and is defined as: ; in, is the optimal solution, y is the input of the proximal operator, is the i-th element of the y vector, z is the soft threshold operator, i is the number of elements, is the sparse weight parameter, is the regularization parameter.

7. The method according to claim 6, characterized in that The overall iterative process in S4 includes: Step 1: Update ; ; Step 2: ; ; Step 3: ; ; Repeat the iteration until , obtaining the high-resolution beam pattern; in, is 10 -5 , t (1) is 1, Y (1) is a zero matrix, is the X value of the kth iteration. The X obtained when the iteration stops is the high-resolution beam pattern. is the value of t at the kth iteration, is the value of t at the k+1th iteration, where t is the momentum parameter used to accelerate the algorithm calculation time. is the value of Y at the k+1th iteration, is the Frobenius norm of the difference between the k+1th and kth iterations, is the Frobenius norm of X at the kth iteration.

8. A high-resolution sound source localization device based on non-convex L1-L2 sparse regularization and deconvolution, characterized in that: The device comprises: A calculation module is used to obtain the sound pressure signal of the sound source using a microphone array with M elements, and calculate the low-resolution beam pattern through a preset delay-sum algorithm; A construction module is configured to rewrite the low-resolution beam pattern based on a sound source propagation model, analyze the generation mechanism of the beam pattern and the influence of the point spread function (PSF), and construct a deconvolution model; the sound source propagation model is a relationship model between the sound source signal and the sound pressure signal; The construction module is further used to add L1-L2 regularization as a sparse constraint to the deconvolution model to construct a non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model; A solution module, configured to iteratively solve the non-convex L1-L2 sparse regularized deconvolution high-resolution positioning model based on the L1-L2 regularized proximal algorithm and the gradient descent algorithm to obtain a high-resolution beam map; A determination module is configured to determine sound source localization information based on the high-resolution beam pattern.

9. An electronic device, characterized in that: include: a memory for storing executable instructions; A processor, configured to implement the deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization as described in any one of claims 1 to 7 when executing the executable instructions stored in the memory.

10. A computer-readable storage medium storing executable instructions for causing a processor to execute the executable instructions to implement the deconvolution high-resolution sound source localization method based on non-convex L1-L2 sparse regularization according to any one of claims 1 to 7.