A multi-sound source positioning method, system and electronic device

By processing data from a multi-channel microphone array and utilizing a high-order cumulant matrix and the spherical harmonic domain MUSIC algorithm, the problems of decreased localization performance and noise robustness of the spherical harmonic domain MUSIC algorithm in multi-sound source scenarios were solved, achieving higher accuracy in sound source location estimation and better multi-sound source localization.

CN119902161BActive Publication Date: 2025-11-21ANHUI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510077798.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-11-21
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

Existing spherical harmonic domain MUSIC algorithms exhibit decreased localization performance in complex scenarios with multiple sound sources in enclosed indoor environments and lack robustness to spatial noise.

Method used

By acquiring time-domain sound pressure data from a multi-channel microphone array, performing short-time Fourier transform and spherical Fourier transform, a spherical harmonic domain sound pressure coefficient model is constructed. Singular value decomposition is then performed using a high-order cumulant matrix to construct a MUSIC spatial spectrum for sound source localization.

Benefits of technology

It improves the algorithm's robustness to noise, enhances the accuracy of sound source location estimation, overcomes the limitations of traditional algorithms when the number of sound sources increases, and improves the overall effect of multi-sound source localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902161B_ABST
    Figure CN119902161B_ABST
Patent Text Reader

Abstract

The application provides a multi-sound source positioning method, system and electronic equipment, and relates to the technical field of sound source positioning. The multi-sound source positioning method comprises the following steps: acquiring time-domain sound pressure collection data of a multi-channel microphone array; performing short-time Fourier transform on the time-domain sound pressure collection data to obtain a frequency-domain sound pressure signal model; performing spherical Fourier transform on the frequency-domain sound pressure signal model to obtain a spherical harmonic domain sound pressure coefficient model; processing the spherical harmonic domain sound pressure coefficient model to obtain a high-order cumulant matrix corresponding to a spherical harmonic domain received signal; and performing singular value decomposition on the high-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain a noise subspace. The method, system and electronic equipment provided by the application expand the dimension of a steering vector matrix through a spherical harmonic domain MUSIC algorithm of a fourth-order cumulant, thereby improving the effective rank of the steering vector matrix and increasing the number of target sound sources that can be positioned.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sound source positioning, in particular to a multi-sound source positioning method, system and electronic device. BACKGROUND

[0002] Sound source positioning technology based on microphone sensor array signal processing is an important acoustics signal processing work. In recent years, it has been applied to spatial filtering, sound source separation, sound source tracking, dereverberation and speech enhancement fields. Among them, the spherical microphone array is widely used in three-dimensional sound field reconstruction, room acoustic analysis, closed cabin sound source positioning and other three-dimensional sound field environments. The spherical microphone array can not only perform sound source positioning in the array element domain, but also can realize sound source positioning in the modal domain. The spherical harmonic domain sound source positioning technology converts the array element domain data obtained by the spherical array into the spherical harmonic domain (SHD) by using the spherical Fourier transform, and then positions the sound source. Compared with the traditional microphone array array element domain processing method, the spherical harmonic domain processing method can decouple the frequency and direction information of the sound source, so that the sound source positioning in the three-dimensional sound field is easier. Therefore, most of the current spherical array sound source positioning methods are processed by the spherical harmonic domain. Most of the modal beamforming of the spherical array needs to be converted to the frequency domain and then processed in the spherical harmonic domain. Yan et al. proposed a real-time implementation method of wideband beamforming of spherical microphone array, and formulated a multi-constraint problem to solve the weight coefficients of the filter to provide appropriate trade-off between multiple conflicting performance indicators. Li et al. combined the traditional MUSIC algorithm with the spherical harmonic domain and proposed a spherical harmonic domain MUSIC algorithm (SH-MUSIC), which successfully applied the subspace method to the spherical array for sound source positioning.

[0003] However, most of the commonly used spherical harmonic domain sound source positioning algorithms are extended from the array element domain microphone array positioning algorithm to the spherical microphone array. When these algorithms are used for sound source positioning in high-noise and high-reverberation environments such as indoor and closed cabins, the spatial resolution and sound source positioning accuracy will be seriously reduced, and they are not robust to spatial noise. SUMMARY

[0004] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a multi-sound source positioning method, system and electronic device, which solves the problem that the conventional spherical harmonic domain MUSIC algorithm in the prior art has reduced positioning performance in a closed indoor multi-sound source complex scene, and is not robust to spatial noise.

[0005] To achieve the above object and other related objects, the present application provides a multi-sound source positioning method, comprising: acquiring time-domain sound pressure collection data of a multi-channel microphone array; performing short-time Fourier transform on the time-domain sound pressure collection data to obtain a frequency-domain sound pressure signal model; performing spherical Fourier transform on the frequency-domain sound pressure signal model to obtain a spherical harmonic domain sound pressure coefficient model; processing the spherical harmonic domain sound pressure coefficient model to obtain a high-order cumulant matrix corresponding to a spherical harmonic domain received signal; performing singular value decomposition on the high-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain a noise subspace; constructing a MUSIC spatial spectrum according to the orthogonality relationship between the corresponding extended steering vector of the high-order cumulant matrix corresponding to the spherical harmonic domain received signal and the noise subspace; and performing grid scanning search on the MUSIC spatial spectrum to obtain a peak position in the MUSIC spatial spectrum as a sound source position.

[0006] In an embodiment of the present application, the acquiring of the time-domain sound pressure collection data of the multi-channel microphone array comprises: obtaining time-domain sound pressure collection data of a plane wave on a spherical surface according to the spatial sampling distribution of corresponding microphone elements in the multi-channel microphone spherical array and the direction of arrival of the plane wave.

[0007] In an embodiment of the present application, the performing of short-time Fourier transform on the time-domain sound pressure collection data to obtain the frequency-domain sound pressure signal model comprises: obtaining a representation model of the time-domain sound pressure collection data in the frequency domain according to the time-domain sound pressure collection data; and converting the representation model into the frequency-domain sound pressure signal model according to the corresponding microphone elements in the multi-channel microphone spherical array which carries noise.

[0008] In an embodiment of the present application, the performing of spherical Fourier transform on the frequency-domain sound pressure signal model to obtain the spherical harmonic domain sound pressure coefficient model comprises: obtaining a deformation matrix of a steering vector matrix according to the steering vector matrix, the corresponding spherical mode intensity and the angular position of the microphone array on the spherical array in the frequency-domain sound pressure signal model; performing spherical Fourier transform on the received sound pressure of the multi-channel microphone array to obtain a spherical harmonic domain receiving matrix; and obtaining the spherical harmonic domain sound pressure coefficient model for sound source positioning according to the spherical harmonic domain receiving matrix, the sampling weight condition corresponding to the spherical harmonic domain receiving matrix and the deformation matrix.

[0009] In an embodiment of the present application, the processing of the spherical harmonic domain sound pressure coefficient model to obtain the high-order cumulant matrix corresponding to the spherical harmonic domain received signal comprises: eliminating an influence factor in the spherical harmonic domain sound pressure coefficient model through plane wave decomposition to obtain a spherical harmonic domain data model; wherein the influence factor comprises a spherical array mode intensity; and performing steering vector matrix expansion on the spherical harmonic domain data model to obtain the high-order cumulant matrix corresponding to the spherical harmonic domain received signal.

[0010] In an embodiment of the present application, the spherical harmonic domain data is subjected to steering vector matrix expansion to obtain a high-order cumulant matrix corresponding to the spherical harmonic domain received signal, comprising: obtaining a high-order cumulant matrix of a spherical harmonic domain signal source and a high-order cumulant matrix of a spherical harmonic domain noise respectively according to a spherical harmonic domain data model; and processing the spherical harmonic domain data model according to the high-order cumulant matrix of the spherical harmonic domain signal source and the high-order cumulant matrix of the spherical harmonic domain noise to obtain the high-order cumulant matrix corresponding to the spherical harmonic domain received signal.

[0011] In an embodiment of the present application, the high-order cumulant matrix is a fourth-order cumulant matrix.

[0012] In an embodiment of the present application, the high-order cumulant matrix corresponding to the spherical harmonic domain received signal is subjected to singular value decomposition to obtain a noise subspace, comprising: performing eigenvalue decomposition on the high-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain eigenvalues and corresponding eigenvectors; classifying the eigenvalues to obtain large eigenvalues and small eigenvalues; and using a space spanned by the eigenvectors corresponding to the small eigenvalues as the noise subspace.

[0013] To achieve the above object and other related objects, the present application further provides a multi-sound source positioning system, comprising: an acquisition unit configured to acquire time-domain sound pressure acquisition data of a multi-channel microphone array; a short-time Fourier transform unit configured to perform short-time Fourier transform on the time-domain sound pressure acquisition data to obtain a frequency-domain sound pressure signal model; a spherical Fourier transform unit configured to perform spherical Fourier transform on the frequency-domain sound pressure signal model to obtain a spherical harmonic domain sound pressure coefficient model; a processing unit configured to process the spherical harmonic domain sound pressure coefficient model to obtain a high-order cumulant matrix corresponding to a spherical harmonic domain received signal; a decomposition unit configured to perform singular value decomposition on the high-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain a noise subspace; a spatial spectrum construction unit configured to construct a MUSIC spatial spectrum according to an orthogonality relationship between a corresponding expanded steering vector of the high-order cumulant matrix corresponding to the spherical harmonic domain received signal and the noise subspace; and a scanning unit configured to perform grid scanning search on the MUSIC spatial spectrum to obtain a peak position in the MUSIC spatial spectrum as a sound source position.

[0014] To achieve the above object and other related objects, the present application further provides an electronic device, comprising: one or more processors; and a storage device configured to store one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the multi-sound source positioning method described above.

[0015] As described above, the multi-source localization method, system, and electronic device of the present invention have the following beneficial effects: By using higher-order cumulants to preprocess the spherical harmonic domain signal, the higher-order cumulants can effectively suppress Gaussian measurement noise, improving the algorithm's robustness to noise. Furthermore, the spherical harmonic domain MUSIC algorithm with higher-order cumulants proposed in this invention also increases the effective rank of the steering vector matrix by expanding its dimension, thereby increasing the number of locatable target sound sources. In complex environments, this method effectively enhances the accuracy of sound source location estimation, overcomes the limitations of traditional algorithms when the number of sound sources increases, and improves the overall effect of multi-source localization. Attached Figure Description

[0016] Figure 1 This is a flowchart illustrating the multi-sound source localization method provided in an embodiment of the present invention.

[0017] Figure 2 The diagram shows the rank and effective rank of the steering vector matrix at different highest modal orders, as provided in an embodiment of the present invention.

[0018] Figure 3 The diagram shows the influence of modal order and source spacing angle on the effective rank, as provided in an embodiment of the present invention.

[0019] Figures 4 to 5 The image shown is a sound source localization diagram for different algorithms when the number of sound sources is less than the spherical harmonic coefficient, as provided in an embodiment of the present invention.

[0020] Figures 6 to 9 The image shown is a sound source localization diagram for different algorithms when the number of sound sources is greater than or equal to the spherical harmonic coefficient, as provided in an embodiment of the present invention.

[0021] Figure 10 The diagram shown is a structural block diagram of a multi-source localization system provided in an embodiment of the present invention.

[0022] Figure 11 The diagram shown is a structural schematic of an electronic device according to an embodiment of the present invention.

[0023] Component designation explanation

[0024] Electronic device 1; multi-source localization system 11; memory 12; processor 13; acquisition unit 111; short-time Fourier transform unit 112; spherical Fourier transform unit 113; processing unit 114; decomposition unit 115; spatial spectrum construction unit 116; scanning unit 117. Detailed Implementation

[0025] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.

[0026] It should be noted that the illustrations provided in the following embodiments are only schematic representations of the basic concept of the present invention. Therefore, the drawings only show the components related to the present invention and are not drawn according to the actual number, shape and size of the components in the actual implementation. In the actual implementation, the form, quantity and proportion of each component can be arbitrarily changed, and the layout of the components may also be more complex.

[0027] In the following description, numerous details are explored to provide a more thorough explanation of embodiments of the invention. However, it will be apparent to those skilled in the art that embodiments of the invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring embodiments of the invention.

[0028] Please see Figure 1 This invention provides a multi-sound source localization method, comprising:

[0029] Step S10: Acquire time-domain sound pressure level data of the multi-channel microphone array;

[0030] Step S20: Perform a short-time Fourier transform on the time-domain sound pressure acquisition data to obtain the frequency-domain sound pressure signal model;

[0031] Step S30: Perform a spherical Fourier transform on the frequency domain sound pressure signal model to obtain the spherical harmonic domain sound pressure coefficient model;

[0032] Step S40: Process the spherical harmonic domain sound pressure coefficient model to obtain the high-order cumulant matrix corresponding to the spherical harmonic domain received signal;

[0033] Step S50: Perform singular value decomposition on the high-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain the noise subspace;

[0034] Step S60: Construct the MUSIC spatial spectrum based on the orthogonality relationship between the extended steering vector and the noise subspace of the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal;

[0035] Step S70: Perform a grid scan search on the MUSIC spatial spectrum to obtain the peak positions in the MUSIC spatial spectrum, which will be used as the sound source positions.

[0036] The above steps clearly demonstrate that in the process of multi-source localization using the spherical harmonic domain (SHN) MUSIC algorithm, the time-domain sound pressure level (SPL) data from a multi-channel microphone array is first acquired. Then, after a short-time Fourier transform (SFT), a spherical Fourier transform (SFT) is performed to obtain the SHN model. This model is then processed to decompose and eliminate the influence of the spherical array modal intensities on the SHN model. Furthermore, by extending the steering vector matrix, the higher-order cumulant matrix corresponding to the received signal in the SHN is calculated. Singular value decomposition (SVD) is then performed on the higher-order cumulant matrix to derive the signal and noise subspaces. Finally, utilizing the orthogonality between the extended steering vector and noise subspaces, a MUSIC spatial spectrum is constructed. By performing a grid scan search on the MUSIC spatial spectrum, the location information corresponding to the peak values ​​obtained from the scan is the source location. The steering vector matrix with the addition of higher-order cumulants has a higher effective rank, which enables the localization of more target sound sources. This effectively enhances the accuracy of sound source location estimation, overcomes the limitations of traditional algorithms when the number of sound sources increases, and improves the overall effect of multi-source localization.

[0037] Figure 1 A flowchart illustrating a multi-source localization method in an exemplary embodiment of this application is shown, including steps S10-S70. The following will be combined with… Figure 1 The technical solution of this application will be described in detail below.

[0038] First, execute step S10 to acquire time-domain sound pressure data of the multi-channel microphone array.

[0039] It is worth noting that after acquiring the time-domain sound pressure data of the multi-channel microphone array, further preprocessing of the time-domain sound pressure data is required, followed by short-time Fourier transform of the preprocessed time-domain sound pressure data.

[0040] In step S10, the time-domain sound pressure level acquisition data of the multi-channel microphone array is obtained, including:

[0041] Based on the spatial sampling distribution of the corresponding microphone elements in the multi-channel microphone spherical array and the direction of arrival of the harmonic plane wave, the time-domain sound pressure acquisition data of the plane wave on the sphere is obtained.

[0042] In this embodiment, it is assumed that there are S far-field sound sources in a three-dimensional free sound field, and these S far-field sound sources generate harmonic plane waves with a wave number of k that radiate onto a spherical array. Among them, the spherical microphone consists of I microphone elements distributed according to a certain spatial sampling. The position of the i-th microphone element on the spherical array is r , , , imφ , , ,

[0048] ,

[0050] ,

[0046] , ,

[0049] , ,

[0045] , , , ,

[0044] , m , , ,

[0047] , =(r i cosφ i sinθ i , r i sinφ i sinθ i , r i cosθ i ) T , φ i is the azimuth angle, θ i is the elevation angle, and the direction of arrival of L harmonic plane waves is (θ l , φ l )(l = 1, …, L). The sound pressure expression of the plane wave at the spherical surface r = (r, θ, φ) is expressed by the sum of spherical harmonic functions and spherical Bessel functions as follows:

[0043]

[0044] In an actual situation, a maximum spherical harmonic order N is selected to approximate the infinite summation in Equation (1) by a finite summation, and the approximate representation of the time-domain sound pressure acquisition data is obtained as:

[0045]

[0046] Among them, N is the maximum modal order, and the product of the beam k and the radius r satisfies kr < N, j m (kr) is the spherical Bessel function of the first kind, represents the spherical harmonic function of the n-th order and m-th degree, and its expression is as follows:

[0047]

[0048] In the formula represents the associated Legendre function; i represents the imaginary unit; e imφ represents the complex exponential function, which provides the periodicity of the azimuth angle φ;<00​​​​​​​​

[0051] In step S20, a short-time Fourier transform is performed on the time-domain sound pressure acquisition data to obtain a frequency-domain sound pressure signal model, including:

[0052] Step S201: Based on the time-domain sound pressure acquisition data, obtain the representation model of the time-domain sound pressure acquisition data in the frequency domain;

[0053] Step S202: Based on the corresponding microphone elements in the multi-channel microphone spherical array carrying its own noise, convert the representation model into a frequency domain sound pressure signal model.

[0054] In this embodiment, the frequency domain representation of the sound pressure of the i-th microphone element on the spherical microphone array, which receives L far-field harmonic plane waves in three-dimensional space, is as follows:

[0055]

[0056] Where S l (k)(l=1,…,L) represents the amplitude of L harmonic plane waves, a i (r a ,Ψ l ) represents the steering vector from the l-th sound source signal to the i-th microphone element in the spherical microphone array, r a Let be the radius of the spherical microphone array. The steering vector contains the incident angle of the sound source and the relative positional relationship between the microphone array elements, as well as the spatial characteristics of the microphone array. Considering the noise carried by the microphone array elements themselves, the model represented by equation (4) can be further expressed in a compact form as a frequency domain sound pressure signal model:

[0057] P(k)=A(r a ,Ψ)S(k)+V(k) (5)

[0058] Where A(r) a Ψ) is represented as an I×L dimensional guiding vector matrix, and S(k) is an L×N matrix. S A one-dimensional signal data matrix, V(k) is an I×N matrix. S An uncorrelated sensor noise data matrix of dimension 1. Assume the noise components are spatial white noise with zero mean and Gaussian distribution, and are uncorrelated with the signal S(k). The steering vector matrix is ​​represented as:

[0059] A(r a ,Ψ)=[a(r a ,Ψ1) a(r a ,Ψ2) … a(r a ,Ψ L (6)

[0060] Where a(r) a ,Ψl )=[a1(r a ,Ψ1)a2(r a ,Ψ2)…a I (r a ,Ψ L )],a(r a The i-th term in Ψ1) represents the sound pressure generated by the l-th planar sound source at the i-th microphone, which can be expressed as:

[0061]

[0062] Where, Φ i b represents the angular position of the microphone array element on the spherical array. n (k,r a Let be the intensity of the nth-order spherical mode, and its expression is:

[0063]

[0064] Where, j n (kr) and j′ n (kr a ) are the first kind of spherical Bessel functions and their derivatives, h n (kr) and h′ n (kr a These are the second-kind spherical Bessel functions and their derivatives. The radius of the spherical array is r. a In general, r = r a .

[0065] Substitute equation (7) into a(r) a ,Ψ l The guidance matrix is ​​rewritten as:

[0066] A(r a ,Ψ)=Y(Φ)[B(kr a )y H (Ψ1)…B(kr a )y H (Ψ L (9)

[0067] Where Y(Φ) is (N+1) 2 A dimensional matrix, whose i-th row vector matrix is

[0068] Next, step S30 is executed: a spherical Fourier transform is performed on the frequency domain sound pressure signal model to obtain the spherical harmonic domain sound pressure coefficient model.

[0069] In step S30, a spherical Fourier transform is performed on the frequency domain sound pressure signal model to obtain the spherical harmonic domain sound pressure coefficient model, including:

[0070] Step S301: Based on the steering vector matrix in the frequency domain sound pressure signal model, the corresponding spherical mode intensity, and the angular position of the microphone array on the spherical array, obtain the deformation matrix of the steering vector matrix;

[0071] Step S302: Perform a spherical Fourier transform on the received sound pressure of the multi-channel microphone array to obtain the spherical harmonic domain receiving matrix;

[0072] Step S303: Based on the spherical harmonic domain receiver matrix, the sampling weight conditions corresponding to the spherical harmonic domain receiver matrix, and the deformation matrix, obtain the spherical harmonic domain sound pressure coefficient model for sound source localization.

[0073] In this embodiment, based on the frequency domain sound pressure signal model P(k)=A(r) a ,Ψ)S(k)+V(k), to obtain the transformed matrix A(r) of the guiding matrix. a ,Ψ)=Y(Φ)[B(kr a )y H (Ψ1)…B(kr a )y H (Ψ L The sound pressure received by the microphone at (r,Φ)=(r,θ,φ) is subjected to a spherical Fourier transform (SFT) (or spherical harmonic decomposition) to obtain the spherical harmonic domain receiver matrix:

[0074]

[0075] In practice, the received sound pressure level is not continuous. It is spatially sampled at the microphone location. Therefore, the spherical Fourier transform (SFT) of the sound pressure level is approximately:

[0076]

[0077] Where n∈[0,N], m∈[-n,n], a i Sample weights for each microphone. This can be written in matrix form as P. nm (k,r)=Y H (Φ)ΓP(k,r,Φ) (12)

[0078] In the formula P nm =[P 00 ,P 1(-1) ,P 10 ,P 11 ,…,P NN ] T And Γ=diag(a1,a2,…,a I Let Y be an I×I dimensional sampling weight matrix, and let the sampling weights satisfy Y H (Φ)ΓY(Φ)=I, where I represents the identity matrix;

[0079] Substituting equation (8) into equation (5), and combining equation (11) and the sampling weight condition, we obtain the spherical harmonic domain sound pressure coefficient model for sound source localization as follows:

[0080] P nm =[B(kr1)y H (Ψ1)…B(kr L )y H (Ψ L )]S(k)+V nm (k) (13)

[0081] P nm =A nm (r,Ψ)S(k)+V nm (k) (14)

[0082] Among them, V nm (k)=Y H (Φ)ΓV(k), A nm (r,Ψ) is the spherical harmonic domain guiding vector matrix, i.e.

[0083] A nm (r,Ψ)=[B(kr1)y H (Ψ1)…B(kr L )y H (Ψ L (15)

[0084] Next, step S40 is executed to process the spherical harmonic domain sound pressure coefficient model to obtain the high-order cumulant matrix corresponding to the spherical harmonic domain received signal.

[0085] In step S40, the spherical harmonic domain sound pressure coefficient model is processed to obtain the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal, including:

[0086] Step S401: Eliminate the influence factors in the spherical harmonic domain sound pressure coefficient model by plane wave decomposition to obtain the spherical harmonic domain data model; wherein, the influence factors include the spherical array mode intensity;

[0087] Step S402: Extend the spherical harmonic domain data model by the guide vector matrix to obtain the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal.

[0088] In this embodiment, the higher-order cumulant matrix can be a fourth-order cumulant matrix. From the aforementioned equations (14) and (15), the spherical harmonic domain data model obtained by further eliminating the influence of the spherical array mode intensity on the spherical harmonic domain sound pressure coefficient through plane wave decomposition is as follows:

[0089]

[0090] in, B(kr) is (N+1) 2 ×(N+1) 2 A 3D modal intensity matrix;

[0091] Generally, the fourth-order cumulant of the received signal satisfies the following equation:

[0092]

[0093] In the formula, cum(·) represents the cumulant, and * represents the complex conjugate; substituting the spherical harmonic domain data model (16) after eliminating the influence of the spherical array mode intensity into the fourth-order cumulant matrix of the above formula (17), it can be further expressed as:

[0094]

[0095]

[0096] In step S402, the spherical harmonic domain data is extended by a steering vector matrix to obtain a higher-order cumulant matrix corresponding to the spherical harmonic domain received signal, including:

[0097] Step S4021: Based on the spherical harmonic domain data model, obtain the higher-order cumulant matrix of the spherical harmonic domain signal source and the higher-order cumulant matrix of the spherical harmonic domain noise, respectively;

[0098] Step S4022: Based on the higher-order cumulant matrix of the spherical harmonic domain signal source and the higher-order cumulant matrix of the spherical harmonic domain noise, process the spherical harmonic domain data model to obtain the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal.

[0099] Preferably, the higher-order cumulant matrix is ​​a fourth-order cumulant matrix.

[0100] In equation (18), Y is the spherical harmonic domain steering vector matrix after eliminating the influence of the spherical array mode intensity, S is the incident signal matrix, and V is the noise matrix. Let H be the Kronecker product, H be the Hermitian transpose, and * be the complex conjugate. To simplify the expression of equation (17), based on the properties of the Kronecker product, we can obtain the following equation:

[0101]

[0102] Among them, C S It is the fourth-order cumulant matrix of the spherical harmonic domain signal source, C V It is the fourth-order cumulant matrix of spherical harmonic domain noise.

[0103] make Then the fourth-order cumulant matrix R corresponding to the spherical harmonic domain received signal can be obtained. nm4 as follows:

[0104]

[0105] Among them, A cum4 It is the expanded guide vector, with a dimension of (N+1). 4 ×(N+1) 4 The dimension of the steering vector in the conventional spherical harmonic domain algorithm based on the cross-power spectrum matrix is ​​(N+1). 2 ×(N+1) 2 This significantly increases the dimension of the steering vector matrix, thus theoretically enabling the processing of more target sound sources.

[0106] Next, step S50 is executed to perform singular value decomposition on the high-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain the noise subspace.

[0107] In step S50, singular value decomposition is performed on the higher-order cumulant matrix corresponding to the spherically harmonic domain received signal to obtain the noise subspace, including:

[0108] Step S501: Perform eigenvalue decomposition on the high-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain eigenvalues ​​and corresponding eigenvectors;

[0109] Step S502: Classify the feature values ​​to obtain large feature values ​​and small feature values;

[0110] Step S503: Use the space spanned by the eigenvectors corresponding to the small eigenvalues ​​as the noise subspace.

[0111] In this embodiment, the higher-order cumulant matrix can be a fourth-order cumulant matrix. For the fourth-order cumulant matrix R... nm4 Perform eigenvalue decomposition and arrange the eigenvalues ​​in descending order as follows: The corresponding feature vectors are respectively And because Then R nm4 There is L 2 Large eigenvalues ​​and I 2 -L 2 Let I represent the number of microphone array elements and L represent the number of sound sources. The space spanned by the eigenvectors corresponding to these smaller eigenvalues ​​is called the noise subspace Ω. n ,Right now

[0112]

[0113] Simultaneously define the signal subspace Ω s for:

[0114]

[0115] Since the signal subspace and the noise subspace are orthogonal, the orthogonality condition can be obtained as follows:

[0116]

[0117] The noise matrix is ​​defined as follows:

[0118]

[0119] Next, step S60 is executed: based on the orthogonality relationship between the extended steering vector and the noise subspace of the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal, the MUSIC spatial spectrum is constructed.

[0120] In this embodiment, the higher-order cumulant matrix can be a fourth-order cumulant matrix. Using the guiding vector matrix A... cum4 and noise subspace U n Based on the orthogonality property, the MUSIC space spectrum output by the spherical harmonic domain MUSIC algorithm based on fourth-order cumulants is as follows:

[0121]

[0122] Next, step S70 is executed to perform a grid scan search on the MUSIC spatial spectrum to obtain the peak positions in the MUSIC spatial spectrum, which are used as the sound source locations. By adding higher-order cumulants, the steering vector matrix has a higher effective rank, thus enabling sound source localization for more target sound sources; this effectively enhances the accuracy of sound source location estimation, overcomes the limitations of traditional algorithms when the number of sound sources increases, and improves the overall effect of multi-source localization.

[0123] In this embodiment, the effective rank of a matrix is ​​defined as follows:

[0124] For matrix A, whose matrix dimension is M×N, performing singular value decomposition yields A = UDV, where matrices U and V are both unitary matrices with dimensions of M×M and N×N, respectively. D is the diagonal matrix composed of eigenvalues ​​after decomposition, with a matrix dimension of M×N, and the eigenvalues ​​are related by the order σ₁ ≥ σ₂ ≥ … ≥ σ₀. Q ≥0, Q=min(M,N). Define the eigenvector σ=(σ1,σ2,..,σ Q ) T The arrangement of eigenvalues ​​can be represented by the following formula:

[0125] p k =σ k / ||σ||1,k=1,2,…,Q (28)

[0126] Where ||·||1 represents the l1 norm, using the parameter p kFind the effective rank of the matrix:

[0127] erank(A) = exp{H(p1,p2,…,p Q )}(29)

[0128] in This is represented as information entropy.

[0129] Generally, the rank of a matrix is ​​a set of integers and is always discrete. However, the effective rank is not necessarily an integer because it is a parameter based on the consistency of the matrix's eigenvalues. It can be a continuous value, and therefore can be used to analyze systems with continuously changing parameters.

[0130] Please see Figure 2 and Figure 3 , Figure 2 and Figure 3 In one embodiment given, Figure 2 The figure shows a comparison of the changes in the rank and effective rank of the steering vector matrix before and after adding the fourth-order cumulant at different highest mode orders in the spherical harmonic domain. As can be seen from the figure, with the increase of the spherical harmonic domain mode order N, the steering vector matrix A before and after the improvement... cum4 Both the rank and effective rank of the improved steering vector matrix A increase when the highest modal order N = 3. cum4 The improved algorithm has achieved full rank in both its rank and effective rank, whereas the original guiding vector matrix A required a maximum modal order of N=5 to achieve full rank. Furthermore, the improved guiding vector matrix has a higher effective rank for the same modal order, thus allowing the improved algorithm to handle more targets at the same modal order. Figure 3 The influence of sound source location at different angular intervals on the effective rank under different modal orders was analyzed. As shown in the figure, under the same modal order and the same sound source angular interval, the effective rank of the improved steering vector matrix is ​​significantly greater than that of the steering vector matrix of the conventional SH-MUSIC algorithm before improvement. Since the steering vector matrices of the SH-MUSIC algorithm and the conventional SH-MUSIC algorithm are identical, the effective value of the improved steering vector matrix is ​​also greater than the effective rank of the steering vector matrix of the conventional SH-MUSIC algorithm. Therefore, the spherical harmonic domain MUSIC algorithm based on fourth-order cumulants, due to the extension of the steering vector matrix, has a higher effective rank than the traditional spherical harmonic domain sound source localization algorithm based on cross-power spectrum, thus improving the accuracy of sound source localization and enabling the array to detect more target sound sources.

[0131] Please see Figures 4 to 9 , Figures 4 to 9 In one embodiment given, Figure 4 The sound source localization map is for the SH-MUSIC algorithm with 3 sources. Figure 5The improved CUM4-SH-MUSIC algorithm and the sound source localization map with 3 sources; Figure 6 The sound source localization map is for the SH-MVDR algorithm with 4 sources. Figure 7 The improved CUM4-SH-MUSIC algorithm and the sound source localization map with 4 sources; Figure 8 The sound source localization map for the SH-MVDR algorithm with 6 sources; Figure 9 The improved CUM4-SH-MUSIC algorithm is used to locate sound sources with a number of 6 sources. Comparative analysis of the SH-MUSIC algorithm and the SH-MVDR algorithm with the improved fourth-order cumulant SH-MUSIC algorithm under different source numbers shows that the simulation results can be divided into two cases: when the number of sound sources is less than the spherical harmonic coefficient and when the number of sound sources is greater than or equal to the spherical harmonic coefficient.

[0132] In the simulation comparison when the number of sound sources is less than the spherical harmonic coefficient, the SH-MUSIC algorithm in the traditional spherical harmonic domain algorithm and the fourth-order cumulant SH-MUSIC algorithm were selected for comparative analysis. The sound sources were set to 3 sources, with incident angles of Ψ1 = (80°, 135°), Ψ2 = (125°, 65°) and Ψ3 = (185°, 130°), respectively. A 50-element rigid spherical microphone array with a radius of 9.75cm was used in the experiment. Figure 4 and Figure 5 The simulation diagrams of normalized sound source localization using the SH-MUSIC algorithm and the fourth-order cumulant SH-MUSIC algorithm are shown respectively when the number of sound sources is less than the spherical harmonic coefficient. The actual sound source locations are represented by black circles "o" and the estimated sound source locations are represented by black plus signs "+".

[0133] When the number of sound sources is greater than or equal to the spherical harmonic coefficients, the SH-MUSIC algorithm fails to obtain the noise subspace during simulation comparison. Therefore, when the number of sound sources is greater than or equal to the spherical harmonic coefficients, the SH-MVDR algorithm in the traditional spherical harmonic domain algorithm is selected for comparative analysis with the fourth-order cumulant SH-MUSIC algorithm. The sound sources are set to 4 sound sources and 6 sound sources respectively. The incident angles of the 4 sound sources are Ψ1=(80°,135°), Ψ2=(125°,65°), Ψ3=(185°,130°) and Ψ4=(260°,65°). The 6 sound sources add two incident angles: Ψ5=(10°,60°) and Ψ6=(95°,10°). Figure 6 and Figure 8 The simulation diagram of normalized sound source localization using the SH-MVDR algorithm is shown when the number of sound sources is greater than or equal to the spherical harmonic coefficient. Figure 7 and Figure 9 The simulation diagram of normalized sound source localization using the fourth-order cumulant SH-MUSIC algorithm is shown when the number of sound sources is greater than or equal to the spherical harmonic coefficient.

[0134] pass Figures 4 to 9 It can be seen that when the number of sound sources is less than the spherical harmonic coefficients, the traditional spherical harmonic domain algorithm can identify and locate all sound sources, but its spatial resolution is significantly worse than that of the CUM4-SH-MUSIC algorithm. However, when the number of sound sources is greater than the spherical harmonic coefficients, the traditional spherical harmonic domain algorithm based on the cross-power spectrum matrix can no longer accurately locate sound sources and cannot detect all sound source targets, while the CUM4-SH-MUSIC algorithm can accurately detect all sound source targets and perform relatively accurate sound source localization. These results demonstrate that the CUM4-SH-MUSIC algorithm can accurately identify and locate sound source targets regardless of whether the number of sound sources is less than or greater than the spherical harmonic coefficients. The CUM4-SH-MUSIC algorithm outperforms the traditional spherical harmonic domain algorithm based on the cross-power spectrum matrix in both localization resolution and the number of targets identified.

[0135] Please refer to 10. This invention also provides a multi-source localization system 11, comprising: an acquisition unit 111 for acquiring time-domain sound pressure data from a multi-channel microphone array; a short-time Fourier transform unit 112 for performing a short-time Fourier transform on the time-domain sound pressure data to obtain a frequency-domain sound pressure signal model; a spherical Fourier transform unit 113 for performing a spherical Fourier transform on the frequency-domain sound pressure signal model to obtain a spherical harmonic domain sound pressure coefficient model; and a processing unit 114 for processing the spherical harmonic domain sound pressure coefficient model to obtain a spherical harmonic domain received signal. The system includes: a higher-order cumulant matrix; a decomposition unit 115, used to perform singular value decomposition on the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain a noise subspace; a spatial spectrum construction unit 116, used to construct a MUSIC spatial spectrum based on the orthogonality relationship between the corresponding extended steering vector of the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal and the noise subspace; and a scanning unit 117, used to perform grid scanning search on the MUSIC spatial spectrum to obtain the peak positions in the MUSIC spatial spectrum as the sound source positions.

[0136] It should be noted that the multi-source localization system 11 provided in the above embodiments and the multi-source localization method provided in the above embodiments belong to the same concept. The specific operation methods of each module and unit have been described in detail in the method embodiments, and will not be repeated here. In practical applications, the multi-source localization system 11 provided in the above embodiments can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. This is not a limitation here.

[0137] Please see Figure 11The electronic device 1 may include a memory 12, a processor 13 and a bus, and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a multi-source localization program.

[0138] The memory 12 includes at least one type of readable storage medium, such as flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 12 can be an internal storage unit of the electronic device 1, such as a portable hard drive. In other embodiments, the memory 12 can be an external storage device of the electronic device 1, such as a plug-in portable hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 1. Furthermore, the memory 12 can include both internal and external storage units of the electronic device 1. The memory 12 can be used not only to store application software and various types of data installed on the electronic device 1, such as multi-source localization code, but also to temporarily store data that has been output or will be output.

[0139] In some embodiments, the processor 13 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits packaged with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 13 is the control unit of the electronic device 1, connecting various components of the electronic device 1 via various interfaces and lines. It executes programs or modules (such as multi-source localization programs) stored in the memory 12, and calls data stored in the memory 12 to perform various functions and process data of the electronic device 1.

[0140] The processor 13 executes the operating system of the electronic device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the multi-source localization method described above.

[0141] For example, the computer program may be divided into one or more modules, which are stored in the memory 12 and executed by the processor 13 to complete this application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which describe the execution process of the computer program in the electronic device 1. For example, the computer program may be divided into an acquisition unit 111, a short-time Fourier transform unit 112, a spherical Fourier transform unit 113, a processing unit 114, a decomposition unit 115, a spatial spectrum construction unit 116, and a scanning unit 117.

[0142] The integrated unit implemented as a software functional module described above can be stored in a computer-readable storage medium, which can be non-volatile or volatile. The software functional module stored in the storage medium includes several instructions to cause a computer device (which may be a personal computer, computer equipment, or network device, etc.) or processor to execute some functions of the multi-source localization method described in the various embodiments of this application.

[0143] In summary, the multi-source localization method, system, and electronic device disclosed in this invention preprocesses the spherical harmonic domain signal using higher-order cumulants. These higher-order cumulants effectively suppress Gaussian measurement noise, improving the algorithm's robustness to noise. Furthermore, the proposed higher-order cumulants-based spherical harmonic domain MUSIC algorithm increases the effective rank of the steering vector matrix by expanding its dimension, thereby increasing the number of locatable target sound sources. In complex environments, this method effectively enhances the accuracy of sound source location estimation, overcomes the limitations of traditional algorithms when the number of sound sources increases, and improves the overall performance of multi-source localization. Therefore, this invention effectively overcomes the various shortcomings of existing technologies and has high industrial application value.

[0144] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.

Claims

1. A multi-source localization method, characterized in that, include: Acquire time-domain sound pressure level data from a multi-channel microphone array; A short-time Fourier transform is performed on the time-domain sound pressure acquisition data to obtain a frequency-domain sound pressure signal model; Perform a spherical Fourier transform on the frequency domain sound pressure signal model to obtain the spherical harmonic domain sound pressure coefficient model; The spherical harmonic domain sound pressure coefficient model is processed to obtain the high-order cumulant matrix corresponding to the spherical harmonic domain received signal; Singular value decomposition is performed on the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain the noise subspace; Based on the orthogonality relationship between the extended steering vector of the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal and the noise subspace, a MUSIC spatial spectrum is constructed. A grid scan search is performed on the MUSIC spatial spectrum to obtain the peak positions in the MUSIC spatial spectrum, which are used as the sound source locations.

2. The multi-source localization method according to claim 1, characterized in that: Acquire time-domain sound pressure level data from a multi-channel microphone array, including: Based on the spatial sampling distribution of the corresponding microphone elements in the multi-channel microphone spherical array and the direction of arrival of the harmonic plane wave, the time-domain sound pressure acquisition data of the plane wave on the sphere is obtained.

3. The multi-source localization method according to claim 1, characterized in that: The time-domain sound pressure acquisition data is subjected to a short-time Fourier transform to obtain a frequency-domain sound pressure signal model, including: Based on the time-domain sound pressure acquisition data, a representation model of the time-domain sound pressure acquisition data in the frequency domain is obtained; Based on the corresponding microphone elements in the multi-channel microphone spherical array carrying its own noise, the representation model is converted into the frequency domain sound pressure signal model.

4. The multi-source localization method according to claim 1, characterized in that: Performing a spherical Fourier transform on the frequency domain sound pressure signal model yields a spherical harmonic domain sound pressure coefficient model, including: Based on the steering vector matrix in the frequency domain sound pressure signal model, the corresponding spherical mode intensity, and the angular position of the microphone array on the spherical array, the deformation matrix of the steering vector matrix is ​​obtained; A spherical Fourier transform is performed on the received sound pressure of the multi-channel microphone array to obtain the spherical harmonic domain receiver matrix. Based on the spherical harmonic domain receiver matrix, the sampling weight conditions corresponding to the spherical harmonic domain receiver matrix, and the deformation matrix, the spherical harmonic domain sound pressure coefficient model for sound source localization is obtained.

5. The multi-source localization method according to claim 1, characterized in that: The spherical harmonic domain sound pressure coefficient model is processed to obtain the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal, including: The influence factors in the spherical harmonic domain sound pressure coefficient model are eliminated by plane wave decomposition to obtain the spherical harmonic domain data model; wherein, the influence factors include the spherical array mode intensity; The spherical harmonic domain data model is extended by a guide vector matrix to obtain the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal.

6. The multi-source localization method according to claim 5, characterized in that: The spherical harmonic domain data is extended by a steering vector matrix to obtain a higher-order cumulant matrix corresponding to the spherical harmonic domain received signal, including: Based on the spherical harmonic domain data model, the higher-order cumulant matrix of the spherical harmonic domain signal source and the higher-order cumulant matrix of the spherical harmonic domain noise are obtained respectively. Based on the higher-order cumulant matrix of the spherical harmonic domain signal source and the higher-order cumulant matrix of the spherical harmonic domain noise, the spherical harmonic domain data model is processed to obtain the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal.

7. The multi-source localization method according to claim 1, characterized in that: The higher-order cumulant matrix is ​​a fourth-order cumulant matrix.

8. The multi-source localization method according to claim 1, characterized in that: Singular value decomposition is performed on the higher-order cumulant matrix corresponding to the spherically harmonic domain received signal to obtain the noise subspace, including: Eigenvalue decomposition is performed on the higher-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain eigenvalues ​​and corresponding eigenvectors. The feature values ​​are classified into large feature values ​​and small feature values; The space spanned by the eigenvectors corresponding to the small eigenvalues ​​is used as the noise subspace.

9. A multi-source localization system, characterized in that, include: The acquisition unit is used to acquire time-domain sound pressure data from a multi-channel microphone array. The short-time Fourier transform unit is used to perform short-time Fourier transform on the time-domain sound pressure acquisition data to obtain a frequency-domain sound pressure signal model. A spherical Fourier transform unit is used to perform a spherical Fourier transform on the frequency domain sound pressure signal model to obtain a spherical harmonic domain sound pressure coefficient model. The processing unit is used to process the spherical harmonic domain sound pressure coefficient model to obtain the high-order cumulant matrix corresponding to the spherical harmonic domain received signal. The decomposition unit is used to perform singular value decomposition on the high-order cumulant matrix corresponding to the spherical harmonic domain received signal to obtain the noise subspace. The spatial spectrum construction unit is used to construct the MUSIC spatial spectrum based on the orthogonality relationship between the corresponding extended steering vector of the higher-order cumulant matrix of the spherical harmonic domain received signal and the noise subspace. as well as The scanning unit is used to perform a grid scan search on the MUSIC spatial spectrum to obtain the peak positions in the MUSIC spatial spectrum as the sound source positions.

10. An electronic device, characterized in that: The electronic device includes: One or more processors; A storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the multi-source localization method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Multi-sound-source locating method based on spherical microphone array

    CN102866385A

  • Near-field acoustic emission source positioning method based on orthogonal matching pursuit under sparse array

    CN117110992A