A multi-sound source localization method using signal time-frequency angle distribution information
By dividing the sub-bands in the frequency domain and performing angle interval screening, the problem of decreased sound source localization accuracy in high reverberation and multi-sound source environments is solved, and a more stable multi-sound source localization effect is achieved.
Patent Information
- Application Number
- CN202211243565.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-28
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-09-28
AI Technical Summary
In highly reverberant and multi-sound source environments, the accuracy of existing multi-sound source localization methods decreases, making it difficult to accurately estimate the direction of arrival of the sound source.
By dividing the recorded signal into sub-bands in the frequency domain, constructing a set of azimuth and elevation angles of time-frequency points, and performing interval division, the angle sub-interval containing the most time-frequency points is screened out, and the direction of arrival of the sound source is obtained using kernel density estimation.
It improves the accuracy of multi-sound source localization and enhances the positioning performance under reverberation conditions, making it suitable for complex scenes.
Smart Images

Figure CN115656927B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of sound source localization in the field of speech signal processing, and in particular relates to the problem of multi-sound source localization in high reverberation scenes. Background Art
[0002] Multiple sound source localization has long been a hot topic in the field of acoustic signal processing. The primary goal of this technology is to estimate the direction of arrival (DOA) of a sound source using only a microphone array. Due to its widespread application, DOA estimation plays a vital role in numerous applications, such as cochlear hearing aids, sound monitoring, localizing dominant noise in machines, and speech enhancement.
[0003] In the study of multiple sound source localization, researchers and institutions both domestically and internationally have proposed numerous methods. The earliest proposed method is based on time difference of arrival (TDOA). Its main idea is to obtain DOA estimates of the sound source signals by calculating the time difference between the sound source arriving at different microphone arrays. Another classic method is multiple signal classification (MUSIC). This method uses the orthogonality of the signal and noise subspaces to obtain a spatial spectrum function, and then uses spectral peak search to obtain DOA estimates of the sound source signals. In addition, sparse component analysis has taken the study of multiple sound source localization to a new level. Its main idea is to assume that in most time-frequency regions, there is always a single source that dominates the others—that is, the energy of this source is much stronger than the others. Therefore, this method transforms the problem of multiple sound source localization into the problem of localizing multiple single sound sources. Considering that there is always a single dominant source in a time-frequency region, researchers in this field have proposed a single source region detection method. The core idea of this method is to determine whether a time-frequency region is a single source region by calculating the correlation between microphone arrays. Finally, the single source is retained to obtain DOA estimates for multiple sound sources. Similarly, there are also multi-source localization methods based on single-source time-frequency point detection, which utilize phase consistency or clustering to perform single-source time-frequency point detection. However, as the recording signal environment becomes increasingly complex, such as with high reverberation, multiple sound sources, and small distances between sound sources, the accuracy of multi-source localization decreases. Summary of the Invention
[0004] The present invention designs a multi-sound source localization method that utilizes the angular distribution information of the signal's time-frequency points in a high-reverberation and multi-sound source environment. This method first divides each frame of speech signal into subbands in the frequency domain, constructing an azimuth set and an elevation set of the time-frequency points within each subband; secondly, the azimuth value range and the elevation value range are divided into intervals, respectively, to establish azimuth subintervals and elevation subintervals; thirdly, the azimuth and elevation angles of the time-frequency points within each subband are divided into corresponding azimuth subintervals and elevation subintervals, and the subinterval with the largest number of time-frequency point azimuth and elevation angles within each subband is selected; finally, a new azimuth set and elevation set are constructed, and the DOA estimation values of the multiple sound sources are obtained through kernel density estimation and peak search.
[0005] The overall design process is briefly described as follows:
[0006] In the first step, a four-channel sound field microphone is used to collect scene data. Recorded signals from multiple sound sources are captured, referred to as A-format signals. The A-format signals are then converted to B-format signals using the A / B-format signal conversion formula. The time-frequency domain coefficients are then obtained using a short-time Fourier transform (SFT). The activity intensity vectors of the B-format signal's time-frequency points are then calculated. The corresponding formulas provide the azimuth and elevation information for each time-frequency point. In the second step, each frame of the recorded signal is divided into subbands in the frequency domain, constructing sets of azimuth and elevation angles for the time-frequency points within each subband. In the third step, the azimuth and elevation angle ranges are partitioned into intervals, respectively, to establish azimuth and elevation subintervals. In the fourth step, the azimuth angles of each subband of the recorded signal's time-frequency points are divided into corresponding azimuth subintervals, constructing sets of azimuth angles for different angle subintervals within each subband. The number of azimuth angles within each set of azimuth subintervals is determined, and the set containing the largest number of azimuth angles is retained. In the fifth step, the elevation angle of each sub-band time-frequency point is divided into the corresponding elevation angle sub-interval, and a set of elevation angles of different angle sub-intervals in each sub-band is constructed. The number of elevation angles in the different elevation angle sub-interval sets is determined, and the set of elevation angle sub-intervals containing the largest number of elevation angles is retained. In the sixth step, the set of azimuth angles retained in each sub-band of each frame is used to construct a set of filtered azimuth angles; at the same time, the set of elevation angles retained in each sub-band of each frame is used to construct a set of filtered elevation angles. Finally, kernel density estimation is performed on the angles in these two sets to obtain the angular distribution curves of azimuth and elevation angles. The angle value corresponding to the peak point in the curve is the estimated direction of arrival of the sound source.
[0007] The technical solution of the present invention is to solve the problem of multi-sound source localization under reverberation conditions, which is mainly divided into the following steps:
[0008] Step 1: Obtain the recorded signal and calculate the azimuth and elevation information of each time-frequency point.
[0009] Step 2: Divide each frame of recorded signal into sub-bands in the frequency domain, and construct an azimuth angle set and an elevation angle set of time-frequency points in each sub-band.
[0010] Step 3: Divide the azimuth angle value range and the elevation angle value range into intervals, and establish azimuth angle sub-intervals and elevation angle sub-intervals.
[0011] Step 4: Divide the azimuth angle of each subband time-frequency point in the recorded signal into corresponding azimuth angle subintervals, and construct azimuth angle sets for different angle subintervals in each subband. Determine the number of azimuth angles in each azimuth angle subinterval set, and retain the azimuth angle subinterval set containing the largest number of azimuth angles.
[0012] Step 5: Divide the azimuth angle of each subband time-frequency point in the recorded signal into corresponding azimuth angle subintervals, constructing azimuth angle sets for different angle subintervals in each subband. Determine the number of azimuth angles in each azimuth angle subinterval set, and retain the azimuth angle subinterval set containing the largest number of azimuth angles.
[0013] In step 6, the set of azimuth angles retained for each subband in each frame is used to construct a set of filtered azimuth angles. Simultaneously, the set of elevation angles retained for each subband in each frame is used to construct a set of filtered elevation angles. Finally, kernel density estimation is performed on the angles in these two sets to obtain azimuth and elevation angle distribution curves. The angle corresponding to the peak point in the curve is the estimated direction of arrival of the sound source.
[0014] 1. The implementation of step 1 is to use a four-channel sound field microphone to collect scene data and collect recording signals from multiple sound sources. The signals collected by the four channels are called M UFL , M DFR , M DBL , M UBR , such signals are recorded as A format signals. B format signals are obtained by linear conversion of A format signals, and the four channel signals of B format signals are recorded as M W , M X , M Y , M Z , where M W Indicates the omnidirectional channel signal, M X , M Y , M Z Represents the channel signals of the three coordinate axes X, Y, and Z of the Cartesian coordinate system. The conversion formula is as follows:
[0015]
[0016] Considering the short-term stability of the sound signal, the recorded signal needs to be framed before positioning processing. The present invention divides the microphone collected signal to be processed into N frames, and the frame length of each processing frame is K, that is, each frame contains K samples.
[0017] After obtaining the B-format signal, the time-frequency domain coefficient M of the time-frequency point (n, k) is obtained by performing a K-point short-time Fourier transform on each frame of the B-format signal. W (n, k), M X (n, k), M Y (n, k), M Z (n, k), where n represents the frame index, n=1, 2...N, and k represents the frequency index, k=1, 2...K. The time-frequency domain coefficients of the B-format signal are calculated using the following formula to obtain the activity intensity vector of the time-frequency point (n, k); T X (n, k), T Y (n, k), T Z (n, k) are the components of the activity intensity vector of the time-frequency point (n, k) in the x, y, and z directions respectively;
[0018]
[0019] Where Re{·} represents the real part of the calculation, (·)* represents the conjugate of the calculation, and the weight coefficient Q is obtained by the following formula:
[0020]
[0021] Where ρ and c represent the density and speed of sound of the medium, respectively. The azimuth angle μ(n, k) and elevation angle η(n, k) of the time-frequency point (n, k) are calculated as follows:
[0022]
[0023]
[0024] 2. Step 2 is implemented by evenly dividing each frame signal into L subbands in the frequency domain. Each subband contains C time-frequency points, where C = K / L. The azimuth angle set of the time-frequency points in the lth subband of the nth frame is constructed by the following formula: and elevation angle set
[0025]
[0026]
[0027] Wherein, k represents a frequency index, n represents a frame index, n=1, 2...N, l represents a subband index, l=1, 2,...,L.
[0028] 3. The implementation of step 3 is as follows: since the range of azimuth angle is [0°, 360°] and the range of elevation angle is [-90°, 90°], the present invention divides the azimuth angle into 18 azimuth angle sub-intervals at intervals of 20°, and records them as follows: In addition, the elevation angle is divided into 18 elevation angle sub-intervals at intervals of 10°.
[0029] 4. Step 4 is implemented by dividing the angle values of the time-frequency points in each sub-band of each frame of the microphone recording signal into the corresponding azimuth sub-intervals. Construct the azimuth set that falls in the i-th azimuth sub-interval in the l-th sub-band of the n-th frame
[0030]
[0031] Then, the following formula is used to calculate The number of internal time-frequency points
[0032]
[0033] in, Indicates the calculation of the azimuth sub-interval set The number of time-frequency points in the azimuth. The index of the azimuth subinterval set containing the largest number of time-frequency point azimuths is calculated using the following formula:
[0034]
[0035] Finally, the azimuth angle set of the time-frequency points retained in the lth subband of the nth frame is obtained That is to say, It is the set of azimuth angle subintervals containing the largest number of time-frequency angle values in the lth subband of the nth frame.
[0036] 5. The implementation of step 5 is to divide the angle values of the time-frequency points in each sub-band of each frame of the microphone recording signal into the corresponding elevation angle sub-intervals in turn. Construct a set of elevation angle sub-intervals in which the elevation angle in the lth sub-band of the nth frame is in the ith elevation angle sub-interval.
[0037]
[0038] Then, the following formula is used to calculate The number of internal time-frequency points
[0039]
[0040] in, Indicates the calculation of the elevation angle sub-interval set The number of elevation angles within the time-frequency points.
[0041] according to The index of the azimuth subinterval set containing the largest number of time-frequency point azimuths is calculated using the following formula:
[0042]
[0043] Finally, the elevation angle set of the time-frequency points retained in the lth subband of the nth frame is obtained
[0044] That is to say, It is the set of elevation subintervals containing the largest number of time-frequency angle values in the lth subband of the nth frame.
[0045] 6. The implementation of step 6 is to obtain the azimuth angle set of the reserved time-frequency points of the lth sub-band of the nth frame Construct a collection of filtered azimuths:
[0046]
[0047] At the same time, the elevation angle set of the retained time-frequency points of the lth subband of the nth frame is obtained Construct a collection of filtered elevation angles:
[0048]
[0049] Finally, S μ and S η The angle of the time-frequency points within the range is estimated by kernel density estimation, and the angular distribution curves of azimuth and elevation are obtained. The angle value corresponding to the peak point in the curve is the estimated direction of arrival of the sound source.
[0050] Beneficial effects
[0051] This method uses the angular distribution of signal time-frequency points to localize multiple sound sources. Compared to traditional detection methods based on single-source regions, this method filters points within each time-frequency domain, identifying a large number of single-source regions and removing numerous outliers. This improves the accuracy of multi-source localization, making it more stable under reverberant conditions and applicable to complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This is the overall framework diagram of this design method. DETAILED DESCRIPTION
[0053] This example detects the direction of arrival of multiple sound sources under 600ms of reverberation. The sound sources are located in a 6.0m × 4.0m × 3.0m anechoic chamber. The sound field microphone is 1m above the ground, and the sound sources and microphones are located on the same horizontal plane. The distance between the sound sources and the microphones is 1m, and adjacent sound sources are spaced 90° apart. The number of sound sources is set to three. The signal processing software used is Matlab 2014a.
[0054] During implementation, the present invention embeds the algorithm into the software to realize the automatic operation of each process. The present invention is further explained below with specific implementation steps and accompanying drawings: The specific workflow is as follows:
[0055] Step 1: Obtain the azimuth and elevation information of each time-frequency point
[0056] A four-channel sound field microphone is used to collect scene data and record signals from multiple sound sources. The signals collected from the four channels are called M UFL , M DFR , M DBL , M UBR , such signals are recorded as A format signals. B format signals are obtained by linear conversion of A format signals, and the four channel signals of B format signals are recorded as M W , M X , M Y ,M Z , where M W Indicates the omnidirectional channel signal, M X , M Y ,M Z Represents the channel signals of the three coordinate axes X, Y, and Z of the Cartesian coordinate system. The conversion formula is as follows:
[0057]
[0058] Considering the short-term stability of the sound signal, the recorded signal needs to be framed before positioning processing. The present invention divides the microphone collected signal to be processed into N frames, and the frame length of each processing frame is K, that is, each frame contains K samples.
[0059] After obtaining the B-format signal, the time-frequency domain coefficient M of the time-frequency point (n, k) is obtained by performing a K-point short-time Fourier transform on each frame of the B-format signal. W (n, k), M X (n, k), M Y (n, k), M Z (n, k), where n represents the frame index, n=1, 2...N, and k represents the frequency index, k=1, 2...K. The time-frequency domain coefficients of the B-format signal are calculated using the following formula to obtain the activity intensity vector of the time-frequency point (n, k); T X(n, k), T Y (n, k), T Z (n, k) are the components of the activity intensity vector of the time-frequency point (n, k) in the x, y, and z directions respectively;
[0060]
[0061] Where Re{·} represents the real part of the calculation, (·)* represents the conjugate of the calculation, and the weight coefficient Q is obtained by the following formula:
[0062]
[0063] Where ρ and c represent the density and speed of sound of the medium, respectively. The azimuth angle μ(n, k) and elevation angle η(n, k) of the time-frequency point (n, k) are calculated as follows:
[0064]
[0065]
[0066] Step 2: Divide each frame of recorded signal into sub-bands in the frequency domain, and construct the azimuth angle set and elevation angle set of the time-frequency points in each sub-band.
[0067] Each frame signal is evenly divided into L subbands in the frequency domain. Each subband contains C time-frequency points, where C = K / L. The azimuth angle set of the time-frequency points in the lth subband of the nth frame is constructed by the following formula: and elevation angle set
[0068]
[0069]
[0070] Wherein, k represents a frequency index, n represents a frame index, n=1, 2...N, l represents a subband index, l=1, 2,...,L.
[0071] Step 3: Divide the azimuth angle value range and the elevation angle value range into intervals, and establish azimuth angle sub-intervals and elevation angle sub-intervals.
[0072] Since the azimuth angle ranges from [0° to 360°] and the elevation angle ranges from [-90° to 90°], the present invention divides the azimuth angle into 18 azimuth angle sub-intervals at intervals of 20°. In addition, the elevation angle is divided into 18 elevation angle sub-intervals at intervals of 10°.
[0073] Step 4: Divide the azimuth angles of each subband time-frequency point in the recorded signal into corresponding azimuth angle subintervals, and construct azimuth angle sets for different angle subintervals in each subband. Determine the number of azimuth angles in each azimuth angle subinterval set, and retain the azimuth angle subinterval set containing the largest number of azimuth angles.
[0074] Divide the angle values of the time-frequency points in each sub-band of each frame of microphone recording signal into the corresponding azimuth sub-intervals. Construct the azimuth set that falls in the i-th azimuth sub-interval in the l-th sub-band of the n-th frame
[0075]
[0076] Then, the following formula is used to calculate The number of internal time-frequency points
[0077]
[0078] in, Indicates the calculation of the azimuth sub-interval set The number of time-frequency points in the azimuth. The index of the azimuth subinterval set containing the largest number of time-frequency point azimuths is calculated using the following formula:
[0079]
[0080] Finally, the azimuth angle set of the time-frequency points retained in the lth subband of the nth frame is obtained That is to say, It is the set of azimuth angle subintervals containing the largest number of time-frequency angle values in the lth subband of the nth frame.
[0081] Step 5: Divide the elevation angle of each sub-band time-frequency point into corresponding elevation angle sub-intervals, and construct elevation angle sets for different angle sub-intervals in each sub-band. Determine the number of elevation angles in each elevation angle sub-interval set, and retain the elevation angle sub-interval set with the largest number of elevation angles.
[0082] Similarly, the angle values of the time-frequency points in each sub-band of each frame of the microphone recording signal are divided into the corresponding elevation angle sub-intervals in turn. Construct a set of elevation angle sub-intervals with the elevation angle in the lth sub-band of the nth frame in the ith elevation angle sub-interval.
[0083]
[0084] Then, the following formula is used to calculate The number of internal time-frequency points
[0085]
[0086] in, Indicates the calculation of the elevation angle sub-interval set The number of elevation angles within the time-frequency points.
[0087] according to The index of the azimuth subinterval set containing the largest number of time-frequency point azimuths is calculated using the following formula:
[0088]
[0089] Finally, the elevation angle set of the time-frequency points retained in the lth subband of the nth frame is obtained
[0090] That is to say, It is the set of elevation subintervals containing the largest number of time-frequency angle values in the lth subband of the nth frame.
[0091] Step 6: Construct a set of filtered azimuth and elevation angles, perform kernel density estimation on the angles in these two sets, and obtain the estimated direction of arrival of the sound source.
[0092] According to the obtained azimuth set of the reserved time-frequency points of the lth subband of the nth frame Construct a collection of filtered azimuths:
[0093]
[0094] At the same time, the elevation angle set of the retained time-frequency points of the lth subband of the nth frame is obtained Construct a collection of filtered elevation angles:
[0095]
[0096] Finally, S μ and S η The angle of the time-frequency points within the range is estimated by kernel density estimation, and the angular distribution curves of azimuth and elevation are obtained. The angle value corresponding to the peak point in the curve is the estimated direction of arrival of the sound source.
[0097] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.
Claims
1. A method for localizing multiple sound sources using the angle distribution information of the time-frequency points of the signal, characterized in that The following steps are involved: Step 1: Obtain the recorded signal and calculate the azimuth and elevation information of each time-frequency point; Step 2: Divide each frame of recorded signal into sub-bands in the frequency domain, and construct an azimuth angle set and an elevation angle set of time-frequency points in each sub-band; Step 3, dividing the azimuth angle value range and the elevation angle value range into intervals, and establishing azimuth angle sub-intervals and elevation angle sub-intervals; Step 4: Divide the azimuth angle of each sub-band time-frequency point in the recorded signal into corresponding azimuth angle sub-intervals, and construct azimuth angle sets of different angle sub-intervals in each sub-band; determine the number of azimuth angles in the different azimuth angle sub-interval sets, and retain the azimuth angle sub-interval set containing the largest number of azimuth angles; Step 5: Divide the elevation angle of each sub-band time-frequency point into corresponding elevation angle sub-intervals, and construct elevation angle sets of different angle sub-intervals in each sub-band; determine the number of elevation angles in different elevation angle sub-interval sets, and retain the elevation angle sub-interval set containing the largest number of elevation angles; In step 6, the set of azimuth angles retained in each sub-band of each frame is used to construct a set of filtered azimuth angles. At the same time, the set of elevation angles retained in each sub-band of each frame is used to construct a set of filtered elevation angles. Finally, kernel density estimation is performed on the angles in these two sets to obtain the angular distribution curves of azimuth and elevation angles. The angle value corresponding to the peak point in the curve is the estimated direction of arrival of the sound source.
2. The method for localizing multiple sound sources using signal time-frequency angle distribution information according to claim 1, wherein: Obtain azimuth and elevation information for each time-frequency point: Use a four-channel sound field microphone to collect scene data and capture recording signals from multiple sound sources; The signals collected by the four channels are called M UFL ,M DFR ,M DBL ,M UBR , such signals are recorded as A format signals; B format signals are obtained by linear conversion of A format signals, and the four channel signals of B format signals are recorded as M W ,M X ,M Y ,M Z , where M W Indicates the omnidirectional channel signal, M X ,M Y ,M Z Respectively represent the channel signals of the three coordinate axes X, Y, and Z of the Cartesian coordinate system; the conversion formula is as follows: The microphone signal to be processed is divided into N frames, and the frame length of each processing frame is K, that is, each frame contains K samples; After obtaining the B-format signal, the time-frequency domain coefficient M of the time-frequency point (n, k) is obtained by performing a K-point short-time Fourier transform on each frame of the B-format signal. W (n,k),M X (n,k),M Y (n,k),M Z (n, k), where n represents the frame index, n = 1, 2...N, and k represents the frequency index, k = 1, 2...K; the time-frequency domain coefficients of the B-format signal are calculated using the following formula to obtain the activity intensity vector of the time-frequency point (n, k): T X (n,k), T Y (n,k), T Z (n, k) are the components of the activity intensity vector of the time-frequency point (n, k) in the x, y, and z directions respectively; Where Re{·} represents the real part of the calculation, (·)* represents the conjugate of the calculation, and the weight coefficient Q is obtained by the following formula: Where ρ and c represent the density and sound speed of the medium, respectively. The azimuth angle μ(n,k) and elevation angle η(n,k) of the time-frequency point (n,k) are calculated as follows: Divide each frame of recorded signal into sub-bands in the frequency domain, and construct the azimuth angle set and elevation angle set of the time-frequency points in each sub-band; Each frame signal is evenly divided into L subbands in the frequency domain. Each subband contains C time-frequency points, where C = K / L. The azimuth angle set of the time-frequency points in the lth subband of the nth frame is constructed by the following formula: and elevation angle set Wherein, k represents the frequency index, n represents the frame index, n=1, 2...N, l represents the subband index, l=1, 2,..., L; The azimuth angle range and the elevation angle range are divided into intervals respectively to establish the azimuth angle sub-interval and the elevation angle sub-interval. Since the azimuth angle range is [0°, 360°] and the elevation angle range is [-90°, 90°], the azimuth angle is divided into 18 azimuth angle sub-intervals at intervals of 20°. In addition, the elevation angle is divided into 18 elevation angle sub-intervals at intervals of 10°. Divide the azimuth angle of each sub-band time-frequency point in the recorded signal into corresponding azimuth angle sub-intervals, and construct azimuth angle sets of different angle sub-intervals in each sub-band; determine the number of azimuth angles in different azimuth angle sub-interval sets, and retain the azimuth angle sub-interval set containing the largest number of azimuth angles; Divide the angle values of the time-frequency points in each sub-band of each frame of microphone recording signal into the corresponding azimuth sub-intervals; construct the azimuth set that falls in the i-th azimuth sub-interval in the l-th sub-band of the n-th frame Then, the following formula is used to calculate The number of internal time-frequency points in, Indicates the calculation of the azimuth sub-interval set The number of internal time-frequency point azimuths; according to The index of the azimuth subinterval set containing the largest number of time-frequency point azimuths is calculated using the following formula: Get the azimuth angle set of the time-frequency points retained in the lth subband of the nth frame That is to say, It is the set of azimuth angle subintervals containing the largest number of time-frequency angle values in the lth subband of the nth frame; Divide the elevation angle of each sub-band time-frequency point into corresponding elevation angle sub-intervals, and construct elevation angle sets of different angle sub-intervals in each sub-band; determine the number of elevation angles in different elevation angle sub-interval sets, and retain the elevation angle sub-interval set containing the largest number of elevation angles; Divide the angle values of the time-frequency points in each sub-band of each frame of the microphone recording signal into the corresponding elevation angle sub-intervals in turn; construct a set of elevation angles in the ith elevation angle sub-intervals of the ith sub-band of the nth frame. Then, the following formula is used to calculate The number of internal time-frequency points in, Indicates the calculation of the elevation angle sub-interval set The number of elevation angles of the inner time-frequency points; according to The index of the azimuth subinterval set containing the largest number of time-frequency point azimuths is calculated using the following formula: Get the elevation angle set of the time-frequency points retained in the lth subband of the nth frame That is to say, It is the set of elevation subintervals containing the largest number of time-frequency angle values in the lth subband of the nth frame; Construct a set of filtered azimuth and elevation angles, perform kernel density estimation on the angles in these two sets, and obtain the estimated direction of arrival of the sound source; According to the obtained azimuth set of the reserved time-frequency points of the lth subband of the nth frame Construct a collection of filtered azimuths: The obtained elevation angle set of the reserved time-frequency points of the lth subband of the nth frame Construct a collection of filtered elevation angles: Finally, S μ and S η The angle of the time-frequency points within the range is estimated by kernel density estimation, and the angular distribution curves of azimuth and elevation are obtained. The angle value corresponding to the peak point in the curve is the estimated direction of arrival of the sound source.
Citation Information
Patent Citations
Multi-sound-source positioning method using dominant sound source component removal
CN110275138A
Sound source positioning method based on double-microphone array
CN111798869A