Multi-source direction finding method based on vector microphone in strong reverberation environment

By combining frequency smoothing and subspace projection algorithms with the MUSIC algorithm, the accuracy problem of multi-source localization of vector microphones in strong reverberation environments is solved, and the single-source time-frequency point is accurately extracted in the time-frequency domain, thus improving the accuracy of multi-source localization.

CN115754898BActive Publication Date: 2026-03-17WUHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-31
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing sound source localization methods based on vector microphones have difficulty accurately detecting the time and frequency points of multiple sound sources in strong reverberation environments, resulting in inaccurate localization.

Method used

Frequency smoothing technology is used to remove reverberation interference. Single-source time-frequency points unaffected by reverberation interference are extracted by constructing a steering vector dictionary and a subspace projection algorithm. Multi-source direction finding and localization are achieved by combining the MUSIC algorithm.

Benefits of technology

It can accurately extract single-source time-frequency points in the time-frequency domain without reverberation interference, thus improving the accuracy and precision of multi-source localization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115754898B_ABST
    Figure CN115754898B_ABST
Patent Text Reader

Abstract

The application discloses a multi-sound source direction-finding positioning method based on a vector microphone in a strong reverberation environment, which comprises the following steps: converting mixed voice signals collected by the vector microphone from a waveform domain to a time-frequency domain to obtain a spectrogram of an observation signal; removing virtual source single time-frequency points by using a frequency smoothing method and iteratively generating low-reverberation self-term source time-frequency point sets; constructing a steering vector dictionary; finding a steering vector most likely to appear in any time-frequency point in the self-term source time-frequency point sets from the steering vector dictionary, and then calculating the STFT coefficient value of the video point corresponding to the source according to the steering vector, determining whether the time-frequency point is a single source time-frequency point not interfered by reverberation according to the STFT coefficient value, and generating a single source time-frequency point set not interfered by reverberation; estimating the number of sources by using a smoothing histogram method; and obtaining multi-source direction-finding results by using MUSIC. The application can accurately extract single source time-frequency points not interfered by reverberation in the time-frequency domain, and then realize accurate multi-source positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the technical field of acoustics, specifically relating to a multi-source direction finding and localization method based on a vector microphone in a strong reverberation environment. Background Technology

[0002] A vector microphone consists of two or three orthogonal directional sensors and an omnidirectional microphone. It measures particle velocity along orthogonal axes and sound pressure levels at the same location. The components in a vector microphone are positioned very close together in space, approximating each other as a single point. This assumption ignores the distance between sensors, thus assuming no phase difference between the signals received by each sensor—a characteristic known as frequency independence. Utilizing the rich acoustic information output by vector microphones and their unique frequency independence, they significantly outperform traditional omnidirectional microphone arrays in array signal processing. Furthermore, vector microphones are small and lightweight, attracting considerable attention in acoustic research and finding widespread application in emerging industries such as target tracking, underwater acoustic communication, and sound source localization.

[0003] Direction of Arrival (DOA) estimation, or source localization, determines the location of a target sound source by processing and analyzing sound signals received by a microphone array. It holds great promise for applications such as far-field speech recognition and intelligent video conferencing, and is an important speech front-end processing technology. With the rise of vector microphones in array signal processing, many source localization algorithms have been extended from traditional omnidirectional microphone arrays to vector microphones, achieving more accurate source direction finding. In recent years, vector microphone-based source localization has primarily focused on multi-source localization in reverberant indoor environments. This involves converting the signals received by the vector microphone from the waveform domain to the time-frequency domain, designing algorithms to detect single-source time-frequency points unaffected by reverberation, and then using the Multiple Signal Classification (MUSIC) algorithm to estimate the direction of multiple sound sources in space. However, as the number of sound sources increases and indoor reverberation intensifies, existing methods are insufficient to accurately detect these reverberation-free single-source time-frequency points, thus hindering accurate source localization. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by providing a multi-source direction finding and localization method based on a vector microphone in a strong reverberation environment. This method can accurately extract single-source time-frequency points that are not affected by reverberation in the time-frequency domain, thereby achieving accurate multi-source localization.

[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0006] A multi-source direction finding and localization method based on a vector microphone in a strongly reverberant environment includes the following steps:

[0007] Step 1: Convert the mixed speech signal collected by the vector microphone from the waveform domain to the time-frequency domain to obtain the spectrogram X(t,f) of the observed signal;

[0008] Step 2: Use the frequency smoothing method to remove the virtual source single-source time-frequency points from the spectrogram of the observed signal obtained in Step 1, and iteratively generate a set of low-reverberation self-source time-frequency points;

[0009] Step 3: Cluster the direction vectors corresponding to the self-source time-frequency point set generated in Step 2, and integrate the cluster centers of each cluster as the guiding vector dictionary Θ. A ;

[0010] Step 4: For any time-frequency point in the self-source time-frequency point set generated in Step 2, find the steering vector most likely to appear at that time-frequency point from the steering vector dictionary generated in Step 3, then calculate the STFT coefficient value of the source corresponding to that video point based on the steering vector, determine whether the time-frequency point is a single-source time-frequency point that is not subject to reverberation interference based on the STFT coefficient value, and generate a set of single-source time-frequency points that are not subject to reverberation interference.

[0011] Step 5: Calculate the estimated correlation angle information for each time-frequency point in the single-source time-frequency point set obtained in Step 4 and generate a histogram. Smooth the histogram and select the peaks with peak values ​​greater than the threshold for counting. The final number of peaks is the number of sources.

[0012] Step 6: Cluster the single-source time-frequency point set generated in Step 4 that is not affected by reverberation to obtain a single-source time-frequency point subset for each source. Use the MUSIC algorithm iteratively to estimate the location of multiple single sources for each subset. Finally, integrate all single-source location estimation results to obtain the multi-source direction finding results.

[0013] Furthermore, in step 1, the method for converting the mixed speech signal from the waveform domain to the time-frequency domain is as follows:

[0014] The mixed speech signal collected by the vector microphone is obtained by convolving the signal transmission response between the source and the microphone with the original signal, and is represented as:

[0015]

[0016] Q represents the number of sources in the reverberant environment, h q (t) represents the transmission response from the q-th source to the vector microphone, s q (t) represents the q-th source signal;

[0017] The spectrum X(t,f) of the observed signal x(t) is obtained after undergoing the Fourier Short Time Fourier Transform:

[0018]

[0019] Wherein d(Ψ) j ) is the steering vector of the j-th source, S j (t,f) represents the STFT coefficient corresponding to the j-th source, where J is the total number of sources.

[0020] Furthermore, the specific method for step 2 is as follows:

[0021] The spectrogram of the observed signal in step 1 includes single-source time-frequency points of the actual source, single-source time-frequency points of the virtual source, and multi-source time-frequency points. The neighborhood time-frequency correlation matrix is ​​calculated by setting a time-frequency sliding window. The neighborhood time-frequency correlation matrix is ​​decomposed into eigenvalues ​​to obtain a series of eigenvalues. When the ratio of the largest eigenvalue to the second largest eigenvalue is greater than a preset threshold, the time-frequency point (t,f) is considered to be a self-source time-frequency point without reverberation interference. The set of low-reverberation self-source time-frequency points is generated iteratively to remove virtual single-source time-frequency points on the spectrogram of the observed signal to the maximum extent.

[0022] Furthermore, the specific calculation method for step 2 is as follows:

[0023] Calculate the neighborhood time-frequency correlation matrix using a time-frequency sliding window:

[0024]

[0025] J t and J f J represents the size of the rectangular sliding window. t Slide J on the timeline f Slide along the frequency axis. Due to the frequency independence of vector microphones, the source's steering vector does not change with frequency during frequency smoothing; therefore, the neighborhood time-frequency correlation matrix... This will be represented as:

[0026]

[0027] In the formula, d(Ψ) J ) is the steering vector of the J-th source. Represents the time-frequency correlation matrix of the source's neighborhood;

[0028] Neighborhood time-frequency correlation matrix Perform eigenvalue decomposition:

[0029]

[0030] Select time-frequency points (t, f) where the ratio of the largest eigenvalue to the second largest eigenvalue is greater than a threshold:

[0031]

[0032] σ max (t,f) and σ submax (t,f) are the largest and second largest feature values, respectively, and λ1 represents a pre-set threshold.

[0033] Furthermore, the specific method for constructing the direction vector dictionary in step 3 is as follows:

[0034] Perform k-means clustering on the steering vectors corresponding to the self-source time-frequency point set obtained in step 2 to obtain N clusters of steering vectors, and calculate the cluster center of each cluster. And integrate them to obtain the guide vector dictionary Θ A :

[0035]

[0036] Ω k It contains the time-frequency points of the self-source in the k-th cluster, Card(Ω) k () represents the number of time-frequency points contained in the k-th class. This represents the steering vector corresponding to each time frequency point.

[0037] Furthermore, the guiding vector dictionary Θ A It includes the steering vector of each actual source and the composite steering vector of multiple sources.

[0038] Furthermore, in step 4, a subspace projection algorithm is used to select single-source time-frequency points. The specific method is as follows:

[0039] Step 4.1: Find the two sources most likely to appear at each time frequency point using the subspace projection algorithm, and randomly select the steering vector dictionary Θ generated in Step 3. A Select two column vectors and Stack two column vectors into a matrix:

[0040]

[0041] Calculate the noise subspace of matrix A2:

[0042]

[0043] Select the optimal combination of steering vectors:

[0044]

[0045] Step 4.2: Calculate the STFT coefficients of the two sources by matrix inversion:

[0046]

[0047] Calculate the ratio of the larger to the smaller of the two coefficient values:

[0048]

[0049] In the formula, λ2 is the set threshold;

[0050] When two STFT coefficient values ​​satisfy the above inequality, the time-frequency point is considered to be a single-source time-frequency point that is not affected by reverberation. Repeat the above steps, traverse the set of source time-frequency points, and iteratively select single-source time-frequency points that are not affected by reverberation.

[0051] Furthermore, the specific implementation method of step 5 is as follows:

[0052] Step 5.1: Based on the array manifold of the vector microphone, calculate the angle information corresponding to each time-frequency point in the single-source time-frequency point set generated in Step 4:

[0053]

[0054]

[0055] X p (t,f) represents the signal collected by the sound pressure sensor in the vector microphone, X vx (t,f), X vy (t,f) and X vz (t,f) represent the signals collected by the three particle velocity sensors along the x, y, and z axes, respectively. Indicates the operation of taking the real part, {·} * To indicate taking the complex number, θ°∈(-90°,90°] are the estimated azimuth and elevation angles, respectively;

[0056] Step 5.2: Calculate the estimated angle values ​​obtained from the time-frequency points of all single sources, generate a histogram, and smooth it using a Gaussian filter to obtain a smoothed histogram. Perform a peak search on the smoothed histogram to find the number of peaks with a peak value greater than a set threshold. This number is the number of sources.

[0057] Furthermore, the specific method for implementing multi-source direction finding using the MUSIC algorithm in step 6 is as follows:

[0058] Step 6.1: Perform k-means clustering on all single-source time-frequency points obtained in Step 4 to obtain a Q-group time-frequency point set;

[0059] Step 6.2, for each set of time and frequency points First, calculate the covariance matrix:

[0060]

[0061] The covariance matrix R is decomposed into the signal subspace U s and noise subspace U n :

[0062]

[0063] In the formula, ∑ s and ∑ n The covariance matrices of the signal and noise are respectively, {·} H This represents the conjugate transpose operation;

[0064] Calculate the spatial spectrum:

[0065]

[0066] In the formula, d(Ψ) is the guiding vector;

[0067] Search space for Ψ to find the point where P is made MUSIC The largest Ψ is the q-th single-source time-frequency point set. Find the DOA of the corresponding source; repeat the above operation for each group of single-source time-frequency points until the DOA values ​​of all sources are calculated, thus completing the final multi-source measurement.

[0068] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0069] This invention addresses the problem of multi-source direction finding in environments with strong reverberation. It utilizes the frequency independence of vector microphones and the time-frequency sparsity of signals, employing frequency smoothing techniques to remove the influence of reverberation on the detection of single-source time-frequency points. Then, a subspace projection algorithm is used to extract single-source points unaffected by reverberation. This allows for accurate extraction of reverberation-free single-source time-frequency points in the time-frequency domain, thereby achieving accurate multi-source localization. The greatest advantage of this invention is that it significantly improves the detection accuracy of low-reverberation single-source time-frequency points, resulting in more accurate position estimation when using the MUSIC algorithm for multi-source direction finding. Attached Figure Description

[0070] Figure 1 This is a flowchart of the multi-sound-source direction finding and localization method according to an embodiment of the present invention;

[0071] Figure 2 This is a flowchart illustrating the generation of the steering vector dictionary and the detection of single-source time-frequency points in this embodiment of the invention.

[0072] Figure 3 The test performance of the embodiments of the present invention under different numbers of sources is shown. RMSAE represents the root mean square angle error. Detailed Implementation

[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0074] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0075] The present invention will be further described below with reference to specific embodiments, but these are not intended to limit the scope of the invention.

[0076] Traditional sound source localization methods based on vector microphones focus on exploring the special mathematical properties of the vector microphone array manifold to extract single-source time-frequency points unaffected by reverberation. However, as reverberation intensifies and the number of sources increases, traditional methods become highly limited. This invention fully utilizes the frequency independence of vector microphones and combines it with time-frequency sparsity to design a method for accurately extracting single-source time-frequency points unaffected by reverberation, thereby achieving accurate multi-source localization.

[0077] This invention provides a multi-source direction finding and localization method based on a vector microphone in a strong reverberation environment, comprising the following steps:

[0078] Step 1: Convert the mixed speech signal collected by the vector microphone from the waveform domain to the time-frequency domain to obtain the spectrogram X(t,f) of the observed signal;

[0079] In this embodiment, the sampling rate sr of the speech signal is 16kHz, the frame length is set to 16ms, and the frame shift is 8ms. Therefore, the length of the window function for each frame is sr×16ms, and the step size is sr×8ms. The Fast Fourier Transform (FFT) has 1024 points, and the Hamming window is selected as the windowing function.

[0080] In this embodiment, the signal transmission model in a reverberant environment is a typical convolution model. The mixed speech signal collected by the vector microphone is obtained by convolving the signal transmission response between the source and the microphone with the original signal, and can be represented as:

[0081]

[0082] In the formula, Q represents the number of sources in the reverberant environment, and h q (t) represents the transmission response from the q-th source to the vector microphone, s q (t) represents the q-th source signal.

[0083] The spectrum X(t,f) of the observed signal x(t) is obtained after undergoing a Short-Time Fourier Transform (STFT), and is represented as:

[0084]

[0085] In the formula, H q (f) represents the transfer function from the q-th source to the vector microphone, S q (t,f) represents the STFT coefficient corresponding to the q-th source.

[0086] According to the mirror model theory of signal reflection, after a signal is reflected by a plane, the reflected signal can be equivalent to a signal emitted by a virtual source that is the same as the actual source. In this embodiment, only the virtual source equivalent to a strong reflected signal is considered, and weaker reflected signals are not considered. Therefore, it can be assumed that there are J sources in space, including Q actual sources and JQ virtual sources. X(t,f) will be re-expressed as:

[0087]

[0088] In the formula, S j (t,f) represents the STFT coefficient corresponding to the j-th source, d(Ψ) j Let be the steering vector of the j-th source, which can be represented by the array manifold of the vector microphone as:

[0089]

[0090] In the formula, θ j ∈(-90°,90°] corresponds to the azimuth and elevation angles of the j-th source relative to the vector microphone, respectively.

[0091] Step 2: Use a frequency smoothing method to remove virtual source single-source time-frequency points as much as possible from the spectrogram of the observed signal obtained in Step 1, and iteratively generate a set of low-reverberation self-source time-frequency points;

[0092] Based on X(t,f) in step 1, it can be deduced that the spectrogram of the observed signal contains three types of time-frequency points: single-source time-frequency points of the actual source, single-source time-frequency points of the virtual source, and multi-source time-frequency points. If "point-level" single-source time-frequency point detection is performed directly, the single-source points of the virtual source will inevitably be detected. Therefore, it is necessary to remove the interference of the single-source points of the virtual source beforehand. In this embodiment, a frequency smoothing method is used to eliminate reverberation interference as much as possible at the "regional level," thereby achieving the purpose of removing the single-source time-frequency points of the virtual source.

[0093] The specific implementation of this step includes the following sub-steps:

[0094] Step 2.1: Set up a time-frequency sliding window to calculate the neighborhood time-frequency correlation matrix;

[0095]

[0096] In the formula, J t J f J represents the size of the rectangular sliding window. t Slide J on the timeline f Sliding along the frequency axis, due to the frequency independence of the vector microphone, the source's steering vector does not change with frequency during frequency smoothing. Therefore, the neighborhood time-frequency correlation matrix... This will be represented as:

[0097]

[0098] Represents the time-frequency correlation matrix of the source's neighborhood;

[0099] Step 2.2, perform a neighborhood time-frequency correlation matrix Eigenvalue decomposition is performed, and the time-frequency points (t,f) whose largest eigenvalue is much larger than the second largest eigenvalue are selected to form a set of self-source time-frequency points that are not affected by reverberation.

[0100] Frequency smoothing can effectively decouple coherent sources. Since virtual sources are strongly correlated with actual sources, directly calculating the source covariance matrix would result in a non-full-rank covariance matrix. After neighborhood frequency smoothing, the rank of the resulting neighborhood time-frequency correlation matrix will equal the sum of the number of active actual and virtual sources at that time-frequency point. Elementary matrix transformations yield a diagonal matrix:

[0101]

[0102] σ i (t,f) is Substitute the i-th eigenvalue into have to:

[0103]

[0104] When the time frequency (t,f) is a possible single-source time frequency, then The largest eigenvalue must be much larger than the second largest eigenvalue, that is:

[0105]

[0106] σ max (t,f) and σ submax(t,f) are the largest and second largest eigenvalues, respectively, and λ1 represents a pre-set threshold.

[0107] When the largest eigenvalue of a time frequency point (t,f) is much larger than the second largest eigenvalue and satisfies the above inequality, the time frequency point is considered to be a self-source time frequency point that is not affected by reverberation. Repeat the above steps to traverse all time frequency points and iteratively generate a set of low-reverberation self-source time frequency points.

[0108] Frequency smoothing removes the influence of virtual sources as much as possible. However, since the detection is performed at the "region level", the selected point set contains single-source time-frequency points as well as a large number of multi-source time-frequency points. The composition of this point set is similar to that of self-source time-frequency points, but it contains a small number of single-source time-frequency points with virtual sources as well as multi-source time-frequency points. Therefore, this point set is called the low-reverberation self-source time-frequency point set.

[0109] Step 3, construct the dictionary of guiding vectors; such as Figure 2 As shown, the direction vectors corresponding to the self-source time-frequency point set obtained in step 2 are subjected to k-means clustering to obtain N clusters of guiding vectors. Since the self-source time-frequency points include single-source time-frequency points and a large number of multi-source time-frequency points, and the number of sources Q is unknown, N > Q is set, and the cluster center of each class is calculated. Obtain the guide vector dictionary Θ A :

[0110]

[0111]

[0112] In the formula, Ω k It contains the time-frequency points of the self-source in the k-th cluster, Card(Ω) k () represents the number of time-frequency points contained in the k-th class. This represents the steering vector corresponding to each time frequency point. Steering vector dictionary Θ A It includes the steering vector of each actual source, as well as the composite steering vector of multiple sources.

[0113] Step 4: Select single-source time-frequency points using a subspace projection algorithm. Based on the sparsity of signals in the time-frequency domain, most time-frequency points are generally dominated by at most two high-energy sources. Time-frequency points containing three or more high-energy sources are relatively few and can be ignored. Therefore, this embodiment assumes that at each single-source time-frequency point, there are at most two strong-energy sources, and other weak-energy sources can be ignored. The STFT coefficient values ​​of the two most likely sources at each time-frequency point are calculated using the subspace projection method. When the STFT coefficient value of one source is much larger than that of the other source, the time-frequency point is considered a single-source time-frequency point unaffected by reverberation. Step 4 specifically includes the following two sub-steps:

[0114] Step 4.1: Find the two sources most likely to appear at each time frequency point using the subspace projection algorithm. The STFT coefficients of these two sources can be expressed as:

[0115]

[0116] In the formula, and These are the steering vectors corresponding to the two sources. This indicates the operation of finding the pseudo-inverse of a matrix. To obtain the STFT coefficients of two sources, the steering vectors of the two sources need to be determined first. This embodiment uses the subspace projection method to determine the steering vectors of the two sources most likely to appear at that time-frequency point. Two column vectors are randomly selected from the steering vector dictionary generated in step 3. and Stack two column vectors into a matrix:

[0117]

[0118] Calculate the noise subspace of matrix A2:

[0119]

[0120] When the product of the noise subspace and the observed signal, PX(t,f), is minimized, the two selected column vectors are considered optimal, and their corresponding sources are the two most likely sources to appear at that time-frequency point. The selection of the optimal steering vector can be expressed as:

[0121]

[0122] Each time, randomly from Θ A Choose two column vectors and iteratively calculate PX(t,f) until the best combination that minimizes PX(t,f) is found. This combination is the steering vector of the two sources most likely to appear at that time frequency point.

[0123] Step 4.2: Calculate the STFT coefficients of the two sources by matrix inversion, and determine whether the absolute ratio of the two STFT coefficients is greater than a set threshold λ2. When one STFT coefficient is significantly larger than the other, the point is considered a single-source time-frequency point. The operation of calculating the STFT coefficients can be expressed as:

[0124]

[0125] Calculate the ratio of the larger to the smaller of the two coefficient values:

[0126]

[0127] When the two STFT coefficient values ​​satisfy the above inequality, the time-frequency point is considered to be a single-source time-frequency point that is not affected by reverberation. Step 4 is repeated to traverse the set of self-source time-frequency points and iteratively select single-source time-frequency points that are not affected by reverberation.

[0128] Step 5: Estimate the number of sources using the smoothed histogram method. This step specifically includes:

[0129] Step 5.1: Based on the array manifold of the vector microphone, calculate the directional angle corresponding to each time-frequency point in the single-source time-frequency point set generated in Step 4. For a single-source time-frequency point, the observation vector of that time-frequency point can be expressed as:

[0130]

[0131] X p (t,f) represents the signal collected by the sound pressure sensor in the vector microphone, X vx (t,f), X vy (t,f) and X vz (t,f) represent the signals collected by the three particle velocity sensors along the x, y, and z axes, respectively. Let θ be the azimuth angle. q Let be the elevation angle. For a known single-source time-frequency point, its corresponding direction angle can be estimated as:

[0132]

[0133]

[0134] In the formula, Indicates the operation of taking the real part, {·} * To indicate taking the complex number, θ°∈(-90°,90°] corresponds to the estimated azimuth and elevation angles, respectively;

[0135] Step 5.2: Calculate the estimated angle values ​​obtained from the time-frequency points of all single sources and generate a histogram. Larger peaks will appear on the histogram within the angle range where sources appear. The number of peaks displayed on the histogram indicates the number of sources. To facilitate peak count, a Gaussian filter is used for smoothing, resulting in a smoothed histogram. A peak search is then performed on the smoothed histogram to find the number of peaks with a value greater than a set threshold; this number represents the number of sources.

[0136] Step 6: Implement multi-source direction finding using the MUSIC algorithm. In Step 5, when using a smoothed histogram to statistically analyze the angle information of all single-source time-frequency points, a rough estimate of the DOA for each source can be obtained. To further improve the direction finding accuracy, k-means clustering is first performed on all single-source time-frequency points obtained in Step 4, resulting in Q sets of time-frequency points. For each set of time-frequency points... First, calculate the covariance matrix:

[0137]

[0138] The covariance matrix R can be decomposed into the signal subspace U s and noise subspace U n :

[0139]

[0140] In the formula, ∑ s and ∑ n The covariance matrices of the signal and noise are respectively, {·} H This represents the conjugate transpose operation;

[0141] Calculate the spatial spectrum:

[0142]

[0143] In the formula, d(Ψ) is the guiding vector;

[0144] Search space for Ψ to find the point where P is made MUSIC The largest Ψ is the q-th single-source time-frequency point set. Calculate the DOA of the corresponding source. Repeat the above operation for each group of single-source time-frequency points until the DOA values ​​of all sources are calculated, thus completing the final multi-source measurement.

[0145] To illustrate the effectiveness of this embodiment, the method described in this embodiment was tested with different numbers of sources. RMSAE represents the root mean square angle error. The experimental results are as follows: Figure 3 As shown, the positioning effect is best with dual sources. As the number of sources increases, the positioning performance gradually decreases, but positioning with low error can still be achieved.

[0146] The above are merely preferred embodiments of the present invention and are not intended to limit the implementation methods and protection scope of the present invention. Those skilled in the art should recognize that any equivalent substitutions and obvious changes made based on the content of this specification should be included within the protection scope of the present invention.

Claims

1. A method for multi-sound source direction finding based on vector microphones in a reverberation environment, characterized in that, The method comprises the following steps: Step 1: convert the mixed speech signal collected by the vector microphone from the waveform domain to the time-frequency domain to obtain a spectrogram of the observation signal ; Step 2: removing the virtual source single-source time-frequency points from the spectrogram of the observation signal obtained in step 1 by using a frequency smoothing method, and iteratively generating a low-reverberation self-item source time-frequency point set; Step 3: Clustering the direction vectors corresponding to the self-sourced time-frequency point set generated in Step 2, and integrating the cluster center of each class as a steering vector dictionary ; Step 4: for any time-frequency point in the self-item source time-frequency point set generated in step 2, finding the steering vector most likely to appear in the time-frequency point from the steering vector dictionary generated in step 3, and then calculating the STFT coefficient value of the source corresponding to the time-frequency point according to the steering vector, determining whether the time-frequency point is a single-source time-frequency point not interfered by reverberation according to the STFT coefficient value, and generating a single-source time-frequency point set not interfered by reverberation; Step 5: calculating the estimated angle information of each time-frequency point in the single-source time-frequency point set calculated in step 4 and generating a histogram, smoothing the histogram, selecting the number of sharp peaks with a peak value greater than a threshold value, and obtaining the final number of sharp peaks, that is, the number of sources; Step 6: clustering the single-source time-frequency point set not interfered by reverberation generated in step 4 to obtain a single-source time-frequency point subset corresponding to each source, and iteratively implementing multiple single-source position estimation on each subset by using a MUSIC algorithm, and finally integrating all single-source position estimation results to obtain a multiple-source direction finding result.

2. The method of claim 1, wherein, In step 1, the method for converting the mixed speech signal from the waveform domain to the time-frequency domain is: The signal transmission response between the source and the microphone is convolved with the original signal to obtain the mixed speech signal collected by the vector microphone, denoted as: the number of sources in the reverberant environment, the transfer response of the thsource to the vector microphone, the thsource signal; Observed signal The spectrogram of the observed signal after short-time Fourier transform : ; wherein, is the steering vector of the th source, denotes the STFT coefficient corresponding to the th source, is the total number of sources.

3. The method of claim 1, wherein, The specific method of step 2 is: The spectrogram of the observation signal in step 1 contains single-source time-frequency points of actual sources, single-source time-frequency points of virtual sources and multi-source time-frequency points, and a neighborhood time-frequency correlation matrix is calculated by setting a time-frequency sliding window Eigenvalue decomposition is performed on the neighborhood time-frequency correlation matrix to obtain a series of eigenvalues, and when the ratio of the largest eigenvalue to the second largest eigenvalue is greater than a pre-set threshold value, the time-frequency point is considered to be a self-term source time-frequency point not interfered by reverberation, and a low-reverberation self-term source time-frequency point set is iteratively generated, so that the virtual source single-source time-frequency points on the spectrogram of the observation signal are removed to the maximum extent.

4. The method of claim 3, wherein, The specific calculation method of step 2 is: The time-frequency sliding window is set to calculate the neighborhood time-frequency correlation matrix: and is a rectangular sliding window size, sliding on the time axis, sliding on the frequency axis, according to the frequency independence of the vector microphone, the steering vector of the source does not change with the frequency in the frequency smoothing process, so the neighborhood time-frequency correlation matrix is represented as: wherein is the steering vector of the J th source, denotes the local time-frequency correlation matrix of the source. Neighbourhood time-frequency correlation matrix Eigenvalue decomposition is performed: In the formula, is the first eigenvalue, i= 1 -J ; Selecting a time-frequency point whose ratio of the largest eigenvalue to the second largest eigenvalue is greater than a threshold value : and are the largest and second largest eigenvalues, respectively, denotes a pre-set threshold value.

5. The method of claim 1, wherein, The specific method of constructing the direction vector dictionary in step 3 is: The guide vectors corresponding to the time-frequency point set from the source in step 2 are subjected to k-means clustering to obtain cluster guide vectors, and the cluster centers of each class are calculated and integrated to obtain a guide vector dictionary : , contains the first self-term source time-frequency point in the cluster, represents the number of time-frequency points contained in the first class, represents the steering vector corresponding to each time-frequency point.

6. The method of claim 1, wherein, Dictionary of steering vectors The dictionary of steering vectors comprises a steering vector for each actual source and a composite steering vector for the plurality of sources.

7. The method of claim 1, wherein, The specific method of selecting single-source time-frequency points by using the subspace projection algorithm in step 4 is: Step 4.1 Find the two sources most likely to be present at each time-frequency bin by the subspace projection algorithm, randomly select from the steering vector dictionary generated in Step 3 Two column vectors are chosen and Stack the two column vectors into a matrix: Computing the matrix of noise subspaces: Selecting the combination of the best steering vector: Step 4.2, calculating the STFT coefficient values of the two sources by matrix inversion: Calculating the ratio of the larger value to the smaller value between the two coefficient values: In the formula, When the two STFT coefficient values satisfy the above inequality, the time-frequency point is considered to be a single-source time-frequency point not interfered by reverberation, and the above steps are repeated to traverse the self-item source time-frequency point set and iteratively select the single-source time-frequency points not interfered by reverberation. 2 is a set threshold value; The specific implementation method of step 5 is:

8. The method of claim 1, wherein, Step 5.1, according to the array manifold of the vector microphone, calculating the angle information corresponding to each time-frequency point in the single-source time-frequency point set generated in step 4: Step 5.2, counting the estimated angle values calculated from all single-source time-frequency points to generate a histogram, and smoothing the histogram by using a Gaussian filter to obtain a smoothed histogram, searching for sharp peaks on the smoothed histogram, and finding the number of sharp peaks with a peak value greater than a set threshold value, which is the number of sources. the signals collected by the pressure sensors in the vector microphone, , and correspond to the signals collected by the three particle velocity sensors along the three axes, respectively, denotes the real part operation, denotes the complex number, and are the estimated azimuth and elevation angles, respectively; The specific method of implementing multiple-source direction finding by using the MUSIC algorithm in step 6 is:

9. The method of claim 1, wherein, Calculating the spatial spectrum: Step 6.

1. Perform k-means clustering on all single-source time-frequency points obtained in Step 4 to obtain group time-frequency point sets; Step 6.2, for each set of time-frequency points First, the covariance matrix is calculated: covariance matrix decomposition into signal subspace and noise subspace : ; wherein and are the covariance matrices of the signal and noise, respectively, denotes the conjugate transpose operation; ​ ; wherein ) is a steering vector; Searching from space , find the value that makes maximum , which is the first group of single-source time-frequency point set corresponding to the source DOA; for each group of single-source time-frequency points, repeat the above operation until the DOA values of all sources are calculated, and the final multi-source measurement is completed.