Sound source direction estimation method based on small linear microphone array
By using a small linear microphone array to estimate the direction of sound sources, and employing the LSDD mechanism to screen high-confidence time-frequency units and perform confidence-weighted fusion, the problems of accuracy and robustness of sound source localization in complex environments are solved, achieving high-precision multi-source resolution and pseudo-peak suppression.
Patent Information
- Application Number
- CN202510926965.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-28
AI Technical Summary
Existing sound source localization algorithms lack robustness and accuracy in complex environments, especially in the presence of reverberation and noise interference, making it difficult to accurately distinguish the directions of multiple sound sources.
A sound source direction estimation method based on a small linear microphone array is adopted. The high-confidence time-frequency unit is screened by introducing the Local Spatial Domain Distance (LSDD) mechanism, and the confidence is evaluated by using the relative spectral peak intensity. The direction spectrum is constructed by combining confidence weighted fusion, and reverberation interference is eliminated to improve the multi-source resolution capability.
It significantly improves the accuracy and robustness of sound source localization, reduces false peak errors, achieves high-precision multi-source resolution, and is suitable for embedded deployment.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
Technical Field
[0001] This invention relates to the field of sound source localization and signal processing technology, and in particular to a sound source direction estimation method based on a small linear microphone array. Technical Background
[0002] With the development of intelligent voice interaction systems, sound source localization technology, as a key support, has been widely used in smart homes, in-vehicle voice systems, human-computer interaction, and industrial monitoring. By accurately estimating the direction of attack (DOA) of the sound source, the system can effectively perform advanced functions such as speech enhancement, beamforming, and sound source separation, thereby significantly improving the robustness of speech recognition and interaction. However, in practical applications, the performance of traditional sound source localization algorithms faces severe challenges due to reverberation effects, multipath propagation, and background noise interference. Currently, common sound source localization algorithms can be broadly classified into three types according to their principles: sound source localization algorithms based on time difference of arrival; sound source localization algorithms based on beamforming; and sound source localization algorithms based on high-resolution spectral estimation.
[0003] 1. Time Difference of Arrival (TDOA) is one of the most classic and widely used methods for sound source localization. Its core idea is to estimate the location of the sound source by measuring the time difference between the arrival times of signals at different microphones. The TDOA method typically involves two steps: first, obtaining time difference information through a time delay estimation algorithm; second, using this time difference information and sensor positions to perform localization calculations. The advantages of the TDOA algorithm are its simple structure and high computational efficiency; however, in complex environments, its robustness and accuracy are significantly affected by noise, echoes, and other interference.
[0004] 2. Beamforming refers to the process of weighting and summing the signals from each microphone in a microphone array to form a beam pointing in a specific direction, thereby increasing the signal strength in that direction and suppressing noise and interference from other directions. However, it only sharpens the peak values, and while it is effective in separating interference sources from a fixed direction, it is prone to misjudging reverberation, leading to incorrect DOA estimation results. In cases of severe reverberation or multiple sound sources, it is susceptible to problems such as spectral peak diffusion and spurious peak interference, making it difficult to accurately distinguish the directions of adjacent sound sources.
[0005] 3. High-resolution spectrum-based sound source localization algorithms primarily obtain accurate sound source location information through spectrum analysis. Unlike traditional time-delay estimation-based methods, these methods focus on the spectral characteristics of the signal, enabling them to distinguish minute signal changes from noise and reverberation, thus achieving higher sound source localization accuracy, but with a larger computational load.
[0006] Therefore, there is an urgent need for a sound source localization method that can accurately filter out signal components dominated by direct sound and reasonably weight and fuse them based on their confidence levels, so as to improve robustness and resolution in complex sound field environments. Summary of the Invention
[0007] The purpose of this invention is to provide a sound source direction estimation method based on a small linear microphone array, which can accurately identify high-confidence time-frequency units dominated by direct sound, and construct a direction spectrum through weighted fusion, thereby improving the localization accuracy of multiple sound sources and reducing the false peak error caused by reverberation interference.
[0008] This invention provides a sound source direction estimation method based on a small linear microphone array, which mainly includes the following steps:
[0009] (1) Signal acquisition and transformation. Multi-channel speech signals were acquired using a small linear microphone array and short-time Fourier transform was performed to obtain time-frequency domain data.
[0010] (2) Construct SRP-PHAT spectra. Calculate narrowband SRP spectra based on phase transformation on each TF element.
[0011] (3) LSDD screening mechanism. The similarity between each TF unit and all scanning angle steering vectors is calculated, and the confidence level is evaluated based on the relative spectral peak intensity. High confidence TF units dominated by direct sound are screened.
[0012] (4) Confidence-weighted fusion. The retained TF units are weighted according to their confidence levels and their histograms are superimposed to construct a weighted directional spectrum.
[0013] (5) Main peak extraction. Detect the position of the main peak in the directional spectrum and output the corresponding sound source direction angle.
[0014] The innovations of this invention compared to existing technologies include:
[0015] 1) This invention introduces the Local Spatial Domain Distance (LSDD) mechanism for the first time under the SRP-PHAT algorithm framework to calculate the similarity between time-frequency units and steering vectors in different directions. By measuring spatial similarity, it effectively identifies high-confidence TF units dominated by direct sound and suppresses the pollution of the directional spectrum by reverberation interference from the source.
[0016] 2) This invention proposes relative spectral peak intensity as a confidence index, and selects time-frequency units with high positioning reliability based on the sharpness of the spectral peaks. This realizes an adaptive screening mechanism based on statistical characteristics, which overcomes the interference retention problem caused by unreasonable full-frequency weighting or threshold selection in traditional methods.
[0017] 3) This invention introduces a confidence-weighted histogram strategy during the fusion process. By weighting and integrating the DOA estimation results of each high-confidence TF unit, a weighted directional spectrum is constructed. This method significantly improves the focusing and resolution of directional spectrum peaks, effectively distinguishes nearby sound sources, and reduces the risk of spurious peak interference.
[0018] The beneficial effects of this invention are:
[0019] 1) This invention effectively eliminates TF units dominated by reverberation wakes through the LSDD screening mechanism.
[0020] 2) The confidence weighting strategy of this invention improves the focusing of spectral peaks and enhances the ability to resolve multiple sound sources.
[0021] 3) Compared with the SRP-PHAT and WH-PHAT algorithms, the present invention has a lower MAE and a higher success rate in dual-source detection in simulation and actual measurement.
[0022] 4) This invention achieves high-precision sound source direction estimation in small microphone arrays, which is convenient for embedded deployment. Attached Figure Description
[0023] Figure 1 The overall flowchart of the LSDD-WH-PHAT algorithm described in this invention is provided for this invention.
[0024] Figure 2 This is a schematic diagram of a linear microphone array signal receiving model provided by the present invention.
[0025] Figure 3 This is a schematic diagram of the dual sound source arrangement in the simulation experiment provided by the present invention.
[0026] Figure 4 A comparison diagram of the directional spectra of different algorithms provided for this invention. Detailed Implementation
[0027] The specific embodiments of the present invention will be clearly and accurately described below with reference to the accompanying drawings.
[0028] This invention introduces an LSDD spatial similarity analysis mechanism and a relative spectral peak confidence evaluation model, retaining only high-quality time-frequency units dominated by direct sound for directional spectrum construction, significantly reducing the cumulative effects of reverberation and multipath interference. Simultaneously, by combining a confidence-weighted histogram fusion strategy, it effectively improves the sharpness and resolution of the main lobe of the directional spectrum. This method is suitable for small microphone array structures with a limited number of array elements, exhibiting low computational complexity and good engineering feasibility. It is particularly suitable for deployment in in-vehicle voice systems, smart terminals, and low-power embedded platforms, demonstrating excellent positioning accuracy and robustness in real-world complex acoustic scenarios.
[0029] like Figure 1 As shown, this embodiment of the invention provides a sound source direction estimation method based on a small linear microphone array, including the following steps:
[0030] (1) Signal Acquisition and Transformation. A small linear microphone array is used to acquire multi-channel speech signals, and a short-time Fourier transform is performed to obtain time-frequency domain data. Step (1) of the embodiment is implemented as follows:
[0031] like Figure 2 As shown, assuming the linear array has M microphones, let c be the speed of sound, and θ be the direction of the sound wave source. Selecting the first microphone in the array as the reference microphone, the propagation time delay difference Δτ between the sound wave reaching the m-th microphone and the reference microphone is... m (θ) can be expressed as:
[0032]
[0033] The signal received by the m-th microphone in the array can be represented as:
[0034]
[0035] Among them, h m,q (t) is the acoustic transfer function (ATF) from the q-th sound source to the m-th microphone, including the direct path and reverberation effects; s q (t) represents the signal from the q-th sound source; n m (t) represents the additive noise of the m-th microphone.
[0036] Frame-by-frame windowing is applied to equation (2) to obtain the signal in the STFT domain, which can be represented by vector notation as follows:
[0037]
[0038] Where z(l,k)=[Z1(l,k) … Z M (l,k)] T h q (k)=[H q,1 (k) … H q,M (k)] T ; n(l,k)=[N1(l,k) … N M (l,k)] T In the formula, l is the time frame and k is the frequency.
[0039] Define the steering vector in z(l,k), which is the direct path in ATF, as follows:
[0040] d θ (k)=[D1(k),D2(k),…,D M (k)] T (4)
[0041]
[0042] Where K is the number of frequency points of the STFT; i is the imaginary unit.
[0043] (2) Constructing the SRP-PHAT spectrum. Narrowband SRP spectra based on phase transform are calculated on each TF unit. Step (2) of the embodiment is implemented as follows:
[0044] For a pair of microphone sequences (m,n), the time difference of arrival of their received signals is defined as Δτ. mn (θ)=τ m (θ)-τ n (θ), combined with formula (3), the cross-correlation function of the two microphone signals is calculated as follows:
[0045]
[0046] This cross-correlation function, by summing all frequency components, comprehensively evaluates the correlation between signals, reflecting the similarity of signals received by two microphones under different time delays given an azimuth angle θ. The azimuth angle θ-related localization spectrum is obtained by summing the cross-correlation functions of all microphone pairs, i.e.:
[0047]
[0048] To effectively suppress the interference of reverberation effects and environmental noise on DOA angle estimation, the narrowband SRP-PHAT algorithm employs a phase transform weighting (PHAT) method to preprocess the frequency domain signal. The weighting function is defined as follows:
[0049]
[0050] After matching and weighting the signal, the signal amplitude information is eliminated, resulting in the narrowband SRP-PHAT localization spectrum:
[0051]
[0052] PHAT weighting significantly reduces non-stationary amplitude interference caused by reverberation wakes, making the localization spectrum construction rely solely on the phase consistency characteristics between signals. This processing mechanism transforms the traditional time delay estimation problem into a pure phase difference matching optimization problem, effectively overcoming the spurious peak phenomenon caused by amplitude distortion in reverberant environments, thereby improving the robustness of localization estimation in complex acoustic environments.
[0053] (3) LSDD screening mechanism. The similarity between each TF unit and all scanning angle steering vectors is calculated, and the confidence level is assessed based on the relative spectral peak intensity. High-confidence TF units dominated by direct sound are then screened. Step (3) of the embodiment is implemented as follows:
[0054] The LSDD spectrum is calculated on a predefined DOA grid Θ. The formula for calculating the LSDD spectrum is as follows:
[0055]
[0056] Where d(·,·) is a function that measures the similarity between two vectors.
[0057] In this way, the similarity between the time-frequency unit l and the array response vector of a unit amplitude wave transmitted from the direction at frequency k can be reflected. When S l,k (Θ j The larger the value, the stronger the signal component from that direction in that time-frequency unit.
[0058] After calculating the LSDD spectrum, to further measure the relative intensity of the highest peaks in the spectrum and thus more accurately determine the degree to which the time-frequency unit is dominated by direct sound, the LSDD algorithm introduces the concept of relative peak intensity. The formula for calculating relative peak intensity is:
[0059]
[0060] Here, j1 and j2 are the indices corresponding to the maximum and minimum values in the LSDD spectrum.
[0061] From equation (11), we can see that R l,k This represents the ratio between the height of the highest peak and the average spectral height after excluding the peak. The larger this value, the more prominent the highest peak in the LSDD spectrum, which means that in this time-frequency unit, the direct sound accounts for a higher proportion of all signal components, and the less it is affected by reverberation.
[0062] Finally, by setting a threshold R th To filter time-frequency units. The specific filtering criteria are:
[0063]
[0064] That is, when the relative peak intensity of a certain time-frequency unit is greater than the threshold, the time-frequency unit is considered to be dominated by direct sound and is selected into the set. middle.
[0065] The choice of threshold directly affects the quality and quantity of the screening results, and it needs to be set reasonably according to the specific application scenario and system parameters. If the threshold is set too high, although it can screen out time-frequency units that are highly dominated by direct sound, it may result in too few screened units and the loss of some useful information; conversely, if the threshold is set too low, it may retain too many time-frequency units that are greatly affected by reverberation, and it may not be able to effectively remove reverberation interference. In practical applications, typically 10% of the entire TF unit is screened.
[0066] (4) Confidence-weighted fusion. The retained TF units are weighted according to their confidence levels, and their histograms are superimposed to construct a weighted directional spectrum. Step (4) of the embodiment is implemented as follows:
[0067] After filtering out the time-frequency units dominated by direct sound using the LSDD algorithm, a set of time-frequency units dominated by direct sound was obtained. These high-quality time-frequency cells are processed using a weighted histogram method to achieve more accurate DOA estimation. This process fully utilizes the reliable source direction information contained in the screened time-frequency cells. First, for each time-frequency cell in the set, the quality metric corresponding to each angle is calculated. The formula for calculating quality metrics is:
[0068]
[0069] The quality metric reflects the relative dominance of the current narrowband DOA estimate among all possible directions. The larger the value, the more reliable the narrowband DOA estimate is in the current time-frequency unit, meaning that the signal strength in that direction accounts for a higher proportion among all candidate directions Θ.
[0070] By using these narrowband DOAs and their corresponding quality metrics, a weighted histogram is fused to obtain a more accurate broadband DOA estimate.
[0071]
[0072] Where δ(·) is the discriminant function (when If the value is 1, then the value is 0.
[0073] The summation operation applies only to elements in the time-frequency cell set filtered by the LSDD algorithm. Equation (14) is a weighted histogram-based positioning spectrum that integrates the narrowband DOA information of all relevant time-frequency cells and their corresponding quality metrics. This allows the highly reliable narrowband DOA to have a greater impact in the calculation of the broadband positioning spectrum.
[0074] (5) Main Peak Extraction. Detect the position of the main peak in the directional spectrum and output the corresponding sound source direction angle. Step (5) in the embodiment is implemented as follows:
[0075] The final broadband DOA estimate can be obtained by finding the peak value of the broadband positioning spectrum in equation (14).
[0076]
[0077] Because only time-frequency units dominated by direct sound, selected by the LSDD algorithm, are used in the calculation process, the interference of reverberation is reduced. Therefore, this weighted histogram fusion method can estimate the DOA of the sound source more accurately in complex reverberation environments compared to the traditional SRP-PHAT algorithm.
[0078] To verify the effectiveness of the algorithm in different reverberation environments, an indoor impulse response was generated in MATLAB based on the IMAGE reverberation model and convolved with clean speech from the TIMIT dataset. Simultaneously, 20dB Gaussian white noise was added to the signal received by the microphone array in a simulated real-world noise environment. A linear 4-element small microphone array was used, with each microphone spaced 0.033m apart. The sound source was fixed within the array plane, 1m from the array center. Figure 3 As shown. The experiment assumes a fixed reverberation time RT for the two-source DOA estimation. 60 = 300ms, the angles of the two sound sources are 30° and 40° respectively. For example Figure 4 As shown, the experimental results of three different algorithms in DOA estimation for dual sound sources are presented.
[0079] Depend on Figure 4 It is evident that in complex dual-source reverberation environments, the DOA estimation performance of different algorithms exhibits significant differences. The traditional SRP-PHAT algorithm forms a broad peak in the directional spectrum within the 28°-42° range, with the main lobe center shifted to 36°, indicating its inability to effectively distinguish between the two sound sources and its susceptibility to reverberation diffusion effects. While the WH-PHAT algorithm generates a double peak near 30° and 40°, initially achieving separation of the two sound sources, it exhibits an additional significant peak in the 30°-40° range, reflecting the dispersion of the localization spectrum caused by low-reliability time-frequency units. In contrast, the LSDD-WH-PHAT algorithm effectively eliminates time-frequency units dominated by non-direct sound through the LSDD time-frequency filtering mechanism, and its directional spectrum displays a clear double-peak structure near 30° and 40° without redundant peak interference.
[0080] The specific embodiments described herein are merely illustrative of the spirit of the invention. Those skilled in the art to which this invention pertains may make various modifications or additions to the described specific embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.
Claims
1. A method for estimating the direction of a sound source based on a small linear microphone array, characterized in that, include: (1) Acquire multi-channel acoustic signals through a small linear microphone array and perform short-time Fourier transform (STFT) on the signals to obtain time-frequency domain data; (2) Constructing narrowband spatial response power spectrum based on SRP-PHAT method; (3) On the preset azimuth scanning grid, the Local Spatial Domain Distance (LSDD) metric method is used to calculate the similarity between each time-frequency cell and the steering vector in each direction; (4) Based on the relative spectral peak intensity, select high-confidence time-frequency units dominated by direct sound and construct a confidence screening set; (5) Perform confidence weighting on each unit in the selected set to construct a weighted histogram broadband directional spectrum; (6) By detecting the position of the main peak of the directional spectrum, the corresponding sound source direction angle (DOA) is output.
2. The sound source direction estimation method based on a small linear microphone array according to claim 1, characterized in that: In step (3), the LSDD similarity calculation method is to use the cosine similarity between each time-frequency unit and each azimuth directional steering vector as the similarity metric.
3. The sound source direction estimation method based on a small linear microphone array according to claim 1, characterized in that: In step (4), the relative spectral peak intensity is calculated using the following formula: In the formula, j1 and j2 are the indices corresponding to the maximum and minimum values in the LSDD spectrum, respectively.
4. The sound source direction estimation method based on a small linear microphone array according to claim 1, characterized in that: In step (5), the confidence weighting method is to weight the DOA estimate of each time-frequency unit according to the relative peak intensity of its LSDD spectrum. The larger the weight, the more significant its influence on the final positioning spectrum.
5. The sound source direction estimation method based on a small linear microphone array according to claim 1, characterized in that: In step (6), the main peak of the directional spectrum is detected by the maximum value of the weighted histogram, and the corresponding scanning angle is the estimated value of the sound source direction.
6. The sound source direction estimation method based on a small linear microphone array according to claim 1, characterized in that: It is applied to sound source direction estimation in scenarios such as in-vehicle voice interaction, smart home voice recognition, conference sound pickup, human-computer interaction, or industrial monitoring.