Sound Source Localization Using Coherence-to-Diffuseness Ratio Mask
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source localization methods, such as cross-correlation-based methods, suffer from reduced accuracy in environments with noise and echo, as they fail to effectively distinguish between direct and echo components, leading to inaccurate sound source direction estimation.
Innovation Solution
A sound source localization method and apparatus that utilize a diffuseness mask generated by the coherence-to-diffuseness power ratio (CDR) to preprocess input signals from multiple microphones, enhancing the robustness to noise and echo by applying a binarized mask and employing algorithms like GCC-PHAT or SRP-PHAT for accurate direction estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If cross-correlation-based sound source localization method is used, then it can estimate directions of multiple sound sources with stable performance, but the localization accuracy deteriorates in environments with noise or echo
Solution Approach 1:
The patent segments the sound field into distinct components (direct sound, echo, noise) by calculating the coherence-to-diffuseness ratio (CDR) for each time-frequency bin. This segmentation allows the system to process different sound components separately, applying appropriate weighting to enhance direct sound while suppressing echo and noise, thereby resolving the contradiction between stable multi-source estimation and accurate localization in noisy/echoic environments
Solution Approach 2:
The patent dynamically changes the parameter weights in the cross-correlation calculation based on the CDR values. By adjusting the weighting parameters according to the estimated diffuseness and coherence characteristics of the environment, the system adapts the cross-correlation method to maintain both stability and accuracy under varying acoustic conditions
2Device complexity
If traditional cross correlation method is applied, then the algorithm is simple and computationally efficient, but the performance becomes very inaccurate when additive noise distortion or echo components exist
Solution Approach 1:
The patent introduces CDR (coherence-to-diffuseness ratio) as an intermediary parameter that mediates between the simple cross-correlation method and the complex noise/echo conditions. The CDR calculation serves as an intermediate processing step that provides environmental characteristics information, enabling the system to adjust the cross-correlation weighting without requiring complex noise removal algorithms, thus maintaining relative simplicity while improving accuracy
3Measurement precision
If noise removal algorithms are applied to improve localization accuracy in noisy environments, then the accuracy improves, but a large amount of data and computation are demanded
Solution Approach 1:
The patent performs preliminary action by calculating the CDR and diffuseness mask before applying the cross-correlation-based localization algorithm. This preliminary estimation of environmental characteristics allows the system to pre-adjust the weighting parameters, avoiding the need for computationally intensive iterative noise removal processes during the actual localization computation, thus achieving improved accuracy with reduced computational demand
Data Source
AI summary
Provided is a sound source localization method including steps of: (a) receiving a mixed signal of a target sound source signal and noise and echo signals through multiple microphones including at least two microphones; (b) generating a binarized mask based on a diffuseness by using a coherence-to-diffuseness ratio CDR, which is information on the target sound source and the noise source, by using the input signal; (c) pre-processing an input signal to multiple microphones by using the generated binarized mask; and (d) performing a predetermined algorithm such as the GCC-PHAT or the SRP-PHAT on the pre-processed input signal to estimate a direction of the target sound source.


