Acoustic Speaker Localization via Modulation Domain SNR Weighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing acoustic speaker localization techniques face challenges in robustly estimating sound source location in high noise and reverberation conditions, particularly due to sensitivity to broadband noise and sub-optimal estimation of narrowband signal-to-noise ratio (SNR).
Innovation Solution
The method involves analyzing audio signals in the modulation domain, filtering modulator signals, and applying a weight mask based on signal-to-noise ratio (SNR) and probability of speech presence to enhance speech localization, using a combination of carrier and modulator signals to improve cross-correlation calculations and noise robustness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If GCC-PHAT method is used for TDOA estimation, then localization can be performed, but sensitivity to broadband noise increases
Solution Approach 1:
The patent applies non-uniform spectral weighting to the PHAT function, where different frequency bins are weighted differently based on their local signal-to-noise ratio characteristics. This allows the system to emphasize frequency regions with better noise immunity while de-emphasizing regions affected by broadband noise, thereby maintaining TDOA estimation accuracy while reducing noise sensitivity.
Solution Approach 2:
The patent modifies the PHAT weighting parameters by introducing frequency-dependent weighting factors that adapt to local noise conditions. By changing the weighting parameters dynamically across different frequency bins, the system optimizes the balance between utilizing phase information for localization and rejecting broadband noise interference.
2Reliability
If non-uniform PHAT weighting is applied to reduce noise sensitivity, then robustness against noise improves, but false TDOA may be generated in reverberation conditions
Solution Approach 1:
The patent segments the audio signal into multiple frequency bins and applies different weighting strategies to each bin based on local SNR estimates. This segmentation allows the system to identify and weight frequency regions that are less affected by reverberation, reducing the generation of false TDOA while maintaining noise robustness in other frequency regions.
Solution Approach 2:
The patent introduces an intermediary SNR estimation mechanism that acts as a mediator between the raw audio signals and the PHAT weighting function. This intermediary layer provides adaptive weighting factors that balance noise suppression and reverberation mitigation, preventing false TDOA generation while maintaining robustness against broadband noise.
3Reliability
If narrowband SNR estimation is used for weighting, then contribution of low SNR frequencies is reduced, but sub-optimal estimation degrades performance
Solution Approach 1:
The patent implements a feedback mechanism where the SNR estimation is continuously refined based on the observed performance in different frequency bins. The system uses the correlation results and signal characteristics to adjust the SNR estimates, creating a closed-loop system that improves localization accuracy while maintaining noise immunity through adaptive weighting.
Data Source
AI summary
A method, computer program product, and computing system for acoustic speech localization, comprising receiving, via a plurality of microphones, a plurality of audio signals. Modulation properties of the plurality of audio signals may be analyzed. Speech sounds may be localized from the plurality of audio signals based upon, at least in part, the modulation properties of the plurality of audio signals.


