Sound Source Localization Using Interaural Time Difference Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source localization methods in humanoid robotics are prone to noise and reverberation, and are computationally demanding, making them unsuitable for real-time implementation in embedded systems.
Innovation Solution
A method using a set of microphones to calculate generalized intercorrelations and directed response power, considering vectors of interaural time differences, to estimate the location of a sound source, with enhanced robustness to noise and reverberation by including theoretically inadmissible vectors in the optimization process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spectral estimation techniques are used for sound source localization, then localization accuracy is improved, but computing power requirements increase significantly
Solution Approach 1:
The patent extracts only the essential information needed for localization by using generalized cross-correlation to identify time differences of arrival, rather than performing full spectral estimation. This extraction approach retains localization accuracy while eliminating computationally intensive spectral analysis, achieving the desired balance between precision and processing power requirements.
2Power
If PHAT-GCC method is used for sound source localization, then computational load is reduced, but robustness to correlated noise and reverberation deteriorates
Solution Approach 1:
The patent modifies the PHAT-GCC approach by changing the parameter space from continuous time differences to discrete directional bins. This parameter transformation allows the system to maintain computational efficiency while improving robustness through a two-stage verification process that filters out spurious detections and validates results against acoustic geometry constraints, thereby addressing both computational and reliability requirements.
3Reliability
If SRP-PHAT method is used for sound source localization, then robustness to noise is improved, but sensitivity to reverberation increases
Solution Approach 1:
The patent segments the localization process into two distinct stages: first identifying candidate directions using generalized cross-correlation, then verifying these candidates using acoustic geometry constraints and Chasles' relation checks. This segmentation allows the system to benefit from noise robustness of SRP-PHAT while using the verification stage to reject reverberation-induced false positives, thereby resolving the contradiction between noise robustness and reverberation sensitivity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The method provides improved immunity to noise and reverberation while being computationally lightweight, enabling real-time sound source localization in humanoid robots.
Implementation Method 1
capture sound signals from a sound source to be located using a set of at least three microphones
Implementation Method 2
calculate a generalized intercorrelation of the captured sound signals, said calculation being carried out for a plurality of values of a delay - called interaural time difference - between said signals
Implementation Method 3
consists of synthesizing an adjustable acoustic beam by adding the signals picked up by the different microphones to which a variable time shift has been applied
Data Source
Figure 1
Figure 2
Figure 3A~3B
AI summary
The invention relates to a method for locating a sound source by maximizing a directed response strength calculated for a plurality of vectors of the interauricular time differences forming a set (E) that includes: a first subset (E1) of vectors compatible with sound signals from a single sound source at an unlimited distance from said microphones; and a second subset (E2) of vectors that are not compatible with sound signals from a single sound source at an unlimited distance from said microphones. Each vector of said first subset is associated with a direction for locating the corresponding single sound source, and each vector of said second subset is associated with the locating direction associated with a vector of said first subset closest thereto according to a predefined metric. The invention also relates to a humanoid robot including: a set of at least three microphones (M1, M2, M3, M4), preferably arranged on a surface higher than the head thereof; and a processor (PR) for implementing one such method.