Sound Source Localization Using Sparse Interaural Time Difference Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound source localization techniques for robots face challenges in accurately localizing sound sources due to environmental changes and the complexity of measuring impulse responses, especially when using microphone arrays that mimic human hearing, which are sensitive to platform variations and interference from facial structures.
Innovation Solution
A sound source localization system utilizing sparse coding and a self-organized map (SOM) to extract sparse interaural time differences (SITDs) from microphone signals, allowing for adaptation to various platforms and environments without the need for repeated impulse response measurements, and employing a neural network model inspired by human auditory processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If HRTF-based sound source localization is used to account for platform influences, then localization accuracy is improved, but measurement complexity increases due to the need for impulse response measurements in dead rooms for each platform configuration
Solution Approach 1:
The patent creates a virtual model (copy) of the head and ear structures that replicates the acoustic characteristics of the physical platform. This virtual model is used to generate HRTF data through simulation, eliminating the need for physical impulse response measurements in anechoic chambers while maintaining localization accuracy.
Solution Approach 2:
The patent replaces the mechanical measurement system (physical impulse response measurements in dead rooms) with a computational simulation system. By using numerical methods to solve the acoustic wave equation, the system obtains HRTF data without requiring complex physical measurement setups.
2Ease of operation
If traditional sound source localization systems are used, then they can operate in simple environments, but they require program modifications to adapt to environmental changes
Solution Approach 1:
The patent implements a dynamic adaptation mechanism where the system continuously learns from environmental acoustic characteristics and adjusts its localization parameters accordingly. This allows the system to adapt to different environments (rooms, outdoor spaces, etc.) without requiring manual reconfiguration or program modifications.
Solution Approach 2:
The system performs self-calibration by automatically characterizing the acoustic environment and adjusting its HRTF models accordingly. This self-service capability eliminates the need for external intervention or complex setup procedures when deploying the system in new environments.
3Measurement precision
If microphone arrays with specific structures are used for sound source localization, then localization can be achieved in ideal conditions, but performance degrades when platform influences and facial structures interfere with sound signal flow
Solution Approach 1:
The patent creates detailed virtual models of the head, ears, and facial structures that accurately replicate how these physical structures modify sound signals. By simulating the acoustic interactions between sound waves and these structures in the virtual model, the system compensates for their interfering effects and maintains localization accuracy.
Data Source
AI summary
A sound source localization system includes a plurality of microphones for receiving a signal as an input from a sound source; a time-difference extraction unit for decomposing the signal inputted through the plurality of microphones into time, frequency and amplitude using a sparse coding and then extracting a sparse interaural time difference (SITD) inputted through the plurality of microphones for each frequency; and a sound source localization unit for localizing the sound source using the SITDs. A sound source localization method includes receiving a signal as an input from a sound source; decomposing the signal into time, frequency and amplitude using a sparse coding; extracting an SITD for each frequency; and localizing the sound source using the SITDs.


