Sound Source Localization Using Cross-Correlation Vectors
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing devices with only two microphones are limited in determining the location of sound sources across 360° due to the ambiguity in time difference of arrival (TDOA) measurements, which restricts noise reduction and audio enhancement capabilities to specific orientations and directions.
Innovation Solution
Implementing a machine learning model that processes cross-correlation vectors from two microphones, using a shallow neural network to distinguish between sound sources across 360° by considering all values in the cross-correlation vector, not just the peak value, thereby overcoming the limitations of traditional TDOA analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional TDOA analysis with two microphones is used, then device complexity is reduced, but measurement precision of sound source location is insufficient
Solution Approach 1:
The patent transforms the TDOA cross-correlation vector from time domain to frequency domain using Fourier transform, then applies phase extraction to obtain direction-dependent phase values. This parameter transformation enables 360° sound source localization with only two microphones by exploiting the phase information across different frequencies, resolving the ambiguity that limits traditional TDOA methods to specific orientations.
2Adaptability or versatility
If orthogonal beamforming is used for noise reduction, then background noise cancellation is improved, but adaptability to different sound source orientations is limited
Solution Approach 1:
The patent implements dynamic adaptation by using the machine learning model to predict sound source direction based on TDOA characteristics, then dynamically adjusting the beamforming weights in real-time according to the predicted direction. This enables the system to maintain optimal noise reduction performance across all 360° orientations without requiring complex pre-configured beamforming patterns for every possible direction.
Solution Approach 2:
The system employs feedback through the machine learning model that continuously analyzes TDOA measurements and provides directional predictions, which then feed back into the beamforming algorithm to adjust its parameters. This closed-loop feedback mechanism enables the beamforming system to adapt to sound sources from any orientation while maintaining computational efficiency.
3Measurement precision
If larger microphone arrays are deployed, then measurement precision of sound source location is improved, but device complexity and cost increase
Solution Approach 1:
The patent creates a virtual expanded array effect by using signal processing techniques to synthesize additional spatial information from the two physical microphones. Through frequency-domain transformation and phase analysis, the system generates directional cues that mathematically emulate what would be obtained from a larger microphone array, achieving high-precision 360° localization without the physical complexity of deploying multiple microphones.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Achieves greater than 95% accuracy in determining the direction of sound sources across a full 360° space without the need for additional microphones, reducing computational overhead and costs associated with larger microphone arrays.
Implementation Method 1
Many computing devices (e.g., laptops, tablets, smartphones, etc.) include at least one microphone to capture sound generated external to the device
Implementation Method 2
the ambiguity in time difference of arrival (TDOA) measurements
Data Source
AI summary
Methods, apparatus, systems, and articles of manufacture to detect the location of sound sources external to computing devices are disclosed. An apparatus, to determine a direction of a source of a sound relative to a computing device, includes a cross-correlation analyzer to generate a vector of values corresponding to a cross-correlation of first and second audio signals corresponding to the sound. The first audio signal is received from a first microphone of the computing device. The second audio signal is received from a second microphone of the computing device. The apparatus also includes a location analyzer to use a machine learning model and a set of the values of the vector to determine the direction of the source of the sound.


