Microphone Array Sound Localization Using Spatial Feature Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound localization techniques in microphone arrays face challenges in accurately determining the direction of sound in noisy environments, particularly due to increased misjudgment probability and decreased detection success rate with music or babble noise, and existing echo cancellation methods do not adequately address these issues.
Innovation Solution
A device and method incorporating a spatial feature generator, voice detector, and angle retriever to generate and process spatial feature signals from a microphone array, which includes determining the existence of sound sources and voices, and outputs an estimated angle signal for sound direction based on candidate angle signals and voice detection results.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a voice detection technique is used to improve angle estimation accuracy, then the accuracy of angle estimation in noisy environment is improved, but the misjudgement probability increases as the strength of music noise or babble noise increases
Solution Approach 1:
The patent segments the sound source detection process into multiple independent spatial feature signals (GCC-PHAT, GCC-Roomba, SRP-PHAT, SRP-Roomba) corresponding to different microphone pairs. Each spatial feature signal processes specific microphone signal pairs independently, allowing the system to evaluate multiple candidate directions and select the most reliable one, thereby reducing misjudgment in noisy environments
Solution Approach 2:
The patent implements a feedback mechanism where the system detects sound sources using multiple spatial feature signals, evaluates the reliability of each detection, and adjusts the selection of candidate angle signals based on the detected sound source information. The angle retriever uses the sound source detection result to selectively retrieve and output candidate angle signals, creating a closed-loop system that continuously improves accuracy while monitoring for misjudgments
2Measurement precision
If a voice detection technique is used to determine direction of voice, then the direction estimation is enhanced, but the detection success rate decreases in a noisy environment
Solution Approach 1:
The patent divides the voice detection process into multiple parallel spatial feature signal channels, each processing different microphone signal pairs. By segmenting the detection across multiple independent features (GCC-PHAT, GCC-Roomba, SRP-PHAT, SRP-Roomba), the system maintains detection capability even when individual channels fail in noisy conditions
Solution Approach 2:
The patent performs preliminary sound source detection using multiple spatial feature signals before final angle retrieval. The system pre-calculates candidate angle signals from all spatial features and stores them for later retrieval based on detection results, allowing the system to prepare multiple potential answers before needing to commit to a final direction estimate
3Measurement precision
If multiple spatial feature signals are generated from microphone array to improve sound localization, then the direction determination is enhanced, but the device complexity increases
Solution Approach 1:
The patent uses the same four spatial feature signal types (GCC-PHAT, GCC-Roomba, SRP-PHAT, SRP-Roomba) across all microphone pairs, making the processing methodology universal and reusable. This multi-functional approach allows the system to handle different microphone configurations and noise conditions using the same set of spatial features, reducing the need for separate specialized processing paths
Data Source
AI summary
A device for sound localization includes a spatial feature generator, a voice detector, an angle selector, and an angle retriever. The spatial feature generator generates M spatial feature signals according to signals of N microphones of a microphone array. The voice detector generates at least one voice detection signal according to at least one of the signals of the N microphones. The angle selector outputs a candidate angle signal according to the M spatial feature signals to indicate a candidate direction of sound. The angle retriever generates a sound detection result according to the M spatial feature signals to indicate whether any sound source exists, and then outputs an estimated angle signal indicative of a direction of sound according to the sound detection result, the at least one voice detection signal, and the candidate angle signal.


