Processor Onset Detection for Speaker Localization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing DOA estimation methods face challenges in accurately localizing a target speaker in noisy, reverberant environments with low Signal to Noise Ratios (SNRs), particularly when the target speaker is not the dominant sound source, leading to reduced DOA resolution and increased detection delays.
Innovation Solution
A processor that applies localization algorithms to determine source directions and onset detection to attribute scores, providing a direction-output-signal for a beamformer to focus on the target speaker, even if they are not dominant, using a combination of localization and onset detection algorithms to enhance wake word detection in challenging acoustic conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional DOA estimation methods are used to localize the dominant sound source, then the detection is robust in noisy environments, but the target speaker cannot be localized accurately when they are not the dominant source
Solution Approach 1:
The patent segments the sound source detection process into two independent components: (1) onset detection that identifies sudden acoustic events regardless of dominance, and (2) DOA estimation that localizes identified onsets. This segmentation allows the system to detect non-dominant speakers by focusing on onset characteristics rather than overall sound energy, resolving the contradiction between accuracy for non-dominant sources and reliability in noisy environments.
Solution Approach 2:
The patent introduces an intermediary onset detection mechanism that acts as a bridge between raw microphone signals and DOA estimation. The onset detector identifies acoustic events with sudden energy increases, and only these detected onsets are passed to the DOA estimation algorithm. This intermediary filtering enables accurate localization of non-dominant speakers by selecting relevant acoustic events while rejecting continuous background noise, thus improving measurement precision without sacrificing reliability.
2Measurement precision
If the system waits for sufficient speech statistics to improve DOA resolution, then the measurement precision improves, but the detection delay increases
Solution Approach 1:
The patent applies preliminary action by performing onset detection before DOA estimation. The onset detector proactively identifies acoustic events with sudden energy increases in real-time, and immediately triggers DOA estimation for these detected onsets. This preliminary identification eliminates the need to accumulate extended speech statistics, achieving high DOA resolution with minimal detection delay by acting on identified onsets rather than waiting for statistical convergence.
3Measurement precision
If Voice Activity Detection (VAD) and dereverberation are applied to improve speaker localization, then the measurement precision improves, but the device complexity increases
Solution Approach 1:
The patent extracts and removes the VAD and dereverberation processing steps from the speaker localization pipeline. Instead of applying these complex preprocessing operations, the system directly uses onset detection based on sudden energy increases followed by DOA estimation. This extraction of unnecessary components maintains speaker localization accuracy by focusing on onset characteristics while significantly reducing device complexity and processing requirements.
Data Source
AI summary
A processor is configured to receive a plurality of sounds signals from a respective plurality of microphones. One or more localization algorithms is applied to the received plurality of sound signals to determine a plurality of source directions. A plurality of tracking directions is determined based on current values for the determined plurality of source directions and previous values for the determined plurality of source directions. An onset detection algorithm is applied to the received plurality of sound signals, and in response to detecting an onset, the processor is configured to determine an onset direction that represents the direction from the microphones to the source of the detected onset. A score is attributed to each of the determined plurality of tracking directions based on any determined onset directions and the one of the determined plurality of tracking directions that has the highest score is provided as a direction-output-signal.


