Sub-band Acoustic Echo Cancellation for Active Speaker Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In two-way audio conferencing systems, accurately detecting an active speaker is challenging due to contamination by non-stationary background noise and echo, which can lead to incorrect identification of participants and degrade audio performance.
Innovation Solution
The system employs an acoustic echo canceller (AEC) to analyze audio metrics in sub-band domains, decoupling noise and echo, and uses machine learning to weight these metrics for accurate active speaker detection, incorporating a hysteresis model to stabilize speaker status and prevent toggling between active and non-active states.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional energy level and voice activity detection is used, then the detection method is simple, but the accuracy deteriorates due to background noise and echo contamination
Solution Approach 1:
The audio signal is divided into multiple sub-bands, and detection is performed independently in each sub-band domain. This segmentation allows the system to analyze different frequency components separately, improving accuracy by isolating speech from noise and echo in specific frequency ranges while maintaining manageable computational complexity through localized processing.
Solution Approach 2:
An acoustic echo canceller is introduced as an intermediary component that processes the audio signal before active speaker detection. The AEC model estimates and removes echo components, providing a cleaned signal to the detection algorithm. This intermediary processing step significantly improves detection accuracy in two-way conferencing systems without requiring complete system redesign.
2Measurement precision
If acoustic echo canceller based detection is used, then the accuracy improves by decoupling noise and echo, but the computational complexity increases
Solution Approach 1:
The system applies echo cancellation and detection processing selectively in sub-band domains rather than processing the entire audio spectrum uniformly. By focusing computational resources on specific sub-bands where speech activity is detected, the system achieves high accuracy while reducing overall computational power consumption compared to full-band processing.
Solution Approach 2:
The audio frequency spectrum is segmented into multiple sub-bands, allowing parallel processing of different frequency components. This segmentation distributes computational load across multiple smaller processing units, reducing the computational power required for each individual processing step while maintaining overall detection accuracy through comprehensive multi-band analysis.
3Stability of the object's composition
If hysteresis model is applied, then the speaker status stability improves, but the response time to detect state changes increases
Solution Approach 1:
The hysteresis model uses dynamic threshold adjustment based on the current state and history of speaker detection. The thresholds for transitioning between active and non-active states are adjusted dynamically to balance stability and responsiveness. This dynamic approach allows the system to maintain stability during transient fluctuations while responding promptly to genuine speaker state changes.
Data Source
AI summary
Systems, methods, and devices are disclosed for detecting an active speaker in a two-way conference. Real time audio in one or more sub band domains are analyzed according to an echo cancellor model. Based on the analyzed real time audio, one or more audio metrics are determined from output from an acoustic echo cancellation linear filter. The one or more audio metrics are weighted based on a priority, and a speaker status is determined based on the weighted one or more audio metrics being analyzed according to an active speaker detection model. For an active speaker status, one or more residual echo or noise is removed from the real time audio based on the one or more audio metrics.


