Bandwidth Voice-Switching for Echo Reduction in Videoconferencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Videoconferencing systems face issues with imperfect echo cancellation, leading to false positives in sound source localization and audible echo due to far-end sound leakage, which affects active speaker detection and panning/zooming accuracy.
Innovation Solution
Implementing bandwidth voice-switching and subband-based techniques to attenuate or eliminate frequency subbands of far-end voice data that correspond to active near-end speech, thereby reducing far-end sound interference in microphone arrays and improving sound source localization accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If echo cancellation is used to remove far-end sound, then audible echo is reduced, but false positives in sound source localization increase due to imperfect cancellation
Solution Approach 1:
The patent segments the audio frequency spectrum into multiple frequency bins and processes each bin independently. By analyzing the spectral content in different frequency ranges separately, the system can identify and suppress far-end sound in specific frequency regions without affecting other frequencies, thereby maintaining sound source localization accuracy while reducing audible echo.
Solution Approach 2:
The patent applies different processing strategies to different frequency regions. Instead of uniform echo cancellation across all frequencies, the system identifies which frequency bins contain far-end sound and applies suppression only to those specific bins. This localized approach preserves the quality of near-end speech in unaffected frequency regions while eliminating echo in problematic regions.
2Measurement precision
If far-end sound is completely removed, then sound source localization accuracy improves, but audible echo may increase due to incomplete cancellation
Solution Approach 1:
The patent dynamically adjusts the amount of far-end sound suppression applied to each frequency bin based on real-time analysis. The system continuously monitors the spectral content and adapts the suppression level for each frequency region, increasing suppression where far-end sound is detected and reducing or eliminating suppression where near-end speech dominates. This dynamic adaptation resolves the contradiction by optimizing both localization accuracy and echo reduction in different frequency regions simultaneously.
Solution Approach 2:
The patent changes the spectral parameters of the audio signal by analyzing and modifying specific frequency bins. By identifying frequency regions where far-end sound dominates and applying targeted suppression to those bins while preserving other frequencies, the system achieves both improved sound source localization and reduced audible echo through parameter-based selective filtering.
3Object-affected harmful factors
If frequency subbands of far-end voice data are attenuated or eliminated, then far-end sound interference is reduced, but the complexity of audio processing increases
Solution Approach 1:
The patent divides the audio processing task into discrete frequency bins that can be independently analyzed and processed. This segmentation allows the system to apply simple comparison and suppression operations to each bin separately, making the overall complex task of far-end sound removal more manageable and computationally efficient through modular processing.
Data Source
AI summary
By-bandwidth voice-switching is performed during doubletalk between a near-end teleconference device and a far-end teleconference device. This may involve receiving far-end voice data from the far-end teleconference device. Near-end voice data is also received at the near-end teleconference device. Frequency subbands of the near-end voice data that having substantial energy are identified. Before playing the far-end voice data on one or more loudspeakers of the near-end teleconference device, frequency subbands from the far-end voice data frequency subbands thereof that correspond to the identified frequency subbands of the near-end voice data are attenuated and/or eliminated.


