Adaptive Audio Target Selection for Echo Latency Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional double-talk detection in electronic devices suffers from latency and high computational costs due to the need for long analysis windows to detect echo latency, especially in systems with variable delays and wireless loudspeakers, which affects the accuracy of audio processing during communication sessions.
Innovation Solution
The implementation of a system that uses a combination of detectors, including a least mean squares (LMS) adaptive filter and voice activity detection, to determine current system conditions by distinguishing between near-end single-talk, far-end single-talk, and double-talk conditions, and dynamically selects target signals based on signal quality metrics to adaptively manage echo cancellation and noise suppression.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If long analysis windows are used to detect echo latency, then measurement precision is improved, but loss of time and productivity deteriorate due to increased latency and computational costs
Solution Approach 1:
The patent divides the audio signal analysis into multiple shorter analysis windows instead of using a single long analysis window. The system processes audio data in segmented frames, allowing for faster individual analysis while maintaining overall detection accuracy through cumulative processing of multiple segments.
Solution Approach 2:
The system performs preliminary processing of audio data by pre-calculating features and maintaining buffers of recent audio samples. This preliminary action enables faster echo latency detection when needed, as the foundation work is already completed in advance rather than waiting for long analysis windows to accumulate.
2Measurement precision
If long analysis windows are used to detect echo latency, then measurement precision is improved, but computational cost increases
Solution Approach 1:
The patent divides the audio signal analysis into multiple shorter analysis windows instead of using a single long analysis window. The system processes audio data in segmented frames, allowing for faster individual analysis while maintaining overall detection accuracy through cumulative processing of multiple segments.
Solution Approach 2:
The system performs partial analysis on each short window, focusing only on the most critical features for echo latency detection rather than complete spectral analysis. This partial action approach reduces computational burden per window while maintaining sufficient accuracy for the application.
3Reliability
If aggressive audio processing is applied to suppress unwanted signals, then reliability is improved, but harmful factors increase due to distortion of local speech
Solution Approach 1:
The patent implements dynamic adjustment of audio processing intensity based on detected system conditions. The system adapts the aggressiveness of echo cancellation and noise suppression in real-time, using milder processing when local speech is detected and more aggressive processing when only remote speech is present, thereby reducing speech distortion while maintaining reliability.
Solution Approach 2:
The system uses feedback from double-talk detection and quality metrics to continuously adjust processing parameters. By monitoring the presence of local speech and the effectiveness of current processing, the system dynamically tunes the balance between echo suppression and speech preservation, preventing excessive distortion.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach reduces latency and computational requirements, improving the accuracy of audio processing by effectively isolating local speech and suppressing unwanted signals in various system conditions, enhancing the overall quality of voice commands and communication sessions.
Implementation Method 1
The implementation of a system that uses a combination of detectors, including a least mean squares (LMS) adaptive filter and voice activity detection
Data Source
AI summary
A system configured to improve audio processing by adaptively selecting target signals based on current system conditions. For example, a device may select a target signal based on a highest signal quality metric when only the local speech is present (e.g., during near-end single-talk conditions), as this maximizes an amount of energy included in the output audio signal. In contrast, the device may select the target signal based on a lowest signal quality metric when only the remote speech is present (e.g., during far-end single-talk conditions), as this minimizes an amount of energy included in the output audio signal. In addition, the device may track positions of the local speech and the remote speech over time, enabling the device to accurately select the target signal when both local speech and remote speech is present (e.g., during double-talk conditions).


