Adaptive Audio Target Selection for Echo Latency Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional double-talk detection in electronic devices suffers from latency and high computational costs due to the need for long analysis windows to detect echo latency, especially in systems with variable delays and wireless loudspeakers, which affects the accuracy of audio processing during communication sessions.

Innovation Solution

The implementation of a system that uses a combination of detectors, including a least mean squares (LMS) adaptive filter and voice activity detection, to determine current system conditions by distinguishing between near-end single-talk, far-end single-talk, and double-talk conditions, and dynamically selects target signals based on signal quality metrics to adaptively manage echo cancellation and noise suppression.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If long analysis windows are used to detect echo latency, then measurement precision is improved, but loss of time and productivity deteriorate due to increased latency and computational costs

Engineering Contradiction:
Improveecho latency detection accuracyVSAvoiddetection latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent divides the audio signal analysis into multiple shorter analysis windows instead of using a single long analysis window. The system processes audio data in segmented frames, allowing for faster individual analysis while maintaining overall detection accuracy through cumulative processing of multiple segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary processing of audio data by pre-calculating features and maintaining buffers of recent audio samples. This preliminary action enables faster echo latency detection when needed, as the foundation work is already completed in advance rather than waiting for long analysis windows to accumulate.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If long analysis windows are used to detect echo latency, then measurement precision is improved, but computational cost increases

Engineering Contradiction:
Improveecho latency detection accuracyVSAvoidcomputational cost
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent divides the audio signal analysis into multiple shorter analysis windows instead of using a single long analysis window. The system processes audio data in segmented frames, allowing for faster individual analysis while maintaining overall detection accuracy through cumulative processing of multiple segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs partial analysis on each short window, focusing only on the most critical features for echo latency detection rather than complete spectral analysis. This partial action approach reduces computational burden per window while maintaining sufficient accuracy for the application.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If aggressive audio processing is applied to suppress unwanted signals, then reliability is improved, but harmful factors increase due to distortion of local speech

Engineering Contradiction:
Improveecho suppression effectivenessVSAvoidspeech distortion
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The patent implements dynamic adjustment of audio processing intensity based on detected system conditions. The system adapts the aggressiveness of echo cancellation and noise suppression in real-time, using milder processing when local speech is detected and more aggressive processing when only remote speech is present, thereby reducing speech distortion while maintaining reliability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from double-talk detection and quality metrics to continuously adjust processing parameters. By monitoring the presence of local speech and the effectiveness of current processing, the system dynamically tunes the balance between echo suppression and speech preservation, preventing excessive distortion.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach reduces latency and computational requirements, improving the accuracy of audio processing by effectively isolating local speech and suppressing unwanted signals in various system conditions, enhancing the overall quality of voice commands and communication sessions.

Implementation Method 1

The implementation of a system that uses a combination of detectors, including a least mean squares (LMS) adaptive filter and voice activity detection

Methodology Applied
Scientific EffectLeast Mean Squares (LMS) adaptive filtering:

Data Source

PatentUS10937441B1Beam level based adaptive target selection
Publication Date: 2021.03.02 AMAZON TECH INC
  • US10937441B1 patent drawing
  • US10937441B1 patent drawing
  • US10937441B1 patent drawing

AI summary

A system configured to improve audio processing by adaptively selecting target signals based on current system conditions. For example, a device may select a target signal based on a highest signal quality metric when only the local speech is present (e.g., during near-end single-talk conditions), as this maximizes an amount of energy included in the output audio signal. In contrast, the device may select the target signal based on a lowest signal quality metric when only the remote speech is present (e.g., during far-end single-talk conditions), as this minimizes an amount of energy included in the output audio signal. In addition, the device may track positions of the local speech and the remote speech over time, enabling the device to accurately select the target signal when both local speech and remote speech is present (e.g., during double-talk conditions).