ML-Assisted Acoustic Echo Cancellation for Residual Echo Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing acoustic echo cancelation (AEC) technologies struggle with accurately detecting and mitigating echo distortion, particularly in online conferencing systems, due to high reverberation and echo underestimation, which affects speech intelligibility.

Innovation Solution

Implementing a machine-learning assisted AEC system that uses a pre-trained model to identify unitary voice signals and apply aggressive echo cancellation modes based on residual echo thresholds, enhancing echo detection and cancellation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional AEC methods are used, then device complexity is kept low, but echo detection precision and speech intelligibility deteriorate due to high reverberation and echo underestimation

Engineering Contradiction:
Improveecho detection precisionVSAvoidAEC system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces a machine learning model as an intermediary component between the audio input and the traditional AEC processing pipeline. This ML model pre-processes the audio signal to identify and flag echo segments, which then guides the subsequent AEC processing. This intermediary step improves echo detection precision without requiring a complete redesign of the entire AEC system, thus managing complexity effectively.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent segments the AEC processing into distinct stages: (1) ML-based echo detection and segmentation of echo-containing audio frames, (2) traditional AEC processing applied selectively to identified echo segments, and (3) aggregation of processed segments. This segmentation allows the system to apply complex ML techniques only where needed rather than processing entire audio streams uniformly, improving precision while controlling overall complexity.

Inventive Principle:
Principle #1Segmentation

2Object-generated harmful factors

If aggressive echo cancellation modes are applied continuously, then echo removal effectiveness improves, but speech intelligibility deteriorates due to over-cancellation of valid speech signals

Engineering Contradiction:
Improveecho distortionVSAvoidspeech intelligibility
Core Design Contradiction:
Object-generated harmful factorsVSReliability

Solution Approach 1:

The patent implements dynamic adaptation of AEC processing intensity based on real-time echo detection results. The system transitions between different AEC modes (e.g., light, medium, aggressive cancellation) depending on the confidence level and severity of detected echo. This dynamic approach ensures aggressive cancellation is applied only when necessary, maintaining speech intelligibility while effectively removing echo distortion.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback loops where the output of the AEC processing is monitored and fed back to adjust subsequent processing parameters. The ML model continuously evaluates residual echo levels and adjusts the aggressiveness of cancellation in real-time, preventing over-cancellation of valid speech while ensuring sufficient echo removal. This closed-loop control maintains the balance between echo suppression and speech preservation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS12477070B1Machine-learning assisted acoustic echo cancelation
Publication Date: 2025.11.18 ZOOM COMMUNICATIONS INC
  • US12477070B1 patent drawing
  • US12477070B1 patent drawing
  • US12477070B1 patent drawing

AI summary

Example methods and systems provide machine-learning assisted acoustic echo cancellation (AEC). The AEC can be used, as an example, to improve the audio quality for online audio and video conferences. A system according to this disclosure includes a pre-trained, machine-learning, AI model designed to detect, in real time, a unitary voice signal, or a signal representing the speech of a single speaker as opposed to that of multiple speakers. A digital signal processing (DSP) algorithm can then detect the echo state, for example, whether distortion results primarily from an echo. Based on these characteristics, the system can, alternatively and automatically apply either a default mode of AEC to the audio signal, or apply a more aggressive mode of AEC.