Contrastive Training for Complex Neural Acoustic Echo Cancellation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing acoustic echo cancellation methods struggle to effectively suppress echoes in audio communication, especially in conference calls, leading to residual echoes when near-end and far-end speech are present simultaneously, and often introduce side effects.

Innovation Solution

A method and electronic device for training a complex neural model using contrastive learning (CL) to generate anchor, positive, and negative audio pairs, extracting features, and calculating loss functions to tune parameters, enhancing the model's ability to distinguish near-end from far-end speech.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional acoustic echo cancellation methods are used, then the system is simple and easy to implement, but residual echoes remain when near-end and far-end speech are present simultaneously

Engineering Contradiction:
ImproveAEC performanceVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by pre-training the neural model using contrastive learning with anchor, positive, and negative audio pairs before actual echo cancellation operation. This pre-training phase prepares the model to better distinguish near-end from far-end speech, improving AEC performance when both speech types are present simultaneously.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters by introducing a complex neural model with multiple trainable parameters through contrastive learning. The loss function calculates differences between anchor, positive, and negative features, adjusting model parameters to maximize separation between near-end and far-end speech representations, thereby reducing residual echoes.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If a complex neural model with contrastive learning is used, then AEC performance is greatly improved, but additional training costs and computational resources are required

Engineering Contradiction:
ImproveAEC performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs the computationally intensive contrastive learning training in advance as a preliminary action, separate from real-time echo cancellation operation. This allows the model to learn robust feature representations offline, so that during actual use, the pre-trained model can quickly process audio streams without incurring real-time training delays.

Inventive Principle:
Principle #10Preliminary action

3Ease of manufacture

If existing AEC methods are used, then the system is simple to implement, but side effects are introduced in audio communication

Engineering Contradiction:
Improveimplementation easeVSAvoidside effects
Core Design Contradiction:
Ease of manufactureVSObject-generated harmful factors

Solution Approach 1:

The patent replaces traditional mechanical or algorithmic echo cancellation methods with a neural network-based approach. The complex neural model learns to distinguish near-end from far-end speech through contrastive learning, substituting conventional signal processing techniques with a data-driven model that adapts to various acoustic environments without introducing the side effects of traditional methods.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12432493B2Method and electronic device for training complex neural model of acoustic echo cancellation
Publication Date: 2025.09.30 MEDIATEK SINGAPORE PTE LTD
  • US12432493B2 patent drawing
  • US12432493B2 patent drawing
  • US12432493B2 patent drawing

AI summary

A method and an electronic device for training a complex neural model of acoustic echo cancellation (AEC). The method includes: generating an anchor audio pair, a positive audio pair and a negative audio pair according to multiple near-end signals and multiple acoustic echo signals; utilizing the complex neural model to extract an anchor audio feature, a positive audio feature and a negative audio feature from the anchor audio pair, the positive audio pair and the negative pair, respectively; calculating a loss function according to the anchor audio feature, the positive audio feature and the negative audio feature; and tuning at least one parameter of the complex neural model according to the loss function.