Neural Echo Suppressor Time-Frequency Mask Residual Echo

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Acoustic echo cancellation in speech-enabled devices is inadequate, leading to residual echo that interferes with target speech recognition, particularly in environments where synthetic speech is played back.

Innovation Solution

A computer-implemented method using a neural echo suppressor (NES) that processes frequency-domain representations of output audio signals from a linear acoustic echo canceller (LAEC) to determine a time-frequency mask, which is then used to attenuate residual echo in the enhanced audio signal.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If linear acoustic echo canceller (LAEC) is used to cancel acoustic echo, then acoustic echo is reduced, but residual echo remains that interferes with target speech recognition

Engineering Contradiction:
Improveacoustic echoVSAvoidspeech recognition accuracy
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent introduces a neural echo suppressor (NES) as an intermediary component between the LAEC and the speech recognition system. The NES processes the output of the LAEC and generates a time-frequency mask that selectively suppresses residual echo while preserving target speech, thereby improving speech recognition accuracy without completely replacing the LAEC

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms the audio signal from time-domain to frequency-domain representation and applies a time-frequency mask in the frequency domain. This parameter transformation allows for more precise control over echo suppression by operating on frequency components rather than raw audio waves, enabling better separation of echo and target speech

Inventive Principle:
Principle #35Parameter changes

2Object-affected harmful factors

If neural echo suppressor (NES) is added to process output from LAEC, then residual echo is further reduced, but system complexity increases

Engineering Contradiction:
Improveresidual echoVSAvoidsignal processing system
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent divides the echo cancellation task into two separate stages: first the LAEC handles the primary acoustic echo cancellation, then the NES handles the residual echo suppression. This segmentation allows each component to specialize in a specific aspect of echo reduction, improving overall performance while maintaining modular system architecture that can be implemented incrementally

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If time-frequency mask is applied to attenuate residual echo, then speech recognition accuracy is improved, but processing time increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsignal processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent pre-computes the time-frequency mask based on the frequency-domain representation of the audio signal and the reference audio. By preparing the mask in advance before speech recognition occurs, the system minimizes real-time processing delays while ensuring accurate echo suppression is applied to the target speech signal

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250203282A1Acoustic Echo Cancellation For Digital Assistants Using Neural Echo Suppression and Multi-Microphone Noise Reduction
Publication Date: 2025.06.19 GOOGLE LLC
  • US20250203282A1 patent drawing
  • US20250203282A1 patent drawing
  • US20250203282A1 patent drawing

AI summary

A method includes receiving a frequency-domain representation of an output audio signal output from a linear acoustic echo canceller (LAEC). The output audio signal includes target speech captured by an audio capture device of a user device and residual echo of reference audio output by an audio output device of the user device. The method also includes receiving a frequency-domain representation of the reference audio and determining, using a neural echo suppressor (NES), based on the frequency-domain representation of the output audio signal and the frequency-domain representation of the reference audio, a time-frequency mask. The method also includes processing, using the time-frequency mask, the frequency-domain representation of the output audio signal to attenuate the residual echo in an enhanced audio signal.