Neural Echo Suppressor Time-Frequency Mask Residual Echo
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Acoustic echo cancellation in speech-enabled devices is inadequate, leading to residual echo that interferes with target speech recognition, particularly in environments where synthetic speech is played back.
Innovation Solution
A computer-implemented method using a neural echo suppressor (NES) that processes frequency-domain representations of output audio signals from a linear acoustic echo canceller (LAEC) to determine a time-frequency mask, which is then used to attenuate residual echo in the enhanced audio signal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If linear acoustic echo canceller (LAEC) is used to cancel acoustic echo, then acoustic echo is reduced, but residual echo remains that interferes with target speech recognition
Solution Approach 1:
The patent introduces a neural echo suppressor (NES) as an intermediary component between the LAEC and the speech recognition system. The NES processes the output of the LAEC and generates a time-frequency mask that selectively suppresses residual echo while preserving target speech, thereby improving speech recognition accuracy without completely replacing the LAEC
Solution Approach 2:
The patent transforms the audio signal from time-domain to frequency-domain representation and applies a time-frequency mask in the frequency domain. This parameter transformation allows for more precise control over echo suppression by operating on frequency components rather than raw audio waves, enabling better separation of echo and target speech
2Object-affected harmful factors
If neural echo suppressor (NES) is added to process output from LAEC, then residual echo is further reduced, but system complexity increases
Solution Approach 1:
The patent divides the echo cancellation task into two separate stages: first the LAEC handles the primary acoustic echo cancellation, then the NES handles the residual echo suppression. This segmentation allows each component to specialize in a specific aspect of echo reduction, improving overall performance while maintaining modular system architecture that can be implemented incrementally
3Measurement precision
If time-frequency mask is applied to attenuate residual echo, then speech recognition accuracy is improved, but processing time increases
Solution Approach 1:
The patent pre-computes the time-frequency mask based on the frequency-domain representation of the audio signal and the reference audio. By preparing the mask in advance before speech recognition occurs, the system minimizes real-time processing delays while ensuring accurate echo suppression is applied to the target speech signal
Data Source
AI summary
A method includes receiving a frequency-domain representation of an output audio signal output from a linear acoustic echo canceller (LAEC). The output audio signal includes target speech captured by an audio capture device of a user device and residual echo of reference audio output by an audio output device of the user device. The method also includes receiving a frequency-domain representation of the reference audio and determining, using a neural echo suppressor (NES), based on the frequency-domain representation of the output audio signal and the frequency-domain representation of the reference audio, a time-frequency mask. The method also includes processing, using the time-frequency mask, the frequency-domain representation of the output audio signal to attenuate the residual echo in an enhanced audio signal.


