Single-Microphone Echo and Noise Separation for Robust Speech Enhancement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing acoustic echo cancellation (AEC) techniques in hands-free communication devices struggle to effectively suppress echoes and noise due to nonlinearities introduced by amplifiers and mechanical components, and machine learning models perform poorly in untrained environments.
Innovation Solution
A speech enhancement system using a delay estimator and an acoustic echo and noise (AEN) decoupling filter, which includes a neural network to generate masks for separating speech, echo, and noise components, allowing for improved suppression of acoustic echoes and noise even in untrained environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If linear transfer functions (NLMS algorithm) are used for acoustic echo cancellation, then the system complexity is low and ease of manufacture is good, but the convergence rate deteriorates under double-talk conditions and changes to echo path
Solution Approach 1:
The patent segments the acoustic echo cancellation problem into multiple independent adaptive filters, each handling a specific frequency band or echo path component. This segmentation allows each filter to converge faster and more reliably while maintaining overall system manageability and ease of implementation.
Solution Approach 2:
The patent employs dynamic adaptive filtering where filter parameters are continuously adjusted based on real-time acoustic conditions, double-talk detection, and echo path changes. This dynamic adaptation improves convergence rate and reliability while maintaining computational efficiency through selective updating of filter coefficients.
2Device complexity
If linear transfer functions are used, then device complexity is low, but the system cannot account for nonlinearities introduced by amplifiers and mechanical components
Solution Approach 1:
The patent introduces nonlinear distortion models and pre-processing stages as intermediary components between the linear adaptive filters and the acoustic echo path. These intermediaries capture nonlinearities from amplifiers and mechanical components, allowing the main linear filters to focus on linear echo cancellation while maintaining high modeling accuracy.
Solution Approach 2:
The patent replaces purely mechanical/acoustic linear modeling with a hybrid approach that substitutes mathematical nonlinear transformation functions to model the effects of amplifiers and mechanical components. This substitution enables accurate representation of nonlinear behaviors without requiring complex physical models of each component.
3Measurement precision
If machine learning models are used for echo and noise suppression, then speech quality can be improved, but performance deteriorates in untrained environments
Solution Approach 1:
The patent performs preliminary training of machine learning models on diverse acoustic environments, noise types, and speech characteristics before deployment. This preliminary action ensures the models have learned robust features and patterns that generalize well to untrained environments, maintaining high speech quality across varying conditions.
Solution Approach 2:
The patent implements adaptive parameter adjustment where the machine learning model's operating parameters (such as confidence thresholds, mixing ratios, and suppression strengths) are dynamically changed based on environmental sensing and performance monitoring. This allows the system to adapt to untrained environments by adjusting parameters rather than requiring full retraining.
Data Source
AI summary
This disclosure provides methods, devices, and systems for audio signal processing. The present implementations more specifically relate to speech enhancement techniques for separating microphone signals into speech, echo, and noise signals. In some aspects, a speech enhancement system may include a delay estimator and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives a microphone signal via a microphone and a far-end audio signal for output via a speaker and estimates a reference audio signal based on a delay between the microphone signal and the far-end audio signal. In some aspects, the AEN decoupling filter may determine a speech mask, an echo mask, and a noise mask based on the microphone signal and the reference audio signal and may suppress an echo component and a noise component of the microphone signal based on the determined set of masks.


