Signal-Level Normalization for Speech Enhancement and Echo Suppression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing acoustic echo cancellation (AEC) techniques in hands-free communication devices struggle to account for nonlinearities introduced by amplifiers and mechanical components, leading to inaccurate speech enhancement due to varying signal levels and phase deviations in audio signals.
Innovation Solution
A speech enhancement system that includes a delay estimator, input normalizer, and acoustic echo and noise (AEN) decoupling filter, which normalizes loudness and utilizes a neural network to determine masks for suppressing echo and noise components based on normalized audio signals, accounting for nonlinearities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If linear transfer functions (NLMS algorithm) are used for acoustic echo cancellation, then the system complexity is low and ease of manufacture is good, but the speech enhancement quality deteriorates due to inability to account for nonlinearities and varying signal levels
Solution Approach 1:
The patent applies parameter changes by normalizing the signal levels of both the reference audio signal and near-end audio signal before processing. This normalization step transforms the varying signal levels into a consistent range, allowing the subsequent mask determination to operate effectively across different loudness conditions. The signal level normalization directly addresses the limitation of linear transfer functions that cannot adapt to varying signal levels, thereby improving speech enhancement quality without requiring a complete overhaul of the system architecture.
Solution Approach 2:
The patent segments the speech enhancement process into distinct functional modules: signal level normalization, mask determination, and echo/noise suppression. By dividing the processing into these separate stages, the system can apply specialized operations at each step - normalization handles signal level variations, mask determination captures nonlinear relationships, and suppression applies the learned masks. This segmentation allows complex nonlinear processing to be achieved through a series of manageable steps, balancing implementation complexity with enhancement quality.
2Device complexity
If linear transfer functions are used to model acoustic coupling, then the device complexity is low, but the reliability of echo cancellation deteriorates due to double-talk conditions and echo path changes
Solution Approach 1:
The patent introduces dynamic adaptation through mask determination that operates on normalized signals at each processing stage. Unlike static linear transfer functions, the mask determination process dynamically adjusts to current signal conditions, including double-talk scenarios and echo path variations. The system continuously computes masks based on current normalized reference and near-end signals, allowing it to adapt to changing acoustic conditions without requiring complex adaptive filter updates, thereby improving reliability while maintaining reasonable complexity.
3Adaptability or versatility
If signal level normalization is applied, then the adaptability to varying signal levels is improved, but the processing time and computational load increase
Solution Approach 1:
The patent applies preliminary action by performing signal level normalization before the main mask determination and suppression operations. By pre-normalizing the input signals, the system prepares them in an optimal state for subsequent processing, ensuring that all downstream operations work with consistently scaled inputs. This preliminary normalization step, while adding some computational overhead, simplifies subsequent processing by eliminating the need for repeated signal level adjustments, ultimately improving overall processing efficiency and achieving signal level independence.
Data Source
AI summary
This disclosure provides methods, devices, and systems for audio signal processing. The present implementations more specifically relate to speech enhancement techniques that are agnostic to varying signal levels in near-end audio signals. In some aspects, a speech enhancement system may include a delay estimator, an input normalizer, and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives a near-end audio signal via a microphone and a far-end audio signal for output via a speaker and estimates a reference audio signal based on a delay between the near-end audio signal and the far-end audio signal. The input normalizer normalizes a loudness of the near-end audio signal and the reference audio signal. The AEN decoupling filter determines a set of masks based on the normalized audio signals and suppresses an echo component and a noise component of the near-end audio signal based on the set of masks.


