Signal-Level Normalization for Speech Enhancement and Echo Suppression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing acoustic echo cancellation (AEC) techniques in hands-free communication devices struggle to account for nonlinearities introduced by amplifiers and mechanical components, leading to inaccurate speech enhancement due to varying signal levels and phase deviations in audio signals.

Innovation Solution

A speech enhancement system that includes a delay estimator, input normalizer, and acoustic echo and noise (AEN) decoupling filter, which normalizes loudness and utilizes a neural network to determine masks for suppressing echo and noise components based on normalized audio signals, accounting for nonlinearities.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If linear transfer functions (NLMS algorithm) are used for acoustic echo cancellation, then the system complexity is low and ease of manufacture is good, but the speech enhancement quality deteriorates due to inability to account for nonlinearities and varying signal levels

Engineering Contradiction:
Improveease of implementationVSAvoidspeech enhancement quality
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by normalizing the signal levels of both the reference audio signal and near-end audio signal before processing. This normalization step transforms the varying signal levels into a consistent range, allowing the subsequent mask determination to operate effectively across different loudness conditions. The signal level normalization directly addresses the limitation of linear transfer functions that cannot adapt to varying signal levels, thereby improving speech enhancement quality without requiring a complete overhaul of the system architecture.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent segments the speech enhancement process into distinct functional modules: signal level normalization, mask determination, and echo/noise suppression. By dividing the processing into these separate stages, the system can apply specialized operations at each step - normalization handles signal level variations, mask determination captures nonlinear relationships, and suppression applies the learned masks. This segmentation allows complex nonlinear processing to be achieved through a series of manageable steps, balancing implementation complexity with enhancement quality.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If linear transfer functions are used to model acoustic coupling, then the device complexity is low, but the reliability of echo cancellation deteriorates due to double-talk conditions and echo path changes

Engineering Contradiction:
Improvealgorithm complexityVSAvoidecho cancellation reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent introduces dynamic adaptation through mask determination that operates on normalized signals at each processing stage. Unlike static linear transfer functions, the mask determination process dynamically adjusts to current signal conditions, including double-talk scenarios and echo path variations. The system continuously computes masks based on current normalized reference and near-end signals, allowing it to adapt to changing acoustic conditions without requiring complex adaptive filter updates, thereby improving reliability while maintaining reasonable complexity.

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If signal level normalization is applied, then the adaptability to varying signal levels is improved, but the processing time and computational load increase

Engineering Contradiction:
Improvesignal level independenceVSAvoidprocessing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent applies preliminary action by performing signal level normalization before the main mask determination and suppression operations. By pre-normalizing the input signals, the system prepares them in an optimal state for subsequent processing, ensuring that all downstream operations work with consistently scaled inputs. This preliminary normalization step, while adding some computational overhead, simplifies subsequent processing by eliminating the need for repeated signal level adjustments, ultimately improving overall processing efficiency and achieving signal level independence.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12412589B2Signal level-independent speech enhancement
Publication Date: 2025.09.09 SYNAPTICS INC
  • US12412589B2 patent drawing
  • US12412589B2 patent drawing
  • US12412589B2 patent drawing

AI summary

This disclosure provides methods, devices, and systems for audio signal processing. The present implementations more specifically relate to speech enhancement techniques that are agnostic to varying signal levels in near-end audio signals. In some aspects, a speech enhancement system may include a delay estimator, an input normalizer, and an acoustic echo and noise (AEN) decoupling filter. The delay estimator receives a near-end audio signal via a microphone and a far-end audio signal for output via a speaker and estimates a reference audio signal based on a delay between the near-end audio signal and the far-end audio signal. The input normalizer normalizes a loudness of the near-end audio signal and the reference audio signal. The AEN decoupling filter determines a set of masks based on the normalized audio signals and suppresses an echo component and a noise component of the near-end audio signal based on the set of masks.