Neural Audio Enhancement via Time-Domain Speaker Embedding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio enhancement techniques for real-time target speaker audio enhancement face challenges due to high computational complexity, particularly when operating in the short-time Fourier transform (STFT) domain, making them unsuitable for low-complexity, real-time applications.

Innovation Solution

The implementation of a machine learning model-based audio enhancement technique, such as PercepNet, that operates on a perceptually motivated, low-dimensional feature space, using a speaker embedder network to condition the separation network and isolate the target speaker, thereby reducing computational complexity while maintaining high-quality speech enhancement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep learning-based audio enhancement methods operate directly in the STFT domain, then speech enhancement quality is improved, but computational complexity increases

Engineering Contradiction:
Improvespeech enhancement qualityVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces traditional mechanical signal processing methods (spectral subtraction, spectral estimation in STFT domain) with a neural network-based system that operates in the time domain. This substitution allows achieving high speech enhancement quality without the computational burden of STFT-domain processing, as the neural network directly processes time-domain signals and outputs enhanced speech without requiring complex frequency-domain transformations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If traditional spectral subtraction and spectral estimation methods are used, then computational complexity is reduced, but speech enhancement quality deteriorates

Engineering Contradiction:
Improvecomputational complexityVSAvoidspeech enhancement quality
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent changes the fundamental parameters of the audio enhancement approach by transitioning from frequency-domain operations (STFT) to time-domain processing with neural networks. This parameter change enables the system to achieve superior speech enhancement quality comparable to or exceeding traditional methods, while maintaining lower computational complexity through efficient time-domain convolution operations and direct speech extraction without iterative spectral processing.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If high-quality speech enhancement is achieved through deep learning methods, then speech quality is improved, but real-time processing capability is reduced

Engineering Contradiction:
Improvespeech enhancement qualityVSAvoidreal-time processing capability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent segments the audio enhancement task into distinct functional components handled by specialized neural network modules: a speech detector network that identifies speech regions, and an enhancement network that processes only those identified speech portions. This segmentation allows the system to focus computational resources selectively on speech segments rather than processing entire audio streams, enabling high-quality enhancement to be achieved in real-time with reduced overall computational burden.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by using the speech detector to identify and process only the relevant speech portions of the audio signal, rather than applying computationally intensive enhancement algorithms to the entire audio stream. This selective processing approach maintains high speech enhancement quality for actual speech content while significantly reducing computational complexity during non-speech periods, thereby enabling real-time processing capability.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12272371B1Real-time target speaker audio enhancement
Publication Date: 2025.04.08 AMAZON TECH INC
  • US12272371B1 patent drawing
  • US12272371B1 patent drawing
  • US12272371B1 patent drawing

AI summary

Real-time audio enhancement for a target speaker may be performed. An embedding of a sample of speaker audio is created using a trained neural network that performs voice identification. The embedding is then concatenated with the input features of a trained machine learning model for audio enhancement. The audio enhancement model can recognize and enhance a target speaker's speech in a real-time implementation, as the embedding is in the same feature space of the audio enhancement model.