Neural Audio Enhancement via Time-Domain Speaker Embedding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio enhancement techniques for real-time target speaker audio enhancement face challenges due to high computational complexity, particularly when operating in the short-time Fourier transform (STFT) domain, making them unsuitable for low-complexity, real-time applications.
Innovation Solution
The implementation of a machine learning model-based audio enhancement technique, such as PercepNet, that operates on a perceptually motivated, low-dimensional feature space, using a speaker embedder network to condition the separation network and isolate the target speaker, thereby reducing computational complexity while maintaining high-quality speech enhancement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If deep learning-based audio enhancement methods operate directly in the STFT domain, then speech enhancement quality is improved, but computational complexity increases
Solution Approach 1:
The patent replaces traditional mechanical signal processing methods (spectral subtraction, spectral estimation in STFT domain) with a neural network-based system that operates in the time domain. This substitution allows achieving high speech enhancement quality without the computational burden of STFT-domain processing, as the neural network directly processes time-domain signals and outputs enhanced speech without requiring complex frequency-domain transformations.
2Device complexity
If traditional spectral subtraction and spectral estimation methods are used, then computational complexity is reduced, but speech enhancement quality deteriorates
Solution Approach 1:
The patent changes the fundamental parameters of the audio enhancement approach by transitioning from frequency-domain operations (STFT) to time-domain processing with neural networks. This parameter change enables the system to achieve superior speech enhancement quality comparable to or exceeding traditional methods, while maintaining lower computational complexity through efficient time-domain convolution operations and direct speech extraction without iterative spectral processing.
3Measurement precision
If high-quality speech enhancement is achieved through deep learning methods, then speech quality is improved, but real-time processing capability is reduced
Solution Approach 1:
The patent segments the audio enhancement task into distinct functional components handled by specialized neural network modules: a speech detector network that identifies speech regions, and an enhancement network that processes only those identified speech portions. This segmentation allows the system to focus computational resources selectively on speech segments rather than processing entire audio streams, enabling high-quality enhancement to be achieved in real-time with reduced overall computational burden.
Solution Approach 2:
The patent applies partial action by using the speech detector to identify and process only the relevant speech portions of the audio signal, rather than applying computationally intensive enhancement algorithms to the entire audio stream. This selective processing approach maintains high speech enhancement quality for actual speech content while significantly reducing computational complexity during non-speech periods, thereby enabling real-time processing capability.
Data Source
AI summary
Real-time audio enhancement for a target speaker may be performed. An embedding of a sample of speaker audio is created using a trained neural network that performs voice identification. The embedding is then concatenated with the input features of a trained machine learning model for audio enhancement. The audio enhancement model can recognize and enhance a target speaker's speech in a real-time implementation, as the embedding is in the same feature space of the audio enhancement model.


