Neural Network Speech Enhancement via Normalizing Flow
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in enhancing speech data quality, particularly in noisy environments, as they struggle to effectively remove noise and reverberation, leading to degraded performance in Automatic Speech Recognition (ASR) systems.
Innovation Solution
The method involves training a neural network using deep normalizing flow techniques to learn a maximum-likely encryption of high-quality speech data given simulated noisy speech data, which includes minimizing errors through decoding errors and spectral distance, and using this trained network to generate de-noised speech data, thereby enhancing audio quality and improving ASR performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional speech recognition systems are used in noisy environments, then the system can process speech data, but the speech data quality is degraded due to noise and reverberation
Solution Approach 1:
The patent applies preliminary action by training the neural network in advance with simulated noisy speech data that includes various noise types and reverberation characteristics. This pre-training enables the system to learn noise removal patterns before actual deployment, allowing it to effectively enhance speech quality when processing real noisy speech without requiring real-time complex noise modeling
Solution Approach 2:
The patent implements feedback by using ASR decoding error and spectral distance as loss functions during neural network training. The ASR system processes the enhanced speech and the decoding error feedback is used to adjust the training, creating a closed-loop system where the enhancement quality directly influences the training optimization, thereby improving speech recognition performance in noisy environments
2Measurement precision
If deep normalizing flow training is used to minimize errors, then speech data quality is improved, but computational complexity increases
Solution Approach 1:
The patent applies parameter changes by transforming the speech data into a latent space using normalizing flows, where the complex noise removal task is simplified by changing the parameter representation. The neural network learns to manipulate parameters in this transformed space rather than directly processing raw speech signals, making the optimization process more efficient while maintaining high speech quality
Solution Approach 2:
The patent introduces an intermediary latent space through normalizing flows that mediates between the noisy speech input and the enhanced speech output. This intermediary representation simplifies the optimization landscape by separating the noise removal task from the speech content preservation, allowing the system to achieve high speech quality without directly navigating the full complexity of the original signal space
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach results in high-performance, low-latency audio enhancement capable of removing various types of noise, including background speakers and reverberation, significantly improving the quality of speech data and enhancing ASR system performance.
Implementation Method 1
the training is deep normalizing flow training. In an embodiment, during the deep normalizing flow training the errors in the neural network are minimized as described herein
Implementation Method 2
creating the simulated noisy speech data by adding reverberation to the high quality speech data using convolution
Data Source
AI summary
Embodiments improve speech data quality through training a neural network for de-noising audio enhancement. One such embodiment creates simulated noisy speech data from high quality speech data. In turn, training, e.g., deep normalizing flow training, is performed on a neural network using the high quality speech data and the simulated noisy speech data to train the neural network to create de-noised speech data given noisy speech data. Performing the training includes minimizing errors in the neural network according to at least one of (i) a decoding error of an Automatic Speech Recognition (ASR) system processing current de-noised speech data results generated by the neural network during the training and (ii) spectral distance between the high quality speech data and the current de-noised speech data results generated by the neural network during the training.


