Neural Network Speech Enhancement via Normalizing Flow

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in enhancing speech data quality, particularly in noisy environments, as they struggle to effectively remove noise and reverberation, leading to degraded performance in Automatic Speech Recognition (ASR) systems.

Innovation Solution

The method involves training a neural network using deep normalizing flow techniques to learn a maximum-likely encryption of high-quality speech data given simulated noisy speech data, which includes minimizing errors through decoding errors and spectral distance, and using this trained network to generate de-noised speech data, thereby enhancing audio quality and improving ASR performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional speech recognition systems are used in noisy environments, then the system can process speech data, but the speech data quality is degraded due to noise and reverberation

Engineering Contradiction:
Improvespeech recognition performanceVSAvoidnoise and reverberation
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary action by training the neural network in advance with simulated noisy speech data that includes various noise types and reverberation characteristics. This pre-training enables the system to learn noise removal patterns before actual deployment, allowing it to effectively enhance speech quality when processing real noisy speech without requiring real-time complex noise modeling

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements feedback by using ASR decoding error and spectral distance as loss functions during neural network training. The ASR system processes the enhanced speech and the decoding error feedback is used to adjust the training, creating a closed-loop system where the enhancement quality directly influences the training optimization, thereby improving speech recognition performance in noisy environments

Inventive Principle:
Principle #23Feedback

2Measurement precision

If deep normalizing flow training is used to minimize errors, then speech data quality is improved, but computational complexity increases

Engineering Contradiction:
Improvespeech data qualityVSAvoidneural network training complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies parameter changes by transforming the speech data into a latent space using normalizing flows, where the complex noise removal task is simplified by changing the parameter representation. The neural network learns to manipulate parameters in this transformed space rather than directly processing raw speech signals, making the optimization process more efficient while maintaining high speech quality

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary latent space through normalizing flows that mediates between the noisy speech input and the enhanced speech output. This intermediary representation simplifies the optimization landscape by separating the noise removal task from the speech content preservation, allowing the system to achieve high speech quality without directly navigating the full complexity of the original signal space

Inventive Principle:
Principle #24Intermediary (Mediator)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach results in high-performance, low-latency audio enhancement capable of removing various types of noise, including background speakers and reverberation, significantly improving the quality of speech data and enhancing ASR system performance.

Implementation Method 1

the training is deep normalizing flow training. In an embodiment, during the deep normalizing flow training the errors in the neural network are minimized as described herein

Methodology Applied
Scientific EffectDeep normalizing flow training:

Implementation Method 2

creating the simulated noisy speech data by adding reverberation to the high quality speech data using convolution

Methodology Applied
Scientific EffectConvolution:

Data Source

PatentUS11657828B2Method and system for speech enhancement
Publication Date: 2023.05.23 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11657828B2 patent drawing
  • US11657828B2 patent drawing
  • US11657828B2 patent drawing

AI summary

Embodiments improve speech data quality through training a neural network for de-noising audio enhancement. One such embodiment creates simulated noisy speech data from high quality speech data. In turn, training, e.g., deep normalizing flow training, is performed on a neural network using the high quality speech data and the simulated noisy speech data to train the neural network to create de-noised speech data given noisy speech data. Performing the training includes minimizing errors in the neural network according to at least one of (i) a decoding error of an Automatic Speech Recognition (ASR) system processing current de-noised speech data results generated by the neural network during the training and (ii) spectral distance between the high quality speech data and the current de-noised speech data results generated by the neural network during the training.