Deep Generative Speech Enhancement for Low-Bitrate Coded Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech coding technologies suffer from quality issues due to decreasing bitrates, leading to distorted and contaminated speech signals that are difficult to restore using conventional methods.
Innovation Solution
A system utilizing self-supervised deep learning models to extract robust feature vectors from coded audio data, followed by generative deep learning models to generate clean speech signals, addressing distortion and coding artifacts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If speech coding bitrate is decreased to enable low-cost mobile and internet communication, then communication cost is reduced, but speech quality deteriorates with various quality issues
Solution Approach 1:
A deep learning-based speech enhancement system is introduced as an intermediary component between the speech codec and the audio output. The system takes coded speech as input, extracts robust features using a self-supervised model, and generates enhanced speech output, thereby mediating the quality degradation caused by low-bitrate coding.
Solution Approach 2:
The patent replaces conventional signal processing methods with deep learning-based approaches. Specifically, it substitutes traditional speech enhancement algorithms with neural network models including a self-supervised feature extraction model and a generative speech synthesis model, achieving superior quality restoration.
2Device complexity
If conventional methods are used to restore speech from coded audio data, then processing simplicity is maintained, but restoration quality is insufficient due to distortion and coding artifacts
Solution Approach 1:
The speech restoration process is segmented into three distinct stages: (1) feature extraction using a self-supervised deep learning model, (2) speech generation using a generative deep learning model, and (3) output enhancement. This segmentation allows each component to be optimized independently while achieving high overall quality.
Solution Approach 2:
The patent introduces intermediate feature vectors as a mediator between the coded speech input and the final restored speech output. The self-supervised model extracts robust feature vectors that capture essential speech characteristics, which then serve as input to the generative model for high-quality speech synthesis.
3Manufacturing precision
If deep learning models are used to generate enhanced speech data from coded audio, then speech quality is improved, but processing complexity increases
Solution Approach 1:
The self-supervised deep learning model performs automatic feature extraction without requiring manual feature engineering or extensive preprocessing. The model learns relevant speech features directly from the coded audio data, reducing the need for complex preprocessing pipelines and manual intervention.
4Manufacturing precision
If robust features are extracted from coded audio data to achieve high quality improved speech, then speech quality is enhanced, but processing time increases
Solution Approach 1:
The self-supervised feature extraction model performs preliminary processing by extracting robust feature vectors from the coded speech before the main speech generation stage. This preliminary action prepares the data in an optimized format that accelerates the subsequent generative modeling process.
Data Source
Figure 1~2
Figure 3~4
Figure 5
AI summary
A system for generating enhanced speech data using robust audio features is disclosed. In some embodiments, a system is programmed to use a self-supervised deep learning model to generate a set of feature vectors from given audio data that contains contaminated speech and is coded. The system is further programmed to use a generative deep learning model to create improved audio data corresponding to clean speech from the set of feature vectors.