Deep Generative Speech Enhancement for Low-Bitrate Coded Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech coding technologies suffer from quality issues due to decreasing bitrates, leading to distorted and contaminated speech signals that are difficult to restore using conventional methods.

Innovation Solution

A system utilizing self-supervised deep learning models to extract robust feature vectors from coded audio data, followed by generative deep learning models to generate clean speech signals, addressing distortion and coding artifacts.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If speech coding bitrate is decreased to enable low-cost mobile and internet communication, then communication cost is reduced, but speech quality deteriorates with various quality issues

Engineering Contradiction:
Improvecommunication costVSAvoidspeech quality
Core Design Contradiction:
Loss of energyVSManufacturing precision

Solution Approach 1:

A deep learning-based speech enhancement system is introduced as an intermediary component between the speech codec and the audio output. The system takes coded speech as input, extracts robust features using a self-supervised model, and generates enhanced speech output, thereby mediating the quality degradation caused by low-bitrate coding.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces conventional signal processing methods with deep learning-based approaches. Specifically, it substitutes traditional speech enhancement algorithms with neural network models including a self-supervised feature extraction model and a generative speech synthesis model, achieving superior quality restoration.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If conventional methods are used to restore speech from coded audio data, then processing simplicity is maintained, but restoration quality is insufficient due to distortion and coding artifacts

Engineering Contradiction:
Improveprocessing simplicityVSAvoidrestoration quality
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The speech restoration process is segmented into three distinct stages: (1) feature extraction using a self-supervised deep learning model, (2) speech generation using a generative deep learning model, and (3) output enhancement. This segmentation allows each component to be optimized independently while achieving high overall quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces intermediate feature vectors as a mediator between the coded speech input and the final restored speech output. The self-supervised model extracts robust feature vectors that capture essential speech characteristics, which then serve as input to the generative model for high-quality speech synthesis.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If deep learning models are used to generate enhanced speech data from coded audio, then speech quality is improved, but processing complexity increases

Engineering Contradiction:
Improvespeech qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The self-supervised deep learning model performs automatic feature extraction without requiring manual feature engineering or extensive preprocessing. The model learns relevant speech features directly from the coded audio data, reducing the need for complex preprocessing pipelines and manual intervention.

Inventive Principle:
Principle #25Self-service

4Manufacturing precision

If robust features are extracted from coded audio data to achieve high quality improved speech, then speech quality is enhanced, but processing time increases

Engineering Contradiction:
Improvespeech qualityVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The self-supervised feature extraction model performs preliminary processing by extracting robust feature vectors from the coded speech before the main speech generation stage. This preliminary action prepares the data in an optimized format that accelerates the subsequent generative modeling process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP4715811A1Coded speech enhancement based on deep generative model
Publication Date: 2026.03.25 DOLBY LABORATORIES LICENSING CORP
  • EP4715811A1 patent drawingFigure 1~2
  • EP4715811A1 patent drawingFigure 3~4
  • EP4715811A1 patent drawingFigure 5

AI summary

A system for generating enhanced speech data using robust audio features is disclosed. In some embodiments, a system is programmed to use a self-supervised deep learning model to generate a set of feature vectors from given audio data that contains contaminated speech and is coded. The system is further programmed to use a generative deep learning model to create improved audio data corresponding to clean speech from the set of feature vectors.