GAN Audio Decoder with Dynamic Truncation for Speech Enhancement

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing deep learning approaches for speech enhancement, particularly in removing coding artifacts and noise, face challenges in achieving a balance between quality and variety, and are limited by the correlation of coding artifacts with desired sounds.

Innovation Solution

A Generative Adversarial Network (GAN) decoder is pre-configured with an encoder and decoder stage, using a bottleneck layer to concatenate audio features with a random noise vector, and applies truncation modes based on audio content and bitstream parameters to improve the processing of audio data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If deep learning approaches are used for speech enhancement, then quality of audio processing is improved, but flexibility and variety are limited

Engineering Contradiction:
Improveaudio processing qualityVSAvoidflexibility and variety
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic truncation modes that can be adjusted during operation. The system switches between different truncation levels (strong, intermediate, weak, or none) based on audio content analysis, allowing the model to adapt its behavior dynamically rather than being fixed, thus resolving the contradiction between processing quality and flexibility

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If coding artifacts are removed, then audio quality is improved, but the complexity of the processing increases

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent changes the truncation parameter of the GAN model based on audio characteristics and bitrate information. By adjusting this single key parameter dynamically, the system handles different coding artifact scenarios without requiring complex reconfiguration of the entire model architecture, thus improving audio quality while managing processing complexity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12518768B2Method and apparatus for processing of audio data using a pre-configured generator
Publication Date: 2026.01.06 DOLBY INTERNATIONAL AB
  • US12518768B2 patent drawing
  • US12518768B2 patent drawing
  • US12518768B2 patent drawing

AI summary

Described herein is a method for setting up a decoder for generating processed audio data from an audio bitstream, the decoder comprising a Generator of a Generative Adversarial Network, GAN, for processing the audio data. The method involves pre-configuring the Generator for processing of audio data with a set of parameters for the Generator, the parameters being determined by training the Generator using the full concatenated distribution. The method involves pre-configuring the decoder to determine a truncation mode for modifying the concatenated distribution and to apply the determined truncation mode to the concatenated distribution. Also described are a method of generating processed audio data from an audio bitstream using the GAN for processing audio data and a respective apparatus, as well as respective systems and computer program products.