Dynamic Range-Reduced Multi-Channel GAN for Codec Artifact Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing low-bitrate audio coding technologies introduce coding artifacts and noise in multi-channel audio signals, which current deep learning approaches have not effectively addressed, particularly in the context of spatial enhancement and codec-agnostic restoration.

Innovation Solution

A method using a multi-channel Generator trained in a Generative Adversarial Network (GAN) setting to jointly enhance multi-channel audio signals in a dynamic range reduced domain, incorporating companding techniques and metadata for selective enhancement based on companding modes, while utilizing a multi-channel Discriminator for training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If low-bitrate audio coding is used to reduce bandwidth and storage requirements, then transmission and storage efficiency is improved, but coding artifacts and quantization noise are introduced degrading audio quality

Engineering Contradiction:
Improvebandwidth and storage requirementsVSAvoidcoding artifacts and quantization noise
Core Design Contradiction:
Loss of energyVSObject-affected harmful factors

Solution Approach 1:

The patent applies a Generative Adversarial Network (GAN) where the Generator is trained to convert coded audio signals containing quantization noise into enhanced audio signals. The Generator learns to map from the noisy coded domain to the clean audio domain, effectively converting the harmful quantization noise into beneficial audio restoration. The adversarial training with the Discriminator further refines this conversion process to produce high-fidelity enhanced audio.

Inventive Principle:
Principle #22Blessing in disguise (Convert harm into benefit)

Solution Approach 2:

The patent introduces a multi-channel Generator as an intermediary component between the coded audio signal and the enhanced audio output. This Generator acts as a mediator that processes the coded signal through learned transformations, applying companding techniques and spatial enhancement to bridge the gap between low-bitrate coded audio and high-quality reconstructed audio, thereby reducing the perceptual impact of coding artifacts.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If multi-channel approaches with spatial information are used to enhance audio quality, then audio quality is improved, but processing complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing complexity
Core Design Contradiction:
Object-affected harmful factorsVSDevice complexity

Solution Approach 1:

The patent merges multiple enhancement functions into a single multi-channel Generator network. Instead of applying separate single-channel enhancement and spatial processing, the Generator jointly processes all audio channels simultaneously, learning both spectral and spatial relationships in one unified model. This consolidation reduces the overall system complexity while maintaining multi-channel audio quality enhancement.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The multi-channel Generator serves multiple functions simultaneously: it performs denoising, spatial enhancement, and companding operations on all audio channels in a single processing pass. This multi-functional approach eliminates the need for multiple separate processing stages, thereby reducing computational complexity while achieving comprehensive audio enhancement.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Object-affected harmful factors

If deep learning approaches are applied to restore audio from coding noise, then audio quality is improved, but the approaches are mostly limited to speech denoising rather than codec-agnostic restoration

Engineering Contradiction:
Improveaudio qualityVSAvoidcodec-agnostic restoration capability
Core Design Contradiction:
Object-affected harmful factorsVSAdaptability or versatility

Solution Approach 1:

The patent designs a codec-agnostic multi-channel Generator that can process and enhance audio from various coding formats (AC-3, AAC, AC-4) without requiring format-specific processing. The Generator learns universal patterns of coding artifacts and restoration transformations that apply across different codecs, making the system versatile and adaptable to multiple audio coding standards while maintaining high audio quality restoration.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies companding (compression-expansion) parameter transformations within the Generator to adapt the enhancement process to different dynamic range requirements. By dynamically adjusting compression and expansion parameters based on the input signal characteristics, the system achieves effective restoration across various coding conditions and formats, enhancing both speech and music content universally.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4143825B1Method, apparatus and system for enhancing multi-channel audio in a dynamic range reduced domain
Publication Date: 2025.08.13 DOLBY INTERNATIONAL AB
  • EP4143825B1 patent drawingFigure 1
  • EP4143825B1 patent drawingFigure 2
  • EP4143825B1 patent drawingFigure 3

AI summary

Described herein is a method of generating, in a dynamic range reduced domain, an enhanced multi-channel audio signal from an audio bitstream including a multi-channel audio signal, wherein the multi-channel audio signal comprises two or more channels, and wherein the method includes jointly enhancing the two or more channels of the dynamic range reduced raw multi-channel audio signal using a multi-channel Generator of a Generative Adversarial Network setting. Described herein are further a method for training a multi-channel Generator in a dynamic range reduced domain in a Generative Adversarial Network setting, an apparatus for generating, in a dynamic range reduced domain, an enhanced multi-channel audio signal from an audio bitstream including a multi-channel audio signal, respective systems and a computer program product.