Dynamic Range-Reduced Multi-Channel GAN for Codec Artifact Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing low-bitrate audio coding technologies introduce coding artifacts and noise in multi-channel audio signals, which current deep learning approaches have not effectively addressed, particularly in the context of spatial enhancement and codec-agnostic restoration.
Innovation Solution
A method using a multi-channel Generator trained in a Generative Adversarial Network (GAN) setting to jointly enhance multi-channel audio signals in a dynamic range reduced domain, incorporating companding techniques and metadata for selective enhancement based on companding modes, while utilizing a multi-channel Discriminator for training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If low-bitrate audio coding is used to reduce bandwidth and storage requirements, then transmission and storage efficiency is improved, but coding artifacts and quantization noise are introduced degrading audio quality
Solution Approach 1:
The patent applies a Generative Adversarial Network (GAN) where the Generator is trained to convert coded audio signals containing quantization noise into enhanced audio signals. The Generator learns to map from the noisy coded domain to the clean audio domain, effectively converting the harmful quantization noise into beneficial audio restoration. The adversarial training with the Discriminator further refines this conversion process to produce high-fidelity enhanced audio.
Solution Approach 2:
The patent introduces a multi-channel Generator as an intermediary component between the coded audio signal and the enhanced audio output. This Generator acts as a mediator that processes the coded signal through learned transformations, applying companding techniques and spatial enhancement to bridge the gap between low-bitrate coded audio and high-quality reconstructed audio, thereby reducing the perceptual impact of coding artifacts.
2Object-affected harmful factors
If multi-channel approaches with spatial information are used to enhance audio quality, then audio quality is improved, but processing complexity increases
Solution Approach 1:
The patent merges multiple enhancement functions into a single multi-channel Generator network. Instead of applying separate single-channel enhancement and spatial processing, the Generator jointly processes all audio channels simultaneously, learning both spectral and spatial relationships in one unified model. This consolidation reduces the overall system complexity while maintaining multi-channel audio quality enhancement.
Solution Approach 2:
The multi-channel Generator serves multiple functions simultaneously: it performs denoising, spatial enhancement, and companding operations on all audio channels in a single processing pass. This multi-functional approach eliminates the need for multiple separate processing stages, thereby reducing computational complexity while achieving comprehensive audio enhancement.
3Object-affected harmful factors
If deep learning approaches are applied to restore audio from coding noise, then audio quality is improved, but the approaches are mostly limited to speech denoising rather than codec-agnostic restoration
Solution Approach 1:
The patent designs a codec-agnostic multi-channel Generator that can process and enhance audio from various coding formats (AC-3, AAC, AC-4) without requiring format-specific processing. The Generator learns universal patterns of coding artifacts and restoration transformations that apply across different codecs, making the system versatile and adaptable to multiple audio coding standards while maintaining high audio quality restoration.
Solution Approach 2:
The patent applies companding (compression-expansion) parameter transformations within the Generator to adapt the enhancement process to different dynamic range requirements. By dynamically adjusting compression and expansion parameters based on the input signal characteristics, the system achieves effective restoration across various coding conditions and formats, enhancing both speech and music content universally.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Described herein is a method of generating, in a dynamic range reduced domain, an enhanced multi-channel audio signal from an audio bitstream including a multi-channel audio signal, wherein the multi-channel audio signal comprises two or more channels, and wherein the method includes jointly enhancing the two or more channels of the dynamic range reduced raw multi-channel audio signal using a multi-channel Generator of a Generative Adversarial Network setting. Described herein are further a method for training a multi-channel Generator in a dynamic range reduced domain in a Generative Adversarial Network setting, an apparatus for generating, in a dynamic range reduced domain, an enhanced multi-channel audio signal from an audio bitstream including a multi-channel audio signal, respective systems and a computer program product.