Two-Channel Audio Decoding With Intelligent Gap Filling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs face limitations in coding efficiency and quality at low bitrates due to restricted bandwidth extension techniques, which fail to accurately align tonal harmonics and introduce computational complexity, synchronization issues, and artifacts, especially in scenarios with varying correlation between source and target audio channels.
Innovation Solution
The implementation of Intelligent Gap Filling (IGF) technology, which analyzes and identifies different two-channel representations for spectral portions, uses parametric data to regenerate spectral content, and applies frequency regeneration in the same spectral domain as the core decoder, allowing for efficient filling of spectral gaps without transforming data into a new domain, thereby reducing computational complexity and preserving tonal components.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If bandwidth extension techniques are used to enable extended audio bandwidth at low bitrates, then coding efficiency is improved, but synchronization issues and artifacts are introduced
Solution Approach 1:
The patent replaces traditional time-domain bandwidth extension methods with frequency-domain processing. By working directly with spectral coefficients from the MDCT transform, the system avoids the synchronization issues that arise from time-domain signal manipulation and filtering operations.
Solution Approach 2:
The system changes the representation domain from time-domain waveforms to frequency-domain spectral coefficients. This parameter transformation enables more accurate control over spectral content and facilitates better alignment of tonal harmonics between source and target bands.
2Length of moving object
If traditional bandwidth extension methods are applied, then audio bandwidth is extended, but computational complexity increases
Solution Approach 1:
The patent merges the bandwidth extension process with the existing MDCT-based audio coding framework. By integrating frequency regeneration into the same transform domain already used for compression, the system avoids additional computational stages and domain transformations.
Solution Approach 2:
The system copies spectral coefficients from source bands to target bands through simple coefficient replication and scaling operations in the frequency domain, replacing complex time-domain filtering and signal generation processes.
3Length of moving object
If spectral patching is used to reconstruct high-frequency content, then bandwidth extension is achieved, but tonal alignment accuracy decreases
Solution Approach 1:
The system uses feedback from the decoded low-frequency signal to guide the regeneration of high-frequency content. By analyzing the spectral characteristics of the decoded signal and using them to inform the frequency regeneration process, the system achieves better alignment of tonal harmonics.
Solution Approach 2:
The patent performs preliminary decoding of the low-frequency portion of the signal to extract spectral characteristics before regenerating the high-frequency content. This preliminary analysis enables more accurate prediction and alignment of tonal components in the extended bandwidth.
4Adaptability or versatility
If data is transformed into a new domain for bandwidth extension, then processing flexibility is improved, but computational overhead increases
Solution Approach 1:
The patent makes the MDCT frequency domain representation serve multiple functions: both audio compression and bandwidth extension. This multi-functionality eliminates the need for separate processing domains while maintaining the flexibility needed for adaptive bandwidth extension.
Data Source
AI summary
An apparatus for generating a decoded two-channel signal includes: an audio processor for decoding an encoded two-channel signal to obtain a first set of first spectral portions; a parametric decoder for providing parametric data for a second set of second spectral portions and a two-channel identification identifying either a first or a second different two-channel representation for the second spectral portions; and a frequency regenerator for regenerating a second spectral portion depending on a first spectral portion of the first set of first spectral portions, the parametric data for the second portion and the two-channel identification for the second portion.


