Adaptive Spectral Tile Selection for Low-Bitrate Audio Gap Filling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs face limitations in maintaining audio quality at low bitrates due to restricted bandwidth extension techniques that fail to accurately align tonal harmonics and introduce computational complexity, synchronization issues, and artifacts, especially in non-tonal signals like unvoiced speech.
Innovation Solution
The implementation of an Intelligent Gap Filling (IGF) scheme that adapts frequency tile filling, allowing for signal-dependent source region selection and spectral envelope adjustment, enabling efficient encoding and decoding in the same spectral domain without the need for downsampling and upsampling, thereby preserving tonal components and reducing computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If bandwidth extension techniques are used to maintain audio quality at low bitrates, then audio quality is improved, but computational complexity increases and synchronization issues arise
Solution Approach 1:
The patent segments the spectral content into multiple frequency tiles, where each tile can be independently processed. This allows the system to apply bandwidth extension only to specific frequency regions that need it, rather than processing the entire spectrum, thereby reducing computational complexity while maintaining audio quality.
Solution Approach 2:
The patent applies different processing strategies to different frequency regions based on their specific characteristics. Tonal frequency tiles undergo phase alignment and spectral envelope adjustment, while non-tonal tiles use different reconstruction methods. This localized approach optimizes audio quality where needed while minimizing unnecessary computational overhead.
2Manufacturing precision
If traditional bandwidth extension with downsampling and upsampling is used, then audio quality is improved, but synchronization issues and artifacts are introduced
Solution Approach 1:
The patent extracts only the necessary spectral parameters and phase information from the low-frequency signal, avoiding the need for complete downsampling and upsampling operations. By extracting and manipulating only the essential components in the spectral domain, the system maintains synchronization accuracy while achieving bandwidth extension.
Solution Approach 2:
The patent introduces spectral envelope adjustment as an intermediary process between the low-frequency and high-frequency components. This intermediary step allows for smooth transitions and proper alignment without the harsh artifacts introduced by direct downsampling/upsampling operations, thereby maintaining synchronization accuracy.
3Manufacturing precision
If spectral patching is used to fill high-frequency regions, then audio quality is improved, but misalignment of tonal harmonics occurs
Solution Approach 1:
The patent performs preliminary phase alignment of tonal harmonics before spectral patching is applied. By pre-aligning the phase information and adjusting the spectral envelope in advance, the system ensures that when spectral patches are applied, the tonal harmonics remain properly aligned, avoiding the misalignment artifacts that would otherwise occur.
4Quantity of substance
If low bitrate encoding is used to reduce transmission costs, then bitrate is reduced, but audio quality deteriorates
Solution Approach 1:
The patent changes the representation parameters from time-domain samples to spectral-domain parameters. By encoding spectral envelopes, phase information, and selected frequency tile coefficients instead of raw audio samples, the system achieves much higher compression ratios while maintaining perceptual audio quality through intelligent gap filling and spectral reconstruction.
Data Source
AI summary
An apparatus for decoding an encoded signal includes: an audio decoder for decoding an encoded representation of a first set of first spectral portions to obtain a decoded first set of first spectral portions; a parametric decoder for decoding an encoded parametric representation of a second set of second spectral portions to obtain a decoded representation of the parametric representation, wherein the parametric information includes, for each target frequency tile, a source region identification as a matching information; and a frequency regenerator for regenerating a target frequency tile using a source region from the first set of first spectral portions identified by the matching information.


