Adaptive Spectral Tile Filling for Low-Bitrate Audio Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs face limitations in bandwidth extension techniques, particularly in maintaining high-frequency detail and timbre accuracy at low bitrates, due to restrictions in replacing perceptually important content above a given cross-over frequency and requiring transformation into new domains, which leads to computational complexity and synchronization issues.
Innovation Solution
The implementation of an Intelligent Gap Filling (IGF) scheme that adapts frequency tile filling, allowing for signal-dependent source region selection and processing, enabling efficient encoding and decoding in the same spectral domain without the need for downsampling or upsampling, and utilizing parametric data for spectral gap filling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If bandwidth extension techniques are used to reduce bitrate, then compression efficiency is improved, but high-frequency detail and timbre accuracy deteriorate
Solution Approach 1:
The patent applies local quality by differentiating between tonal and non-tonal spectral regions. Tonal regions are preserved with high fidelity using waveform coding, while non-tonal regions use parametric coding. This localized approach ensures that perceptually important tonal components maintain their detail and timbre accuracy even at low bitrates, while achieving compression in non-critical regions.
Solution Approach 2:
The spectrum is segmented into multiple frequency tiles, with different coding strategies applied to different tiles. This segmentation allows the system to selectively apply bandwidth extension only where necessary, preserving high-frequency detail in tonal tiles while using compression in non-tonal tiles, thus resolving the contradiction between compression efficiency and frequency detail accuracy.
2Adaptability or versatility
If spectral patching is used for bandwidth extension, then bandwidth extension capability is improved, but computational complexity increases
Solution Approach 1:
The patent implements dynamic spectral selection where the source region for each target frequency tile is adaptively chosen based on spectral similarity metrics. This dynamic approach allows the system to efficiently identify matching spectral regions without exhaustive search, reducing computational complexity while maintaining versatile bandwidth extension capability across different audio signals.
Solution Approach 2:
The system uses parametric representation of spectral regions and changes parameters such as spectral similarity thresholds and tile boundaries adaptively. This allows bandwidth extension to be achieved through parameter manipulation rather than complex signal processing operations, reducing computational complexity while preserving extension capability.
3Manufacturing precision
If transformation into new domains is performed for bandwidth extension, then bandwidth extension accuracy is improved, but synchronization issues arise
Solution Approach 1:
The patent performs bandwidth extension within the same spectral domain (MDCT domain) used for the core audio coding, making the system multi-functional. This universal approach eliminates the need for separate transformation stages, thereby maintaining synchronization reliability while achieving accurate bandwidth extension through spectral replication and parametric synthesis in the unified domain.
4Adaptability or versatility
If downsampling or upsampling is used in bandwidth extension, then bandwidth extension capability is improved, but computational complexity increases
Solution Approach 1:
The patent extracts the essential spectral characteristics from the low-frequency region and directly synthesizes the high-frequency content through parametric modeling and spectral replication. This extraction approach avoids the need for time-consuming downsampling/upsampling operations, reducing computational complexity while maintaining bandwidth extension capability through intelligent spectral synthesis.
Data Source
AI summary
An apparatus for decoding an encoded signal includes: an audio decoder for decoding an encoded representation of a first set of first spectral portions to obtain a decoded first set of first spectral portions; a parametric decoder for decoding an encoded parametric representation of a second set of second spectral portions to obtain a decoded representation of the parametric representation, wherein the parametric information includes, for each target frequency tile, a source region identification as a matching information; and a frequency regenerator for regenerating a target frequency tile using a source region from the first set of first spectral portions identified by the matching information.


