Reconstruction-Band Energy Coding for Low-Bitrate Audio Spectra
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs face limitations in coding high-frequency content at low bitrates, leading to loss of detail and timbre due to the restriction of bandwidth extension techniques, which require transformation into new domains, increasing computational complexity and memory requirements, and fail to accurately align tonal harmonics.
Innovation Solution
The Intelligent Gap Filling (IGF) method allows for encoding and decoding in the same spectral domain, using frequency regeneration and parametric data to fill spectral gaps, enabling the encoding of perceptually important tonal portions across the entire spectrum without additional domain transformations, thus maintaining audio quality at low bitrates.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If bandwidth extension techniques are used to code high-frequency content at low bitrates, then compression efficiency is improved, but audio quality deteriorates due to loss of detail and timbre
Solution Approach 1:
The spectrum is divided into multiple reconstruction bands, each processed independently with dedicated energy information values. This segmentation allows precise control of high-frequency content in each band, preserving timbre and detail while maintaining low bitrate through selective coding of only essential energy parameters per band.
Solution Approach 2:
The invention uses parametric representation of high-frequency content by encoding energy information values for each reconstruction band instead of full spectral data. This parameter-based approach dramatically reduces bitrate while preserving perceptually important characteristics through efficient energy distribution control across frequency bands.
2Loss of energy
If bandwidth extension techniques transform into new domains for high-frequency reconstruction, then coding efficiency is improved, but device complexity increases due to additional transformations
Solution Approach 1:
The invention merges the bandwidth extension process with the existing MDCT decoding framework, eliminating the need for separate domain transformations. High-frequency reconstruction is performed directly in the MDCT spectral domain by adjusting energy information values, integrating HF synthesis with the core decoding pipeline and reducing computational overhead.
Solution Approach 2:
The MDCT decoder itself performs high-frequency reconstruction by utilizing energy information values for each reconstruction band. The existing decoder structure is enhanced to self-generate high-frequency content through parametric control without requiring external transformation stages or additional processing domains, making the system self-sufficient and computationally efficient.
3Loss of energy
If bandwidth extension techniques are used, then low bitrate coding is enabled, but tonal harmonics are not accurately aligned
Solution Approach 1:
The invention dynamically adjusts energy information values for each reconstruction band based on the specific spectral characteristics and tonal content of that band. This dynamic, adaptive approach allows precise alignment of tonal harmonics by individually controlling energy distribution in each band rather than applying uniform transformation, ensuring accurate frequency representation across the entire spectrum.
Solution Approach 2:
Each reconstruction band is treated with local quality control through dedicated energy information values tailored to that specific frequency region. This localized approach preserves tonal harmonics in bands where they are present while allowing noise-like characteristics in other bands, matching the perceptual and spectral properties of each frequency region individually for accurate tonal alignment.
Data Source
AI summary
An apparatus for decoding an encoded audio signal having an encoded representation of a first set of first spectral portions and an encoded representation of parametric data indicating spectral energies for a second set of second spectral portions, has: an audio decoder for decoding the encoded representation of the first set of the first spectral portions to obtain a first set of first spectral portions and for decoding the encoded representation of the parametric data to obtain a decoded parametric data for the second set of second spectral portions indicating, for individual reconstruction bands, individual energies; a frequency regenerator for reconstructing spectral values in a reconstruction band having a second spectral portion using a first spectral portion of the first set of the first spectral portions and an individual energy for the reconstruction band, the reconstruction band having a first spectral portion and the second spectral portion.


