Spectral Audio Decoding Using Frequency Tiles for Bandwidth Extension
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs face limitations in bandwidth extension techniques, as they restrict high-frequency content replacement and require transformation into different spectral domains, leading to increased computational complexity, memory requirements, and artifacts such as spatial segregation and pre-echoes, especially in mobile devices.
Innovation Solution
The Intelligent Gap Filling (IGF) method performs bandwidth extension in the same spectral domain as the core decoder, analyzing audio signals to encode tonal portions with high resolution and noisy components with low spectral resolution, using frequency tiles for reconstruction and noise filling to fill spectral gaps, while maintaining the spectral envelope and energy distribution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of substance
If bandwidth extension techniques are used to reduce bitrate, then compression efficiency is improved, but computational complexity and memory requirements increase
Solution Approach 1:
The patent segments the spectral representation into different frequency bands (low-frequency and high-frequency regions) and applies different processing strategies to each segment. The low-frequency region is coded with higher precision while the high-frequency region uses bandwidth extension techniques, allowing selective application of complex algorithms only where needed.
Solution Approach 2:
The patent applies local quality by maintaining high coding precision in the low-frequency region where perceptual importance is higher, while using approximate bandwidth extension methods in the high-frequency region. This localized differentiation of quality levels reduces overall computational complexity while preserving perceptual audio quality.
2Loss of substance
If spectral patching is used to reconstruct high-frequency content, then bandwidth extension is achieved, but artifacts such as spatial segregation and pre-echoes occur
Solution Approach 1:
The patent performs preliminary action by pre-processing the low-frequency signal to extract spectral characteristics and temporal envelope information before reconstructing the high-frequency content. This preliminary extraction of relevant features allows for more accurate bandwidth extension that preserves temporal continuity and reduces artifacts like pre-echoes.
Solution Approach 2:
The patent implements feedback mechanisms where the reconstructed high-frequency signal is compared with the original low-frequency signal to adjust spectral parameters and minimize artifacts. This feedback loop allows iterative optimization of the bandwidth extension process to reduce spatial segregation and pre-echo artifacts.
3Loss of substance
If transformation into different spectral domains is performed for bandwidth extension, then high-frequency content is reconstructed, but device complexity increases
Solution Approach 1:
The patent makes the spectral representation multi-functional by using the same frequency domain data structure for both low-frequency coding and high-frequency bandwidth extension. This universal approach eliminates the need for separate transformation stages and different memory structures, reducing device complexity while achieving full-bandwidth reconstruction.
Data Source
Figure 1A~1B
Figure 2A
Figure 2B
AI summary
An apparatus for decoding an encoded audio signal, comprises a spectral domain audio decoder (112) for generating a first decoded representation of a first set of first spectral portions, the decoded representation having a first spectral resolution; a parametric decoder (114) for generating a second decoded representation of a second set of second spectral portions having a second spectral resolution being lower than the first spectral resolution; a frequency regenerator (116) for regenerating every constructed second spectral portion having the first spectral resolution using a first spectral portion and spectral envelope information for the second spectral portion; and a spectrum time converter (118) for converting the first decoded representation and the reconstructed second spectral portion into a time representation.