Audio Decoding with Temporal Noise Shaping for High-Frequency Reconstruction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs face limitations in bandwidth extension techniques, particularly in maintaining high-frequency detail and timbre at low bitrates, due to restricted spectral patching and transformation requirements, leading to pre- or post-echoes and increased computational complexity.
Innovation Solution
The implementation of Intelligent Gap Filling (IGF) technology, which performs bandwidth extension in the same spectral domain as the core decoder, using temporal noise shaping and tile shaping to reconstruct high-frequency content, reducing echoes and computational overhead by filling spectral gaps with parametric data and source spectral portions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spectral patching is used for bandwidth extension, then high-frequency content can be reconstructed, but pre-echoes and post-echoes occur and computational complexity increases
Solution Approach 1:
The patent segments the spectral reconstruction process into two distinct parts: (1) copying spectral patches from low-frequency bands to high-frequency bands, and (2) applying temporal noise shaping separately to control quantization noise. This segmentation allows each part to be optimized independently, reducing the harmful echo effects while maintaining reconstruction accuracy.
Solution Approach 2:
The patent introduces temporal noise shaping as an intermediary process between spectral copying and final reconstruction. This intermediary step shapes the quantization noise in the time domain before the spectral patches are applied, preventing the direct transmission of noise artifacts that cause pre- and post-echoes.
2Measurement precision
If bandwidth extension is performed with domain transformation, then audio quality is improved, but computational complexity increases
Solution Approach 1:
The patent makes the temporal noise shaping filter bank universal by designing it to operate directly in the same domain as the core decoder's MDCT coefficients. This eliminates the need for separate domain transformations, allowing the same filter bank structure to handle both spectral copying and noise shaping without additional computational overhead.
Solution Approach 2:
The patent merges the temporal noise shaping operation with the spectral copying operation by applying the shaping filter to the MDCT coefficients directly before or during the patching process. This combining of operations eliminates redundant domain transformations and reduces overall computational complexity.
3Productivity
If spectral copying is used for high-frequency reconstruction, then bandwidth extension is achieved, but timbre and color are not well maintained
Solution Approach 1:
The patent applies local quality control by using temporal noise shaping with different parameters for different frequency regions. The noise shaping filter is designed to adapt to local spectral characteristics, allowing each frequency band to maintain its unique timbre and color properties while still achieving efficient bandwidth extension.
Data Source
AI summary
An apparatus for decoding an encoded audio signal, includes: a spectral domain audio decoder for generating a first decoded representation of a first set of first spectral portions being spectral prediction residual values; a frequency regenerator for generating a reconstructed second spectral portion using a first spectral portion of the first set of first spectral portions, wherein the reconstructed second spectral portion additionally includes spectral prediction residual values; and an inverse prediction filter for performing an inverse prediction over frequency using the spectral residual values for the first set of first spectral portions and the reconstructed second spectral portion using prediction filter information included in the encoded audio signal.


