Audio Bandwidth Extension Using IGF and Temporal Noise Shaping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs face limitations in bandwidth extension techniques, particularly in maintaining high-frequency detail and timbre, due to restricted transformation domains and limited temporal control, leading to pre- or post-echoes and increased computational complexity.
Innovation Solution
The implementation of Intelligent Gap Filling (IGF) technology, which combines Temporal Noise Shaping (TNS) or Temporal Tile Shaping (TTS) with high-frequency reconstruction, allowing bandwidth extension in the same spectral domain as the core decoder, and using frequency tiles for gap filling, thereby reducing echoes and computational complexity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If bandwidth extension techniques are used to reduce bitrate, then coding efficiency is improved, but temporal continuity and audio quality deteriorate due to pre- or post-echoes
Solution Approach 1:
The patent applies temporal noise shaping by predicting and removing noise components before they manifest as audible artifacts. The predictor analyzes the spectral content and predicts noise in upcoming frames, allowing pre-echoes to be removed before they occur. Similarly, post-echoes are predicted and removed in advance, preventing temporal discontinuities that would degrade audio quality.
Solution Approach 2:
The patent implements a feedback mechanism where the predictor continuously monitors the decoded audio signal for noise components and feeds this information back to the noise shaper. The system analyzes the spectral content of previous frames and uses this feedback to adjust noise prediction and removal in current frames, maintaining temporal continuity while preserving coding efficiency.
2Measurement precision
If transformation to second domain is applied for high-frequency reconstruction, then spectral accuracy is improved, but device complexity increases
Solution Approach 1:
The patent merges the noise prediction and spectral shaping operations into a unified temporal noise shaping process that operates within the same transform domain as the audio coder. Instead of transforming to a second domain for high-frequency reconstruction, the system combines predictive noise removal with spectral envelope modification in the MDCT domain, reducing computational complexity while maintaining spectral accuracy.
Solution Approach 2:
The patent introduces a spectral envelope as an intermediary that mediates between the time-domain audio signal and the frequency-domain representation. The temporal noise shaper modifies the spectral envelope to remove noise components without requiring transformation to a second domain. This intermediary approach maintains spectral accuracy while avoiding the computational overhead of multiple transform domains.
Data Source
Figure 1A~1B
Figure 2A
Figure 2B
AI summary
An apparatus for decoding an encoded audio signal, comprises: a spectral domain audio decoder (602) for generating a first decoded representation of a first set of first spectral portions being spectral prediction residual values; a frequency regenerator (604) for generating a reconstructed second spectral portion using a first spectral portion of the first set of first spectral portions, wherein the reconstructed second spectral portion additionally comprises spectral prediction residual values; and an inverse prediction filter (606) for performing an inverse prediction over frequency using the spectral residual values for the first set of first spectral portions and the reconstructed second spectral portion using prediction filter information (607) included in the encoded audio signal.