Spectral Audio Decoding With Intelligent Gap Filling for Bandwidth Extension
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio codecs face limitations in bandwidth extension techniques, which restrict high-frequency content replacement and result in loss of detail or timbre, and require transformation into new domains, leading to computational complexity and memory issues, especially in mobile devices.
Innovation Solution
The proposed solution involves performing bandwidth extension in the same spectral domain as the core decoder, allowing full-rate core decoding and using Intelligent Gap Filling (IGF) to regenerate spectral portions, eliminating the need for downsampling and upsampling, and enabling efficient filling of spectral gaps using parametric data and source spectral ranges.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If bandwidth extension techniques are used to reduce bitrate, then coding efficiency is improved, but high-frequency content detail and timbre are lost
Solution Approach 1:
The spectrum is divided into multiple frequency bands (first spectral portions and second spectral portions), where different bands are processed differently. High-frequency bands are segmented and regenerated using parametric data and spectral shaping, allowing selective reconstruction of frequency content while maintaining overall coding efficiency.
Solution Approach 2:
The patent transforms the audio signal from time domain to spectral domain using MDCT, enabling parameter-based manipulation of frequency components. Spectral envelope parameters and scale factors are used to control the regeneration of high-frequency content, allowing flexible adjustment between compression ratio and frequency detail preservation.
2Reliability
If transformation into new domains is performed for bandwidth extension, then high-frequency reconstruction is enabled, but computational complexity and memory requirements increase
Solution Approach 1:
The patent merges the bandwidth extension process with the existing MDCT-based core decoding process. By performing spectral regeneration in the same spectral domain (MDCT domain) where the core decoder operates, additional domain transformations are eliminated, reducing computational complexity while maintaining high-frequency reconstruction capability.
Solution Approach 2:
The MDCT transform serves multiple functions: it performs the initial time-to-spectral domain conversion for core decoding, and also provides the spectral domain basis for bandwidth extension and high-frequency regeneration. This multi-functionality eliminates the need for separate transform stages, reducing overall computational burden.
3Productivity
If downsampling and upsampling are used in bandwidth extension, then processing efficiency is improved, but audio quality and waveform fidelity are degraded
Solution Approach 1:
The patent replaces the mechanical downsampling/upsampling process with a spectral-domain signal processing approach. Instead of changing the sampling rate in the time domain, the system operates entirely in the spectral domain using MDCT coefficients, applying parametric synthesis and spectral shaping to regenerate high-frequency content while maintaining the original sampling rate and waveform fidelity.
4Manufacturing precision
If spectral gap filling is performed in the spectral domain, then audio quality is improved, but additional processing stages are required
Solution Approach 1:
The spectral gap filling process is merged with the existing MDCT decoding pipeline. The frequency regenerator uses the same MDCT spectral coefficients produced by the core decoder, and the spectrum-time converter integrates with the existing inverse MDCT stage. This merging approach improves audio quality without adding independent processing stages.
Data Source
AI summary
An apparatus for decoding an encoded audio signal, includes a spectral domain audio decoder for generating a first decoded representation of a first set of first spectral portions, the decoded representation having a first spectral resolution; a parametric decoder for generating a second decoded representation of a second set of second spectral portions having a second spectral resolution being lower than the first spectral resolution; a frequency regenerator for regenerating every constructed second spectral portion having the first spectral resolution using a first spectral portion and spectral envelope information for the second spectral portion; and a spectrum time converter for converting the first decoded representation and the reconstructed second spectral portion into a time representation.


