Audio High-Frequency Reconstruction with Reduced Decoder Delay
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio encoding techniques, such as MPEG-4 AAC, face challenges with spectral band replication (SBR) that are not ideal for certain audio types, particularly musical content with low crossover frequencies, necessitating improved spectral band replication methods.
Innovation Solution
The integration of enhanced spectral band replication (eSBR) processing, including harmonic transposition and QMF-patching pre-flattening, is implemented in audio bitstreams to regenerate high frequency bands, adjusting spectral envelopes and adding noise and sinusoidal components for accurate audio reconstruction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If spectral band replication (SBR) is used for high frequency reconstruction, then compression efficiency is improved, but audio quality deteriorates for certain audio types such as musical content with low crossover frequencies
Solution Approach 1:
The patent changes the spectral translation parameters dynamically based on the audio signal characteristics. Specifically, it adjusts the translation start frequency, translation step size, and spectral envelope shaping parameters according to whether the content is music or speech, and according to the detected crossover frequency. This allows optimal high frequency reconstruction for different audio types while maintaining compression efficiency.
Solution Approach 2:
The patent introduces dynamic adaptation in the SBR process by detecting signal characteristics (music/speech classification, crossover frequency) and adjusting the spectral translation parameters accordingly. The system transitions from static parameter settings to dynamic parameter adjustment based on real-time signal analysis, improving audio quality for varying content types.
2Manufacturing precision
If enhanced spectral band replication (eSBR) processing is implemented, then audio quality is improved, but device complexity increases
Solution Approach 1:
The patent segments the spectral translation process into distinct stages: lowband decoding, spectral envelope extraction, harmonic transposition, spectral patching with pre-flattening, and highband synthesis. Each stage is independently optimized and controlled by specific parameters, making the complex eSBR process more manageable and implementable in practical decoders.
Solution Approach 2:
The patent applies pre-flattening to the spectral envelope before spectral patching. This preliminary processing step equalizes the spectral amplitude variations in the lowband signal, which facilitates more accurate high frequency reconstruction during the subsequent spectral translation stage, improving overall audio quality.
3Stability of the object's composition
If spectral patching with pre-flattening is applied, then high frequency signal stability is improved, but processing time increases
Solution Approach 1:
The patent adjusts the spectral translation step size and overlap parameters to optimize the balance between signal stability and processing efficiency. By carefully selecting these parameters based on signal characteristics, the system achieves stable high frequency reconstruction without excessive processing overhead.
Data Source
AI summary
A method for decoding an encoded audio bitstream is disclosed. The method includes receiving the encoded audio bitstream and decoding the audio data to generate a decoded lowband audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded lowband audio signal with an analysis filterbank to generate a filtered lowband audio signal. The method also includes extracting a flag indicating whether either spectral translation or harmonic transposition is to be performed on the audio data and regenerating a highband portion of the audio signal using the filtered lowband audio signal and the high frequency reconstruction metadata in accordance with the flag. The high frequency regeneration is performed as a post-processing operation with a delay of 3010 samples per audio channel.

