Integrated High-Frequency Reconstruction for Low-Delay Audio Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Spectral band replication techniques in audio coding, such as SBR, are not ideal for certain audio types like musical content with low crossover frequencies, necessitating improved methods for high frequency reconstruction.
Innovation Solution
The method involves decoding an encoded audio bitstream, extracting high frequency reconstruction metadata, filtering the lowband audio signal, and regenerating the highband portion using analysis and synthesis filterbanks, with options for spectral translation or harmonic transposition based on metadata flags, and incorporating enhanced SBR processing like harmonic transposition and QMF-patching.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If spectral band replication (SBR) is used for high frequency reconstruction, then compression efficiency is improved, but audio quality deteriorates for certain audio types like musical content with low crossover frequencies
Solution Approach 1:
The patent implements dynamic switching between different HFR techniques (spectral translation, harmonic transposition, QMF-patching) based on the characteristics of the audio content. The system analyzes the audio signal properties and selects the most appropriate reconstruction method in real-time, allowing optimal performance across different audio types while maintaining compression efficiency.
Solution Approach 2:
The patent changes the fundamental parameters of high frequency reconstruction by offering multiple techniques with different characteristics. Spectral translation copies QMF subbands directly, harmonic transposition uses frequency transposition operations, and QMF-patching uses different patching strategies. These parameter changes allow the system to adapt to different audio content requirements.
2Ease of manufacture
If spectral translation is used for high frequency reconstruction, then processing simplicity is maintained, but spectral accuracy deteriorates for musical content
Solution Approach 1:
The system dynamically selects between spectral translation (simpler) and harmonic transposition (more accurate) based on the detected audio content characteristics. For musical content with low crossover frequencies, the system switches to harmonic transposition to maintain spectral accuracy, while for other content it uses spectral translation for simplicity.
3Manufacturing precision
If enhanced HFR processing is implemented, then audio quality is improved, but device complexity increases
Solution Approach 1:
The patent designs a universal decoder architecture that can perform multiple HFR techniques (spectral translation, harmonic transposition, QMF-patching) within a single integrated system. This multi-functional approach allows the decoder to adapt to different audio content requirements without requiring separate hardware implementations for each technique.
Solution Approach 2:
The patent introduces an audio analysis intermediary that characterizes the input audio signal and determines the appropriate HFR technique to use. This intermediary component simplifies the overall system by automatically selecting the optimal processing path based on audio content, reducing the complexity burden on the main decoding architecture.
4Adaptability or versatility
If new HFR techniques are introduced, then adaptability to different audio types is improved, but backward compatibility becomes challenging
Solution Approach 1:
The decoder is designed with universal functionality to handle both legacy SBR bitstreams and enhanced HFR bitstreams. The system can process standard spectral translation requests while also supporting harmonic transposition and QMF-patching when indicated by appropriate metadata flags, ensuring broad compatibility across different audio formats.
Solution Approach 2:
The patent uses metadata flags as intermediaries to communicate the desired HFR technique from the encoder to the decoder. These flags allow the system to maintain backward compatibility with legacy decoders that ignore the flags and use default spectral translation, while enabling enhanced techniques when the flags are properly interpreted.
Data Source
AI summary
A method for decoding an encoded audio bitstream is disclosed. The method includes receiving the encoded audio bitstream and decoding the audio data to generate a decoded lowband audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded lowband audio signal with an analysis filterbank to generate a filtered lowband audio signal. The method also includes extracting a flag indicating whether either spectral translation or harmonic transposition is to be performed on the audio data and regenerating a highband portion of the audio signal using the filtered lowband audio signal and the high frequency reconstruction metadata in accordance with the flag. The high frequency regeneration is performed as a post-processing operation with a delay of 3010 samples per audio channel.

