Audio Decoder High Frequency Reconstruction Metadata
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing technologies, such as the MPEG-4 AAC standard, face challenges in efficiently reconstructing high frequency components of audio signals, particularly for musical content with low crossover frequencies, where spectral band replication techniques may not be ideal.
Innovation Solution
The method involves decoding an encoded audio bitstream, extracting high frequency reconstruction metadata, and filtering the decoded lowband audio signal to regenerate the highband portion using either spectral translation or harmonic transposition, as indicated by a flag in the metadata.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If spectral band replication is used for high frequency reconstruction, then compression efficiency is improved, but audio quality deteriorates for musical content with low crossover frequencies
Solution Approach 1:
The patent applies dynamics by making the HFR technique selectable based on content type. The system dynamically switches between spectral translation (for speech/low complexity content) and harmonic transposition (for musical content with low crossover frequencies), allowing optimal performance for different audio scenarios rather than using a fixed approach
Solution Approach 2:
The patent changes the fundamental parameter of HFR methodology from a fixed spectral translation approach to a variable approach that selects between spectral translation and harmonic transposition based on content characteristics. This parameter change enables adaptation to different audio types, improving both compression efficiency and audio quality for specific content genres
2Quantity of substance
If only low frequency components are encoded and transmitted, then data rate is reduced, but high frequency reconstruction accuracy deteriorates
Solution Approach 1:
The patent uses an intermediary approach by transmitting additional metadata about the original high frequency content characteristics alongside the low frequency encoded audio. This metadata acts as an intermediary that guides the reconstruction process, enabling more accurate high frequency regeneration without increasing the primary audio data rate
Solution Approach 2:
The patent applies preliminary action by having the encoder analyze and store characteristics of the original high frequency content before transmission. This preliminary analysis creates a template or guide that the decoder uses to accurately reconstruct high frequencies, preparing reconstruction information in advance rather than attempting to transmit full high frequency data
3Device complexity
If spectral translation is used for high frequency reconstruction, then decoding complexity is reduced, but reconstruction quality deteriorates for certain audio types
Solution Approach 1:
The patent makes the decoding complexity dynamic by selecting the HFR technique based on content type indicators. For speech and low complexity content, simpler spectral translation is used. For musical content with low crossover frequencies, harmonic transposition is selected, accepting higher decoding complexity in exchange for significantly improved reconstruction quality for that specific content type
Data Source
AI summary
A method for decoding an encoded audio bitstream is disclosed. The method includes receiving the encoded audio bitstream and decoding the audio data to generate a decoded lowband audio signal. The method further includes extracting high frequency reconstruction metadata and filtering the decoded lowband audio signal with an analysis filterbank to generate a filtered lowband audio signal. The method also includes extracting a flag indicating whether either spectral translation or harmonic transposition is to be performed on the audio data and regenerating a highband portion of the audio signal using the filtered lowband audio signal and the high frequency reconstruction metadata in accordance with the flag. The high frequency regeneration is performed as a post-processing operation with a delay of 3010 samples per audio channel.

