Hierarchical Audio Codec for Scalable Multichannel Bit Streams
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional audio compression methods are not scalable to address varying channel capacities, storage needs, or different data rates, as they fix the data rate and audio quality at the time of compression, limiting flexibility in delivering audio over limited bandwidth channels or storing audio content.
Innovation Solution
A method using a hierarchical filterbank to decompose audio signals into tonal and residual components, ranking and quantizing them based on psychoacoustic criteria, and forming a master bit stream that can be scaled to arbitrary data rates by eliminating low-ranking components, with differential coding and joint channel coding for multichannel audio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional audio compression methods are used, then audio quality is fixed at the time of compression, but the data rate cannot be adjusted for varying channel capacities
Solution Approach 1:
The audio signal is decomposed into multiple frequency subbands using a filter bank, and each subband is independently encoded with its own quantization parameters. This segmentation allows selective removal or coarsening of specific subbands to achieve different data rates while maintaining audio quality in critical frequency regions.
Solution Approach 2:
The quantization parameters and bit allocation are made dynamic rather than fixed, allowing the encoder to adjust the number of bits allocated to different frequency subbands based on the desired data rate. This dynamic adaptation enables the same compressed audio representation to serve multiple data rate requirements.
2Adaptability or versatility
If a high data rate bit stream is created, then audio quality is improved, but the bit stream cannot be efficiently scaled to lower data rates
Solution Approach 1:
During the initial compression process, the audio signal is decomposed into multiple frequency subbands with quantization parameters that preserve information at different levels of detail. This preliminary organization of data allows the bit stream to be later scaled to different data rates by selectively removing or coarsening less critical subbands without requiring complete re-encoding.
Solution Approach 2:
Different frequency subbands are assigned different quantization qualities based on their perceptual importance. Critical frequency regions maintain high quality even at lower data rates, while less critical regions can be coarsely quantized or removed. This local quality differentiation enables scalable bit streams that maintain acceptable audio quality across a range of data rates.
3Adaptability or versatility
If multiple compression algorithms are used to create layered bit streams, then a wider range of data rates is supported, but the device complexity increases
Solution Approach 1:
A single compression algorithm with configurable parameters is designed to perform multiple functions: it can produce bit streams at different data rates from the same input signal, and the same algorithm structure can be used for both high and low data rate encoding. This universal approach eliminates the need to manage multiple separate compression algorithms while still supporting a wide range of data rates.
Data Source
AI summary
A method for compressing audio input signals to form a master bit stream that can be scaled to form a scaled bit stream having an arbitrarily prescribed data rate. A hierarchical filterbank decomposes the input signal into a multi-resolution time/frequency representation from which the encoder can efficiently extract both tonal and residual components. The components are ranked and then quantized with reference to the same masking function or different psychoacoustic criteria. The selected tonal components are suitably encoded using differential coding extended to multichannel audio. The time-sample and scale factor components that make up the residual components are encoded using joint channel coding (JCC) extended to multichannel audio. A decoder uses an inverse hierarchical filterbank to reconstruct the audio signals from the tonal and residual components in the scaled bit stream.


