Hierarchical Audio Coding Bit Allocation for Music Quality
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing G.729.1 codec's perceptual weighting filter, based on an energy criterion, is not optimal for coding music signals, leading to insufficient quantization noise shaping and reduced quality in the high band (4000-7000 Hz) due to its reliance on CELP-like speech production models.
Innovation Solution
A method is introduced to hierarchically code digital audio signals by calculating a frequency masking threshold and determining perceptual importance per frequency sub-band, allowing for improved bit allocation in the improvement coding layer, which enhances the quality of coding by considering the signal-to-mask ratio and energy allocation, while maintaining compatibility with existing standards.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If CELP-like speech production models with energy criterion are used for coding, then compatibility with existing standards is maintained, but perceptual quality for music signals deteriorates due to insufficient quantization noise shaping
Solution Approach 1:
The audio signal is divided into frequency sub-bands, with different coding strategies applied to each band. The improvement layer specifically targets high-frequency sub-bands (4000-7000 Hz) where the core CELP codec performs poorly, allowing selective enhancement without compromising overall compatibility
Solution Approach 2:
Different coding approaches are applied to different frequency regions. The improvement layer uses perceptual masking-based bit allocation specifically for high-frequency bands where music signals require better quality, while the core layer maintains speech-oriented CELP coding for lower bands, optimizing each region according to its specific requirements
2Device complexity
If bit allocation is based on energy criterion, then computational simplicity is maintained, but perceptual importance of frequency sub-bands is not optimized leading to poor noise shaping
Solution Approach 1:
The bit allocation criterion is changed from pure energy-based allocation to perceptual masking-based allocation in the improvement layer. By calculating masking thresholds and using signal-to-mask ratios, the system optimizes bit distribution according to perceptual importance rather than just energy content, significantly improving noise shaping quality in critical frequency regions
Solution Approach 2:
The patent introduces a hierarchical dimension to bit allocation, with the core layer using simple energy-based allocation and the improvement layer adding perceptual masking-based allocation. This layered approach allows the system to operate at different complexity levels depending on available bandwidth, transforming a single-dimension problem into a multi-dimensional solution space
3Manufacturing precision
If the codec operates in widened band (50-7000 Hz), then audio quality is improved, but effectiveness for super-widened band (50-14000 Hz) is insufficient due to high band limitations
Solution Approach 1:
The improvement layer pre-processes and enhances the high-frequency band (4000-7000 Hz) using perceptual masking and optimized bit allocation before final synthesis. This preliminary enhancement of the high band ensures that when band extension to super-widened frequencies (7000-14000 Hz) is applied, the foundation is already optimized, enabling effective extension to higher frequencies
Solution Approach 2:
The improvement layer acts as an intermediary between the core CELP codec and the final audio output, specifically enhancing the high-frequency transition region. This intermediary processing ensures smooth spectral continuity and proper noise shaping in the 4000-7000 Hz range, which serves as a bridge for effective extension to super-widened bands
Data Source
AI summary
A method of hierarchical coding of a digital audio frequency input signal into several frequency sub-bands, including a core coding of the input signal according to a first throughput and at least one enhancement coding of higher throughput, of a residual signal. The core coding uses a binary allocation according to an energy criterion. The method includes for the enhancement coding: calculating a frequency-based masking threshold for at least part of the frequency bands processed by the enhancement coding; determining a perceptual importance per frequency sub-band as a function of the masking threshold and as a function of the number of bits allocated for the core coding; binary allocation of bits in the frequency sub-bands processed by the enhancement coding, as a function of the perceptual importance determined; and coding the residual signal according to the bit allocation. Also provided are a decoding method, a coder and a decoder.


