Hierarchical Audio Coding With Residual-Aware Bit Allocation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional hierarchical audio coding schemes suffer from low coding efficiency due to independent coding schemes for the core and extended layers, which do not consider the residual perception distribution characteristics of the signal, leading to suboptimal bit allocation and reduced audio quality, especially at medium to low code rates.
Innovation Solution
A method that divides frequency domain coefficients after Modified Discrete Cosine Transform into core and extended layers, using bit allocation with variable step lengths and quantization techniques like pyramid lattice vector quantization and sphere lattice vector quantization, based on auditory perceptive characteristics, to improve coding efficiency and audio quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If independent coding schemes are used for core layer and extended layer, then device complexity is reduced and ease of manufacture is improved, but coding efficiency deteriorates and audio quality is reduced
Solution Approach 1:
The patent merges the core layer and extended layer coding schemes into a unified framework. The extended layer coding utilizes the residual signal from core layer decoding, creating an integrated system where both layers work together rather than independently. This combining approach resolves the contradiction by maintaining implementation feasibility while significantly improving coding efficiency and audio quality through coordinated processing of both layers.
Solution Approach 2:
The patent applies preliminary action by first performing core layer coding and decoding to generate a residual signal, then using this residual signal as the basis for extended layer coding. This sequential approach where the first layer's output prepares the data for the second layer enables the system to achieve high coding efficiency while maintaining reasonable implementation complexity, as each layer builds upon the previous layer's work.
2Device complexity
If uniform bit allocation is used across all frequency sub-bands, then device complexity is reduced, but coding efficiency deteriorates due to ignoring residual perception distribution characteristics
Solution Approach 1:
The patent applies local quality by implementing non-uniform bit allocation that adapts to the local characteristics of different frequency sub-bands. The system calculates residual perception distribution characteristics for each sub-band and allocates bits accordingly, giving more bits to sub-bands that require higher precision and fewer bits to sub-bands that can tolerate coarser quantization. This localized adaptation resolves the contradiction by improving coding efficiency through characteristic-aware allocation while keeping the overall system complexity manageable.
Solution Approach 2:
The patent introduces dynamics into the bit allocation process by making it adaptive rather than static. The bit allocation scheme dynamically adjusts based on the residual perception distribution characteristics of each frequency sub-band, allowing the system to optimize coding efficiency for varying signal conditions while maintaining a relatively simple overall structure through automated adaptation.
3Ease of operation
If core layer information is not utilized in extended layer coding, then independence between layers is maintained improving ease of operation, but coding efficiency is reduced
Solution Approach 1:
The patent uses preliminary action by having the core layer processing occur first and generate a residual signal that is then fed into the extended layer coding process. This sequential dependency allows the extended layer to utilize core layer information effectively, improving coding efficiency while maintaining a clear operational flow that is relatively easy to implement and manage.
Solution Approach 2:
The patent applies the nested doll principle by embedding the extended layer coding within the overall hierarchical structure that begins with core layer processing. The extended layer is nested within the framework established by the core layer, with the residual signal from core decoding serving as the input foundation for extended layer processing. This nesting allows efficient information utilization while maintaining structured, manageable layer relationships.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A hierarchical audio coding, decoding method and system are provided. Said method includes: dividing frequency domain coefficients of an audio signal after MDCT into a plurality of coding sub-bands, quantizing and coding amplitude envelope values of coding sub-bands; allocating bits to each coding sub-band of the core layer, quantizing and coding core layer frequency domain coefficients to obtain coded bits of core layer frequency domain coefficients; calculating the amplitude envelope value of each coding sub-band of the core layer residual signal; allocating bits to each coding sub-band of the extended layer, quantizing and coding the extended layer coding signal to obtain coded bits of the extended layer coding signal; multiplexing and packing amplitude value envelope coded bits of each coding sub-band composed by core layer and extended layer frequency domain coefficients, core layer frequency coefficients coded bits, and extended layer coding signal coded bits, then transmitting to the decoding end.