Hierarchical Audio Coding with Variable Frame Durations
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional hierarchical audio coding techniques are limited in flexibility, as they only consider a single strategy for formatting bit streams, which restricts the transmission hierarchy and prioritization of audio signal layers, leading to suboptimal compression ratios and bit rate efficiency.
Innovation Solution
A method for hierarchically coding audio signals that allows multiple strategies for formatting bit streams by inserting indications of frame duration order, enabling frames of varying durations across base and enhancement levels, and using sinusoidal breakdown and residual signal coding with analysis filters to achieve higher compression ratios and adaptability.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If conventional hierarchical coding techniques are used with a single bit stream formatting strategy, then the transmission hierarchy is simplified, but the compression ratio and bit rate efficiency are suboptimal
Solution Approach 1:
The patent applies dynamics by enabling multiple bit stream formatting strategies that can be dynamically selected based on transmission conditions. The system transitions from a static single-strategy approach to a dynamic multi-strategy approach where the encoder can adaptively choose different framing modes (e.g., frame-based, packet-based, or custom hierarchies) to optimize compression ratio and bit rate efficiency for specific transmission scenarios.
Solution Approach 2:
The patent changes the parameter of bit stream formatting by introducing multiple configurable parameters such as frame duration, layer hierarchy depth, and packet segmentation size. These parameters can be adjusted to create different formatting strategies, allowing the system to optimize compression performance for various transmission conditions without fixing a single rigid structure.
2Device complexity
If frames of equal duration are used across all levels, then the coding structure is simplified, but the adaptability to different transmission conditions is reduced
Solution Approach 1:
The patent applies segmentation by dividing the audio signal into frames of potentially different durations across various hierarchy levels. Instead of using uniform frame lengths, the system segments the signal adaptively, allowing base layer frames and enhancement layer frames to have different durations optimized for their respective coding requirements and transmission conditions.
Solution Approach 2:
The system dynamically adjusts frame durations based on transmission conditions and signal characteristics. The encoder can select different framing modes where base layer frames may have longer durations while enhancement layers use shorter frames, or vice versa, depending on the available bandwidth, latency requirements, and audio content analysis.
3Device complexity
If a fixed hierarchy of base layer and enhancement layers is used, then the transmission protocol is simplified, but the optimization of bit rate allocation is limited
Solution Approach 1:
The patent introduces dynamic bit rate allocation strategies where the encoder can adaptively determine the optimal number of enhancement layers and their respective bit rates based on available transmission bandwidth and audio quality requirements. This dynamic allocation allows the system to optimize bit rate efficiency by concentrating bits where they provide the most perceptual benefit rather than using a fixed predetermined allocation.
Solution Approach 2:
The system changes the parameter of bit rate allocation by introducing configurable parameters such as the number of enhancement layers, bit rate per layer, and priority weights. These parameters can be adjusted to create different allocation strategies optimized for specific transmission conditions, allowing flexible adaptation to varying bandwidth availability and quality requirements.
4Device complexity
If only single-strategy bit stream formatting is implemented, then the implementation is simpler, but the quality of audio reconstruction at different bit rates is suboptimal
Solution Approach 1:
The patent applies segmentation by separating the audio coding into distinct base layer and enhancement layers with different coding strategies. The base layer uses a simplified coding approach for robust transmission, while enhancement layers use more sophisticated coding for quality improvement. This segmentation allows the system to achieve high audio reconstruction quality at different bit rates by selectively transmitting base layer only for low bandwidth or both layers for high bandwidth conditions.
Solution Approach 2:
The system changes the parameter of coding strategy by introducing configurable parameters such as coding mode selection, quantization precision per layer, and transformation type. These parameters can be adjusted to optimize audio reconstruction quality for different transmission conditions, allowing the system to adapt the coding parameters to match the available bandwidth and quality requirements.
Data Source
AI summary
Hierarchical coding of a source audio signal in the form of a data stream including a base level and at least two hierarchical enhancement levels, each of the levels being organized in successive frames. At least one frame of at least one enhancement level has a duration less than the duration of at least one frame of the base level. At least one indication representative of an order used for a set of enhancement level frames corresponding to the duration of at least one frame of the base level is inserted into the data stream.


