Layered Compressed Sound Coding for Variable Transmission Conditions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing layered coding schemes are inadequate for special types of compressed sound or sound field representations, such as Higher-Order Ambisonics (HOA), particularly in environments with time-varying transmission conditions, leading to signal dropouts and inefficiencies in error protection and bandwidth usage.
Innovation Solution
A method of layered encoding that subdivides components into hierarchical layers, assigns basic and enhancement side information to these layers, ensuring each layer contains sufficient information for reconstructing a sound representation, even if higher layers are not received, while optimizing bandwidth usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If layered coding is applied to compressed sound representations, then adaptability to transmission conditions is improved, but device complexity increases due to the need for hierarchical layer management and component grouping
Solution Approach 1:
The compressed sound representation is segmented into multiple hierarchical layers (base layer and enhancement layers), each containing specific components. This segmentation enables independent transmission and decoding of layers, allowing the system to adapt to varying transmission conditions by selectively decoding layers based on available resources and channel quality, thereby resolving the contradiction between adaptability and complexity through structured organization.
Solution Approach 2:
The patent introduces a hierarchical dimension to the traditional flat structure of audio coding. By organizing components across multiple layers with different priorities and quality levels, the system gains an additional dimension for controlling and adapting the decoding process, enabling flexible adaptation to transmission conditions without proportionally increasing operational complexity.
2Reliability
If error protection is increased for the base layer, then reliability is improved, but bandwidth consumption increases
Solution Approach 1:
Different error protection levels are applied locally to different layers based on their importance and size. The base layer, being critical for basic sound representation, receives high error protection. Enhancement layers receive progressively less protection as their contribution to overall quality diminishes. This local differentiation of protection levels optimizes the bandwidth-reliability tradeoff by concentrating resources where they provide maximum benefit.
Solution Approach 2:
Instead of applying full error protection uniformly across all layers (excessive action), the patent applies partial error protection tailored to each layer's requirements. The base layer receives sufficient protection to ensure reliable decoding, while enhancement layers receive only the necessary protection level, avoiding waste of bandwidth on unnecessary protection for lower-priority data.
3Manufacturing precision
If enhancement layers are added to improve quality, then sound representation quality is improved, but device complexity increases due to additional layer management requirements
Solution Approach 1:
The decoding system is designed to be dynamic, allowing the decoder to adaptively select which layers to decode based on transmission conditions and available bandwidth. Rather than requiring complex fixed management of all possible layers, the system dynamically adjusts its behavior, simplifying the effective complexity at any given moment while maintaining the capability for high-quality reconstruction when conditions permit.
Data Source
AI summary
The present document relates to a method of layered encoding of a compressed sound representation of a sound or sound field. The compressed sound representation comprises a basic compressed sound representation comprising a plurality of components, basic side information for decoding the basic compressed sound representation to a basic reconstructed sound representation of the sound or sound field, and enhancement side information including parameters for improving the basic reconstructed sound representation. The method comprises sub-dividing the plurality of components into a plurality of groups of components and assigning each of the plurality of groups to a respective one of a plurality of hierarchical layers, the number of groups corresponding to the number of layers, and the plurality of layers including a base layer and one or more hierarchical enhancement layers, adding the basic side information to the base layer, and determining a plurality of portions of enhancement side information from the enhancement side information and assigning each of the plurality of portions of enhancement side information to a respective one of the plurality of layers, wherein each portion of enhancement side information includes parameters for improving a reconstructed sound representation obtainable from data included in the respective layer and any layers lower than the respective layer. The document further relates to a method of decoding a compressed sound representation of a sound or sound field, wherein the compressed sound representation is encoded in a plurality of hierarchical layers that include a base layer and one or more hierarchical enhancement layers, as well as to an encoder and a decoder for layered coding of a compressed sound representation.


