Layered Ambisonic Audio Coding for Spatial Resolution and Bandwidth
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional channel-based audio systems lack flexibility in rendering and manipulating sound sources, leading to issues with spatial resolution and bandwidth utilization when trying to layer audio components, especially in object-based audio scenarios.
Innovation Solution
A hybrid audio processing technique that combines Ambisonic audio components with object-based audio, allowing for layered coding where a base layer includes lower-order Ambisonics and additional layers can include higher-order Ambisonics or object-based audio signals, enabling flexible rendering and adaptation to device capabilities and user preferences.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If channel-based audio is used, then spatial coverage is provided, but flexibility in rendering and manipulating sound sources is limited
Solution Approach 1:
The audio signal is segmented into multiple independent layers: a base layer containing lower-order Ambisonic components and enhancement layers containing higher-order Ambisonic components and object-based audio. This segmentation allows selective processing and rendering of different sound sources, providing flexibility while managing complexity through hierarchical organization.
Solution Approach 2:
The patent transitions from traditional 2D channel-based audio to 3D spatial audio by incorporating higher-order Ambisonic components and object-based metadata that encode three-dimensional position information. This dimensional expansion enables sophisticated spatial manipulation and rendering of sound sources in a three-dimensional soundscape.
2Measurement precision
If higher-order Ambisonics are used, then spatial resolution is improved, but bandwidth requirement increases
Solution Approach 1:
The Ambisonic signal is divided into multiple orders with lower-order components placed in the base layer and higher-order components in enhancement layers. This segmentation allows progressive transmission of spatial resolution information, enabling receivers to achieve high spatial resolution when bandwidth permits while maintaining efficient bandwidth usage by transmitting only essential lower-order components when resources are constrained.
Solution Approach 2:
The system dynamically adapts the transmission and rendering of Ambisonic orders based on available bandwidth and receiver capabilities. Enhancement layers containing higher-order components are conditionally transmitted and processed, allowing the system to optimize between spatial resolution and bandwidth consumption in real-time according to network conditions and device capabilities.
3Adaptability or versatility
If object-based audio is used, then independent manipulation of sound sources is enabled, but rendering complexity increases
Solution Approach 1:
Object-based audio signals are segmented into separate enhancement layers with associated metadata, distinct from the base Ambisonic layer. This segmentation isolates the complexity of object processing to specific layers, allowing the base layer to be rendered independently using simpler Ambisonic decoding while object-based manipulation capabilities are added progressively through enhancement layers when computational resources permit.
Solution Approach 2:
The patent introduces an intermediary layered structure that mediates between simple Ambisonic rendering and complex object-based audio processing. The base layer provides a simplified rendering path for devices with limited capabilities, while enhancement layers serve as intermediaries that add object-based manipulation capabilities progressively, allowing receivers to choose appropriate processing complexity based on their capabilities.
4Adaptability or versatility
If a hybrid layered approach is used, then flexibility and spatial resolution are enhanced, but processing complexity increases
Solution Approach 1:
The hybrid audio signal is segmented into a hierarchical structure with a base layer containing lower-order Ambisonic components and enhancement layers containing higher-order Ambisonic components and object-based audio. This segmentation organizes processing complexity into manageable segments that can be independently processed and combined, allowing receivers to process only the layers appropriate to their capabilities.
Solution Approach 2:
The system dynamically adapts the processing complexity based on receiver capabilities and network conditions. Receivers can selectively process only the base layer for simple playback, or progressively process enhancement layers when higher computational resources are available, allowing the processing complexity to scale dynamically with available resources while maintaining flexibility and spatial resolution.
Data Source
AI summary
A first layer of data having a first set of Ambisonic audio components can be decoded where the first set of Ambisonic audio components is generated based on ambience and one or more object-based audio signals. A second layer of data is decoded having at least one of the one or more object-based audio signals. One of the object-based audio signals is subtracted from the first set of Ambisonic audio components. The resulting Ambisonic audio components are rendered to generate a first set of audio channels. The one or more object-based audio signals are spatially rendered to generate a second set of audio channels. Other aspects are described and claimed.


