Audio Scene Entity Delivery for 6DOF Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio scene delivery technologies, such as MPEG-H, do not effectively support six degrees of freedom (6DOF) translation, as they only cover three degrees of freedom, and lack information on which audio scene entities are being delivered, making it difficult to dynamically change audio representations based on user movement and position.
Innovation Solution
The introduction of a new box, 'maeS', in the MPEG-H bitstream provides information about contributing scene elements and their representations, allowing for the generation of miscibility labels and matrices to determine which audio data can be mixed together, enabling efficient delivery of audio scene entities in different representations based on user position and orientation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple audio scene entity combinations are generated and delivered to support 6DOF translation, then adaptability and rendering quality improve, but data transmission volume and bandwidth requirements increase
Solution Approach 1:
The audio scene is segmented into multiple entity combinations, each representing a different spatial configuration suitable for specific user positions and orientations. Instead of transmitting all possible audio data, the system divides the complete audio scene into discrete entities that can be independently selected and rendered based on user movement, thereby reducing overall data transmission while maintaining 6DOF capability
Solution Approach 2:
The system dynamically selects and switches between different audio scene entity combinations based on real-time user position and orientation data. This dynamic adaptation allows the client to receive only the relevant audio entities needed for the current viewing angle, rather than continuously transmitting all audio data, thus optimizing bandwidth utilization while supporting seamless 6DOF translation
2Adaptability or versatility
If audio scene entities are delivered in multiple representations (objects, channels, HOA), then client flexibility and rendering options improve, but processing complexity and device requirements increase
Solution Approach 1:
The server generates multiple representations of audio scene entities (objects, channels, HOA) but the client selectively processes only the representation most suitable for its capabilities and the current rendering context. This partial processing approach allows the system to provide full multi-format availability while avoiding the complexity of simultaneously processing all representations, thus balancing client flexibility with device complexity
3Manufacturing precision
If audio representations are dynamically changed based on user position and orientation, then audio experience quality improves, but computational overhead and processing time increase
Solution Approach 1:
The server pre-generates and delivers audio scene entities in multiple representations and configurations before the client needs them. By preparing the audio entities in advance with various spatial configurations already computed, the system reduces the real-time processing burden on the client, allowing for high-quality dynamic adaptation based on user position and orientation without excessive computational overhead during actual rendering
Data Source
AI summary
In an example embodiment, method, apparatus, and computer program product are provided. The apparatus includes at least one processor; and at least one memory including computer program code; the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus at least to perform: assign one or more audio representations to one or more audio scene entities in an audio scene; generate one or more audio scene entity combinations based on the one or more audio scene entities and the one or more audio representations; and signal the one or more audio scene entity combinations to a client, wherein the one or more audio representations assigned to the one or more audio scene entities cause the client to select an appropriate audio scene entity combination from the one or more audio scene entity combinations to render the audio scene.


