Audio Decoder Scene Packet Segmentation for VR Spatial Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current VR and AR technologies face challenges in providing an immersive audio experience with high definition and efficient data transmission and decoding, particularly in achieving six degrees of freedom audio while maintaining feasible bandwidths.
Innovation Solution
The use of three distinct packet types - scene configuration packets, scene update packets, and scene payload packets - within a bitstream, conforming to MPEG-H MHAS packet definitions, allows for efficient transmission and rendering of spatial audio by defining renderer configurations, updating metadata, and providing necessary scene information.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high definition spatial audio rendering is implemented for immersive VR/AR experience, then audio quality and immersion are improved, but data transmission bandwidth and decoding complexity increase
Solution Approach 1:
The audio scene is segmented into multiple packets of different types (scene configuration packets, scene update packets, scene payload packets). Each packet type carries specific renderer configuration information for different aspects of the audio scene, allowing the decoder to process and render audio in a structured, modular manner that reduces overall decoding complexity while maintaining high definition quality
Solution Approach 2:
The renderer configuration information dynamically adapts to temporal evolution of the audio scene through timestamp information. The system evaluates timestamps and adjusts rendering configurations in real-time, enabling high definition spatial audio rendering that responds to scene changes without requiring complete re-decoding of entire audio sequences
2Measurement precision
If complete renderer configuration information is provided for temporal evolution of audio scenes, then audio rendering accuracy is improved, but data transmission volume increases
Solution Approach 1:
The patent extracts and separates renderer configuration information into distinct packet types that can be independently transmitted and processed. Scene configuration packets contain essential rendering parameters, while scene update packets carry only the changes needed, reducing redundant data transmission while maintaining complete rendering accuracy
Solution Approach 2:
Scene configuration packets are transmitted in advance to provide renderer configuration information before the actual audio rendering is needed. This preliminary provision of configuration data allows the decoder to be pre-configured for accurate rendering without requiring all configuration details to be present simultaneously, optimizing data transmission efficiency
3Productivity
If real-time audio scene rendering is implemented, then user experience responsiveness is improved, but processing latency increases
Solution Approach 1:
The system uses periodic scene update packets with timestamp information to refresh renderer configurations at regular intervals. This periodic updating mechanism enables real-time audio scene rendering by synchronizing configuration updates with the audio playback timeline, maintaining responsiveness while managing processing latency through predictable, periodic processing cycles
Data Source
AI summary
Embodiments create an audio decoder which spatially renders one or more audio signals. The audio decoder receives a plurality of packets of different packet types, comprising one or more scene configuration packets providing a renderer configuration information defining a usage of scene objects and/or a usage of scene characteristics, and comprising one or more scene update packets defining a update of scene metadata for the rendering, and comprising one or more scene payload packets comprising definitions of one or more of the scene objects and/or definitions of one or more of the scene characteristics. The audio decoder selects definitions of one or more scene objects and/or definitions of one or more scene characteristics, which are in included in the scene payload packets, for the rendering in dependence on the renderer configuration information. The audio decoder updates one or more scene metadata in dependence on a content of the one or more scene update packets. Further embodiments are related to encoders, methods and bitstreams.Further embodiments create decoders, encoders, methods and bitstreams with scene update packets with update conditions, with scene configuration packets providing a renderer configuration information defining a temporal evolution of a rendering scenario and with a timestamp information and/or with subscene cell information, wherein the cell information defines an association between the one or more cells and respective one or more data structures.


