Immersive Audio Bitstream Metadata for Dynamic Listening Poses
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in efficiently communicating audio data for immersive applications like VR and AR, requiring improved bitstream formats that balance accuracy, completeness, and dynamic rendering while reducing data rate and computational complexity.
Innovation Solution
A bitstream format is developed that includes audio data and metadata describing audio sources and environmental properties, allowing for flexible and dynamic rendering, with features like frequency grids, orientation representations, and animation indications to support varying listener positions and environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If 3D audio data is generated for each possible listening pose separately, then audio quality for each pose is improved, but data amount and processing complexity increase significantly
Solution Approach 1:
The patent segments the audio environment into discrete audio objects (sound sources) and separates spatial information into metadata parameters. Instead of storing complete 3D audio data for each pose, the system segments the audio scene into independent objects with associated metadata containing position, velocity, and acoustic environment parameters, which can be efficiently processed for any listening pose.
Solution Approach 2:
The patent creates a universal audio object representation that can serve multiple listening poses simultaneously. The audio objects and their metadata are designed to be pose-independent, allowing the same set of audio objects to be rendered for any listening pose by adjusting the rendering parameters, eliminating the need for separate 3D audio data for each pose.
2Measurement precision
If detailed acoustic environment data is captured for all listening poses, then spatial audio accuracy is improved, but metadata complexity and storage requirements increase
Solution Approach 1:
The patent applies local quality by associating specific acoustic environment parameters with individual audio objects based on their spatial characteristics. Each audio object has metadata containing only the acoustic parameters relevant to its location and properties (such as reflection characteristics for objects in reflective environments), rather than storing complete environmental data for all possible positions.
Solution Approach 2:
The patent uses parameter changes to represent acoustic environment variations. Instead of storing complete acoustic impulse responses for each position, the system uses a limited set of acoustic parameters (reflection time, early decay time, clarity) that can be adjusted based on the listening pose and audio object position, reducing metadata complexity while maintaining spatial accuracy.
3Reliability
If real-time 3D audio rendering is implemented for dynamic environments, then audio realism is improved, but computational load and processing time increase
Solution Approach 1:
The patent performs preliminary action by pre-processing audio signals into audio objects with associated metadata during encoding. The spatial parameters and acoustic environment data are prepared in advance as structured metadata, allowing the decoder to efficiently render real-time 3D audio by simply applying the pre-prepared parameters rather than performing complex spatial processing from scratch during playback.
Solution Approach 2:
The patent uses copying by transmitting audio objects as reusable templates that can be instantiated for different listening poses. Instead of computing complete 3D audio fields, the system copies the audio object definition and its metadata to the decoder, which then renders the appropriate spatial audio based on the target listening pose, significantly reducing computational load while maintaining realism.
Data Source
Figure 1~2
Figure 3
Figure 4
AI summary
An (encoding) apparatus comprises a metadata generator (203) generating metadata for audio data for a plurality of audio elements representing audio sources in an environment. The metadata comprises acoustic environment data for the environment where the acoustic environment data describes properties affecting sound propagation for the audio sources in the environment. At least some of the acoustic environment data is applicable to a plurality of listening poses in the environment and the properties include both static and dynamic properties. A bitstream generator (205) generates the bitstream to include the metadata and often also audio data representing the audio elements for the audio sources in the environment. A decoding apparatus may comprise a receiver for receiving the bitstream and a renderer for rendering audio for the audio environment based on the acoustic environment data and on audio data for the audio elements.