Audio Bitstream Metadata for Dynamic Acoustic Environments
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies face challenges in efficiently communicating audio data for immersive applications like VR and AR, requiring improved bitstream formats that balance data rate, complexity, and flexibility for dynamic environments and changing listening positions.
Innovation Solution
A bitstream format that includes metadata for audio sources and acoustic environment properties, allowing for static and dynamic properties applicable to multiple listening poses, with optimized data representation and flexible rendering capabilities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If individualized audio models with separate audio sources and metadata are used, then audio quality and flexibility are improved, but data rate increases
Solution Approach 1:
The audio environment is segmented into discrete audio objects, each with its own metadata properties. This allows selective transmission of only necessary data for each object rather than transmitting complete audio environments, reducing overall data rate while maintaining flexibility.
Solution Approach 2:
The patent transmits partial audio environment data by selecting and transmitting only the most relevant properties for each audio object based on current rendering needs, rather than transmitting all possible properties. This reduces data rate while maintaining sufficient flexibility for quality rendering.
2Reliability
If comprehensive acoustic environment data is transmitted, then audio quality is improved, but device complexity increases
Solution Approach 1:
The patent applies different levels of data detail to different audio objects based on their importance and characteristics. Critical audio objects receive more detailed metadata while less critical objects use simplified representations, maintaining audio quality for important elements while reducing overall complexity.
Solution Approach 2:
The patent dynamically adjusts which metadata properties are transmitted based on current rendering requirements, listener position, and environmental factors. This selective parameter transmission maintains audio quality when needed while reducing complexity during stable or less critical scenarios.
3Adaptability or versatility
If dynamic properties are included for multiple listening poses, then adaptability is improved, but data rate increases
Solution Approach 1:
The patent updates dynamic properties at optimized intervals rather than continuously for all listening poses. By determining when updates are actually needed based on listener movement and environmental changes, the system maintains adaptability while reducing redundant data transmission.
Solution Approach 2:
The patent creates a unified metadata structure that serves multiple listening poses simultaneously through parameter sets that can be applied across different listener positions. This universal representation maintains adaptability while avoiding the need to transmit separate complete datasets for each pose.
Data Source
AI summary
An (encoding) apparatus comprises a metadata generator (203) generating metadata for audio data for a plurality of audio elements representing audio sources in an environment. The metadata comprises acoustic environment data for the environment where the acoustic environment data describes properties affecting sound propagation for the audio sources in the environment. At least some of the acoustic environment data is applicable to a plurality of listening poses in the environment and the properties include both static and dynamic properties. A bitstream generator (205) generates the bitstream to include the metadata and often also audio data representing the audio elements for the audio sources in the environment. A decoding apparatus may comprise a receiver for receiving the bitstream and a renderer for rendering audio for the audio environment based on the acoustic environment data and on audio data for the audio elements.


