Audio Object Metadata for Immersive Playback Adaptation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional multi-channel audio systems require professional rendering and are limited to specific playback settings, leading to degraded performance when played on different systems due to mismatched settings, as audio object properties like position and size cannot be manipulated once created.
Innovation Solution
A method and system for generating metadata associated with audio objects, including estimated trajectories and perceptual sizes, allowing for precise playback across various systems and enabling post-processing manipulation, using techniques like energy-weighted, correspondence, and hybrid approaches for position estimation and ICC-based perceptual size calculation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If audio content is created with a multi-channel format optimized for a specific playback setting, then the listening experience is optimized for that setting, but the performance degrades when played on different playback settings due to mismatch between settings
Solution Approach 1:
The audio content is segmented into individual audio objects with independently manipulatable properties (position, size, trajectory). This segmentation allows each audio object to be processed and adapted separately for different playback configurations, resolving the contradiction between optimized quality for a specific setting and compatibility across multiple settings.
Solution Approach 2:
The audio object properties are made dynamic and manipulatable after creation. The system allows modification of audio object attributes such as position, size, and trajectory metadata, enabling the same audio content to be adaptively rendered for different playback settings (e.g., 5.1, 7.1, stereo) without degrading the listening experience.
2Manufacturing precision
If audio object properties are fixed during creation, then the rendering is precise for the intended format, but the properties cannot be manipulated once created for different playback systems
Solution Approach 1:
The system performs preliminary actions by embedding rich metadata (position, size, trajectory) during audio object creation that preserves all necessary information for future manipulation. This preliminary encoding of properties allows precise initial rendering while simultaneously enabling later adaptation to different playback configurations without loss of accuracy.
Solution Approach 2:
The invention enables parameter changes of audio objects after creation by storing manipulatable properties in metadata. Users can modify parameters such as position coordinates, size dimensions, and trajectory paths, allowing the same audio content to be precisely rendered for different playback formats (5.1, 7.1, stereo, headphones) while maintaining manufacturing precision through controlled parameter adjustment.
3Reliability
If professional rendering is required in a studio using panning tools, then the audio quality is optimized, but the process becomes complex and cannot be tailored for different formats flexibly
Solution Approach 1:
The invention creates a digital copy of audio object properties in the form of metadata that can be independently manipulated without affecting the original audio content. This copying approach allows professional-quality rendering to be replicated and adapted for different formats through metadata manipulation rather than requiring complex re-rendering processes, reducing device complexity while maintaining quality.
Data Source
AI summary
Example embodiments disclosed herein relate to audio object processing. A method for processing audio content, which includes at least one audio object of a multi-channel format, is disclosed. The method includes generating metadata associated with the audio object, the metadata including at least one of an estimated trajectory of the audio object and an estimated perceptual size of the audio object, the perceptual size being a perceived area of a phantom of the audio object produced by at least two transducers. Corresponding system and computer program product are also disclosed.


