Object-Based Audio Rendering With Interactive User Mixing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio technologies lack the ability to provide immersive, personalizable audio experiences that allow users to interactively select and mix audio content, particularly in object-based audio programs, without compromising on spatial accuracy and compatibility with legacy systems.
Innovation Solution
The implementation of object-based audio programs that include metadata for interactive rendering, allowing users to select and mix audio content using a controller, while ensuring compatibility with legacy decoders and systems through the use of durable and non-durable metadata, and employing distributed rendering across multiple devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If object-based audio programs with interactive rendering metadata are implemented, then user control and personalization capability are improved, but device complexity and system compatibility requirements worsen
Solution Approach 1:
The audio program is segmented into multiple independent audio objects, each with its own metadata properties. This allows users to selectively control individual objects while the system processes them independently, reducing overall system complexity through modular handling of audio elements.
Solution Approach 2:
A controller device acts as an intermediary between the user and the audio rendering system. The controller receives user input, processes it according to the audio program's metadata, and generates appropriate control signals, thereby shielding the complex rendering system from direct user interaction while enabling sophisticated user control.
2Adaptability or versatility
If distributed rendering across multiple devices is employed, then audio rendering flexibility and user experience are improved, but synchronization complexity and system coordination requirements worsen
Solution Approach 1:
The system implements feedback mechanisms where each rendering device reports its operational status and output to the controller. The controller uses this feedback to dynamically adjust rendering parameters and maintain synchronization across distributed devices, enabling flexible multi-device rendering without excessive coordination complexity.
Solution Approach 2:
Audio objects are pre-processed and tagged with rendering metadata during program creation, including spatial information and device capability requirements. This preliminary action allows distributed devices to independently process their assigned objects without requiring complex real-time coordination, as the necessary rendering parameters are predetermined.
3Adaptability or versatility
If durable and non-durable metadata are used to ensure legacy compatibility, then backward compatibility is improved, but information processing complexity worsens
Solution Approach 1:
Different metadata properties are assigned different durability characteristics based on their specific function. Critical spatial and identification metadata are marked as durable for legacy compatibility, while user-preference and dynamic control metadata are marked as non-durable. This local differentiation allows the system to process only relevant metadata for each rendering context, reducing overall processing complexity.
Data Source
AI summary
Methods for generating an object based audio program which is renderable in a personalizable manner, e.g., to provide an immersive, perception of audio content of the program. Other embodiments include steps of delivering (e.g., broadcasting), decoding, and/or rendering such a program. Rendering of audio objects indicated by the program may provide an immersive experience. The audio content of the program may be indicative of multiple object channels (e.g., object channels indicative of user-selectable and user-configurable objects, and typically also a default set of objects which will be rendered in the absence of a selection by a user) and a bed of speaker channels. Another aspect is an audio processing unit (e.g., encoder or decoder) configured to perform, or which includes a buffer memory which stores at least one frame (or other segment) of an object based audio program (or bitstream thereof) generated in accordance with, any embodiment of the method.


