Object-Based Audio Rendering with Legacy-Compatible Base Layers
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio decoding and rendering systems struggle to provide a full range audio experience for object-based audio programs, especially when legacy systems are unable to parse object channels and related metadata, limiting the flexibility and personalization of audio content rendering.
Innovation Solution
The invention generates an object-based audio program with a base layer compatible with legacy systems for a default mix and an extension layer for personalized rendering, using a base layer of speaker channels and an extension layer of selectable object channels, along with metadata for alternative mixes, allowing both legacy and non-legacy systems to deliver a full range audio experience.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If object-based audio programs use extension layers with selectable object channels for personalized rendering, then adaptability and personalization are improved, but device complexity increases because legacy systems cannot parse object channels and metadata
Solution Approach 1:
The audio program is divided into a base layer (compatible with legacy systems) and an extension layer (enabling personalized rendering). The base layer contains speaker channels that legacy systems can process, while the extension layer contains object channels and metadata for advanced systems. This segmentation allows the same audio program to serve both legacy and modern systems without requiring complex processing in legacy devices.
Solution Approach 2:
Metadata acts as an intermediary between the object channels and the rendering system. The metadata contains rendering parameters that guide advanced systems on how to process and mix the object channels. This intermediary structure enables personalized rendering without requiring legacy systems to understand or process the object-based audio data, thus managing complexity while enabling adaptability.
2Device complexity
If legacy systems render only the base layer without metadata processing, then device complexity is reduced, but adaptability and personalization capability are limited
Solution Approach 1:
The audio program is divided into a base layer (compatible with legacy systems) and an extension layer (enabling personalized rendering). The base layer contains speaker channels that legacy systems can process, while the extension layer contains object channels and metadata for advanced systems. This segmentation allows the same audio program to serve both legacy and modern systems without requiring complex processing in legacy devices.
3Adaptability or versatility
If both base layer and extension layer are processed to provide full range audio experience, then adaptability is improved, but ease of operation worsens due to increased processing requirements
Solution Approach 1:
The audio program is divided into a base layer (compatible with legacy systems) and an extension layer (enabling personalized rendering). The base layer contains speaker channels that legacy systems can process, while the extension layer contains object channels and metadata for advanced systems. This segmentation allows the same audio program to serve both legacy and modern systems without requiring complex processing in legacy devices.
Solution Approach 2:
Legacy systems perform partial processing by rendering only the base layer, which is sufficient for basic audio playback. Advanced systems perform excessive action by processing both the base layer and the extension layer with metadata to provide personalized rendering. This approach allows systems to process only what they are capable of, improving ease of operation while maintaining adaptability.
Data Source
Figure 1~4
Figure 5~7
Figure 6
AI summary
Methods for generating an object based audio program, renderable in a personalizable manner, and including a bed of speaker channels renderable in the absence of selection of other program content (e.g., to provide a default full range audio experience). Other embodiments include steps of delivering, decoding, and/or rendering such a program. Rendering of content of the bed, or of a selected mix of other content of the program, may provide an immersive experience. The program may include multiple object channels (e.g., object channels indicative of user-selectable and user-configurable objects), the bed of speaker channels, and other speaker channels. Another aspect is an audio processing unit (e.g., encoder or decoder) configured to perform, or which includes a buffer memory which stores at least one frame (or other segment) of an object based audio program (or bitstream thereof) generated in accordance with, any embodiment of the method.