Adaptive Object Audio Selection via Metadata for MPEG-DASH
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
MPEG-DASH standards do not allow for the selection of object audio data based on display states, leading to inefficient audio data transmission and rendering, as audio data is transmitted regardless of the display state, resulting in extra transmission bands and rendering processing.
Innovation Solution
An information processing apparatus and method that generates a content file with object selection information stored in a metadata file, enabling the selection of object audio data based on the display state, allowing for optimal audio data transmission and rendering by storing necessary audio data for each display state and signaling object selection information accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If object audio data is transmitted for all display states, then complete audio coverage is ensured, but transmission bandwidth increases and rendering processing becomes more complex
Solution Approach 1:
The patent extracts and transmits only the necessary object audio data based on the current display state. The server determines which object audio data corresponds to the displayed image region and transmits only that subset, rather than all object audio data. This extraction principle reduces transmission bandwidth while ensuring complete audio coverage for the visible content.
Solution Approach 2:
The patent applies local quality by making the audio transmission adaptive to the specific display state. Different display states (e.g., different zoom levels, pan positions, or visible regions) receive different sets of object audio data tailored to what is currently visible. This ensures audio quality matches the local display context without unnecessary transmissions.
2Reliability
If object audio data is transmitted for all display states, then complete audio coverage is ensured, but rendering processing complexity increases
Solution Approach 1:
The patent extracts only the relevant object audio data needed for the current display state on the client side. By receiving a subset of object audio data rather than all available data, the rendering processing complexity is reduced while maintaining complete audio coverage for what is actually displayed.
3Measurement precision
If object selection information is stored in metadata file, then audio data selection accuracy improves, but content file structure complexity increases
Solution Approach 1:
The patent introduces a metadata file as an intermediary between the content file and the object audio data. The metadata file contains object selection information that precisely identifies which object audio data corresponds to the displayed image region. This intermediary structure enables accurate audio data selection while maintaining a clear separation between content delivery and audio rendering logic.
Data Source
AI summary
There are provided an information processing apparatus, an information processing method, and a program. The information processing apparatus includes a generating unit configured to generate a content file including object selection information for selecting object audio data in accordance with a display state of an image, in which the generating unit stores the object selection information in a metadata file included in the content file.


