Audio Object Metadata for Immersive Playback Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional multi-channel audio systems require professional rendering and are limited to specific playback settings, leading to degraded performance when played on different systems due to mismatched settings, as audio object properties like position and size cannot be manipulated once created.

Innovation Solution

A method and system for generating metadata associated with audio objects, including estimated trajectories and perceptual sizes, allowing for precise playback across various systems and enabling post-processing manipulation, using techniques like energy-weighted, correspondence, and hybrid approaches for position estimation and ICC-based perceptual size calculation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio content is created with a multi-channel format optimized for a specific playback setting, then the listening experience is optimized for that setting, but the performance degrades when played on different playback settings due to mismatch between settings

Engineering Contradiction:
Improvelistening experience qualityVSAvoidplayback setting compatibility
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The audio content is segmented into individual audio objects with independently manipulatable properties (position, size, trajectory). This segmentation allows each audio object to be processed and adapted separately for different playback configurations, resolving the contradiction between optimized quality for a specific setting and compatibility across multiple settings.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The audio object properties are made dynamic and manipulatable after creation. The system allows modification of audio object attributes such as position, size, and trajectory metadata, enabling the same audio content to be adaptively rendered for different playback settings (e.g., 5.1, 7.1, stereo) without degrading the listening experience.

Inventive Principle:
Principle #15Dynamics

2Manufacturing precision

If audio object properties are fixed during creation, then the rendering is precise for the intended format, but the properties cannot be manipulated once created for different playback systems

Engineering Contradiction:
Improveaudio object positioning accuracyVSAvoidpost-creation property manipulation
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary actions by embedding rich metadata (position, size, trajectory) during audio object creation that preserves all necessary information for future manipulation. This preliminary encoding of properties allows precise initial rendering while simultaneously enabling later adaptation to different playback configurations without loss of accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The invention enables parameter changes of audio objects after creation by storing manipulatable properties in metadata. Users can modify parameters such as position coordinates, size dimensions, and trajectory paths, allowing the same audio content to be precisely rendered for different playback formats (5.1, 7.1, stereo, headphones) while maintaining manufacturing precision through controlled parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If professional rendering is required in a studio using panning tools, then the audio quality is optimized, but the process becomes complex and cannot be tailored for different formats flexibly

Engineering Contradiction:
Improveaudio rendering qualityVSAvoidrendering process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The invention creates a digital copy of audio object properties in the form of metadata that can be independently manipulated without affecting the original audio content. This copying approach allows professional-quality rendering to be replicated and adapted for different formats through metadata manipulation rather than requiring complex re-rendering processes, reducing device complexity while maintaining quality.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10362427B2Generating metadata for audio object
Publication Date: 2019.07.23 DOLBY LABORATORIES LICENSING CORP
  • US10362427B2 patent drawing
  • US10362427B2 patent drawing
  • US10362427B2 patent drawing

AI summary

Example embodiments disclosed herein relate to audio object processing. A method for processing audio content, which includes at least one audio object of a multi-channel format, is disclosed. The method includes generating metadata associated with the audio object, the metadata including at least one of an estimated trajectory of the audio object and an estimated perceptual size of the audio object, the perceptual size being a perceived area of a phantom of the audio object produced by at least two transducers. Corresponding system and computer program product are also disclosed.