Immersive Audio Bitstream Metadata for Dynamic Listening Poses

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in efficiently communicating audio data for immersive applications like VR and AR, requiring improved bitstream formats that balance accuracy, completeness, and dynamic rendering while reducing data rate and computational complexity.

Innovation Solution

A bitstream format is developed that includes audio data and metadata describing audio sources and environmental properties, allowing for flexible and dynamic rendering, with features like frequency grids, orientation representations, and animation indications to support varying listener positions and environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If 3D audio data is generated for each possible listening pose separately, then audio quality for each pose is improved, but data amount and processing complexity increase significantly

Engineering Contradiction:
Improveaudio qualityVSAvoiddata amount
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent segments the audio environment into discrete audio objects (sound sources) and separates spatial information into metadata parameters. Instead of storing complete 3D audio data for each pose, the system segments the audio scene into independent objects with associated metadata containing position, velocity, and acoustic environment parameters, which can be efficiently processed for any listening pose.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates a universal audio object representation that can serve multiple listening poses simultaneously. The audio objects and their metadata are designed to be pose-independent, allowing the same set of audio objects to be rendered for any listening pose by adjusting the rendering parameters, eliminating the need for separate 3D audio data for each pose.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If detailed acoustic environment data is captured for all listening poses, then spatial audio accuracy is improved, but metadata complexity and storage requirements increase

Engineering Contradiction:
Improvespatial audio accuracyVSAvoidmetadata complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies local quality by associating specific acoustic environment parameters with individual audio objects based on their spatial characteristics. Each audio object has metadata containing only the acoustic parameters relevant to its location and properties (such as reflection characteristics for objects in reflective environments), rather than storing complete environmental data for all possible positions.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses parameter changes to represent acoustic environment variations. Instead of storing complete acoustic impulse responses for each position, the system uses a limited set of acoustic parameters (reflection time, early decay time, clarity) that can be adjusted based on the listening pose and audio object position, reducing metadata complexity while maintaining spatial accuracy.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If real-time 3D audio rendering is implemented for dynamic environments, then audio realism is improved, but computational load and processing time increase

Engineering Contradiction:
Improveaudio realismVSAvoidcomputational load
Core Design Contradiction:
ReliabilityVSPower

Solution Approach 1:

The patent performs preliminary action by pre-processing audio signals into audio objects with associated metadata during encoding. The spatial parameters and acoustic environment data are prepared in advance as structured metadata, allowing the decoder to efficiently render real-time 3D audio by simply applying the pre-prepared parameters rather than performing complex spatial processing from scratch during playback.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses copying by transmitting audio objects as reusable templates that can be instantiated for different listening poses. Instead of computing complete 3D audio fields, the system copies the audio object definition and its metadata to the decoder, which then renders the appropriate spatial audio based on the target listening pose, significantly reducing computational load while maintaining realism.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP4407436B1Bitstream representing audio in an environment
Publication Date: 2026.05.13 KONINKLIJKE PHILIPS NV
  • EP4407436B1 patent drawingFigure 1~2
  • EP4407436B1 patent drawingFigure 3
  • EP4407436B1 patent drawingFigure 4

AI summary

An (encoding) apparatus comprises a metadata generator (203) generating metadata for audio data for a plurality of audio elements representing audio sources in an environment. The metadata comprises acoustic environment data for the environment where the acoustic environment data describes properties affecting sound propagation for the audio sources in the environment. At least some of the acoustic environment data is applicable to a plurality of listening poses in the environment and the properties include both static and dynamic properties. A bitstream generator (205) generates the bitstream to include the metadata and often also audio data representing the audio elements for the audio sources in the environment. A decoding apparatus may comprise a receiver for receiving the bitstream and a renderer for rendering audio for the audio environment based on the acoustic environment data and on audio data for the audio elements.