Spatial Audio Rendering with Listener Motion Compensation Metadata

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current spatial audio rendering technologies lack flexibility in rendering audio scenes according to listener motion, particularly head orientation and position changes, which limits the complexity of motion compensation and immersive experience.

Innovation Solution

A metadata-driven approach that allows content creators to specify the preferred degrees of freedom for listener motion compensation during spatial audio rendering, using a data structure that includes settings for headphone virtualization, reverberation, head locking, and distance attenuation, enabling precise control over the rendering process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spatial audio rendering is performed with listener motion compensation, then the immersive experience and accuracy of audio scene rendering is improved, but the device complexity and computational requirements increase

Engineering Contradiction:
Improveaccuracy of audio scene renderingVSAvoidcomplexity of spatial audio renderer
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The spatial audio rendering system is segmented into distinct functional modules: metadata processing module, listener motion detection module, audio object rendering module, and head-related impulse response application module. Each module handles a specific aspect of the rendering process, allowing independent optimization and reducing overall system complexity while maintaining high rendering accuracy through specialized processing in each segment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Listener motion compensation parameters and head-related impulse responses are pre-computed and stored in metadata before playback. The system performs preliminary analysis of the audio scene and pre-processes spatial transformation data, so that during actual playback, the renderer can quickly apply pre-prepared compensation parameters without real-time complex calculations, thus improving accuracy while reducing runtime computational complexity.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If metadata is used to control rendering parameters, then the flexibility and adaptability of spatial audio rendering is improved, but the loss of information increases due to compression

Engineering Contradiction:
Improveflexibility of rendering controlVSAvoidinformation loss in metadata
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system extracts only the essential listener motion parameters (head orientation, head position, torso position) and rendering control parameters needed for spatial audio rendering, separating them from the full audio scene description. By extracting only the critical metadata elements required for motion compensation and rendering control, the system achieves high flexibility while minimizing information loss through selective data extraction rather than comprehensive data transmission.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The metadata structure uses parameter-based control where continuous physical parameters (head orientation angles, position coordinates) are transformed into discrete rendering control parameters. This parameter transformation allows the system to maintain flexibility in controlling spatial audio rendering while reducing information loss through efficient parameter encoding and compression suitable for metadata transmission.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240406660A1Spatial Audio Rendering with Listener Motion Compensation using Metadata
Publication Date: 2024.12.05 APPLE INC
  • US20240406660A1 patent drawing
  • US20240406660A1 patent drawing
  • US20240406660A1 patent drawing

AI summary

The various aspects of the disclosure here enable a content creation side to control how discrete audio objects that make up a sound program are rendered by a decoding side to achieve high spatial resolution, while giving the content creator the flexibility to decide how complex the spatial audio rendering should be in the decoding side. Metadata associated with the sound program will instruct a spatial audio renderer on how complex its listener motion compensation should be. Other aspects are also described and claimed.