Spatial Audio Metadata for Legacy-Compatible 6DoF Rendering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio recording and rendering technologies struggle to provide high-quality spatial audio experiences that adapt to user movement, particularly in environments where legacy devices are used, lacking the ability to maintain spatial accuracy and adjust audio rendering accordingly.

Innovation Solution

A method involving the use of spatial metadata, including distance, direction, and energy ratio parameters, is associated with audio signals to enable rendering with or without metadata, allowing for improved spatial audio rendering on both legacy and advanced devices, accommodating user movement in six degrees of freedom.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If spatial metadata is used to improve spatial audio rendering quality and user movement adaptation, then spatial accuracy and adaptability are improved, but device complexity and processing requirements increase

Engineering Contradiction:
Improvespatial accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The spatial audio signal is segmented into multiple components (direct sound, reflected sound, diffuse sound) based on spatial metadata parameters. Each component is processed separately with appropriate spatial transformations applied, allowing complex spatial rendering to be broken down into manageable segments that can be handled by different processing stages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Spatial metadata including distance parameters, direction parameters, and energy ratio parameters is extracted and prepared in advance during the recording or encoding stage. This preliminary processing organizes spatial information before rendering, reducing the computational burden during real-time playback and allowing legacy devices to use pre-computed spatial cues without complex real-time processing.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If spatial metadata is incorporated to enable six degrees of freedom movement adaptation, then adaptability to user movement is improved, but loss of information and data requirements increase

Engineering Contradiction:
Improvemovement adaptationVSAvoiddata overhead
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The spatial metadata structure is designed to be universal and multi-functional, serving multiple purposes: it enables six degrees of freedom movement adaptation, provides spatial accuracy for static rendering, and maintains compatibility with legacy systems that ignore the metadata. The same metadata parameters (distance, direction, energy ratio) support both advanced spatial audio rendering and traditional audio playback.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent creates a simplified copy of spatial information in the form of metadata parameters that can be used by different rendering systems. Instead of requiring full spatial audio field data, essential spatial characteristics are copied into compact metadata structures that can be efficiently stored and transmitted, reducing data overhead while maintaining adaptability.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If dual rendering contexts are supported (with and without spatial metadata), then compatibility with legacy devices is improved, but device complexity increases

Engineering Contradiction:
Improvedevice compatibilityVSAvoidrendering system complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The rendering system is designed to be dynamic and adaptive, automatically selecting between two rendering contexts based on device capabilities. Advanced devices detect the presence of spatial metadata and activate the enhanced rendering path that utilizes distance parameters, direction parameters, and energy ratio parameters for six degrees of freedom movement adaptation. Legacy devices automatically fall back to traditional rendering without spatial metadata processing. This dynamic adaptation allows a single system to serve multiple device types without requiring complex configuration or manual intervention.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3688753B1Recording and rendering spatial audio signals
Publication Date: 2025.11.05 NOKIA TECHNOLOGIES OY
  • EP3688753B1 patent drawingFigure 1
  • EP3688753B1 patent drawingFigure 2
  • EP3688753B1 patent drawingFigure 3~4

AI summary

Examples of the disclosure relate to a method, apparatus and computer program, the method comprising:obtaining audio signals wherein the audio signals represent spatial sound and can be used to render spatial audio using linear methods;obtaining spatial metadata corresponding to the spatial sound represented by the audio signals; and associating the spatial metadata with the obtained audio signals so that in a first rendering context the obtained audio signals can be rendered without using the spatial metadata and in a second rendering context the obtained audio signals can be rendered with using the spatial metadata.