Spatial Audio Metadata for Legacy-Compatible 6DoF Rendering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio recording and rendering technologies struggle to provide high-quality spatial audio experiences that adapt to user movement, particularly in environments where legacy devices are used, lacking the ability to maintain spatial accuracy and adjust audio rendering accordingly.
Innovation Solution
A method involving the use of spatial metadata, including distance, direction, and energy ratio parameters, is associated with audio signals to enable rendering with or without metadata, allowing for improved spatial audio rendering on both legacy and advanced devices, accommodating user movement in six degrees of freedom.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If spatial metadata is used to improve spatial audio rendering quality and user movement adaptation, then spatial accuracy and adaptability are improved, but device complexity and processing requirements increase
Solution Approach 1:
The spatial audio signal is segmented into multiple components (direct sound, reflected sound, diffuse sound) based on spatial metadata parameters. Each component is processed separately with appropriate spatial transformations applied, allowing complex spatial rendering to be broken down into manageable segments that can be handled by different processing stages.
Solution Approach 2:
Spatial metadata including distance parameters, direction parameters, and energy ratio parameters is extracted and prepared in advance during the recording or encoding stage. This preliminary processing organizes spatial information before rendering, reducing the computational burden during real-time playback and allowing legacy devices to use pre-computed spatial cues without complex real-time processing.
2Adaptability or versatility
If spatial metadata is incorporated to enable six degrees of freedom movement adaptation, then adaptability to user movement is improved, but loss of information and data requirements increase
Solution Approach 1:
The spatial metadata structure is designed to be universal and multi-functional, serving multiple purposes: it enables six degrees of freedom movement adaptation, provides spatial accuracy for static rendering, and maintains compatibility with legacy systems that ignore the metadata. The same metadata parameters (distance, direction, energy ratio) support both advanced spatial audio rendering and traditional audio playback.
Solution Approach 2:
The patent creates a simplified copy of spatial information in the form of metadata parameters that can be used by different rendering systems. Instead of requiring full spatial audio field data, essential spatial characteristics are copied into compact metadata structures that can be efficiently stored and transmitted, reducing data overhead while maintaining adaptability.
3Adaptability or versatility
If dual rendering contexts are supported (with and without spatial metadata), then compatibility with legacy devices is improved, but device complexity increases
Solution Approach 1:
The rendering system is designed to be dynamic and adaptive, automatically selecting between two rendering contexts based on device capabilities. Advanced devices detect the presence of spatial metadata and activate the enhanced rendering path that utilizes distance parameters, direction parameters, and energy ratio parameters for six degrees of freedom movement adaptation. Legacy devices automatically fall back to traditional rendering without spatial metadata processing. This dynamic adaptation allows a single system to serve multiple device types without requiring complex configuration or manual intervention.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
Examples of the disclosure relate to a method, apparatus and computer program, the method comprising:obtaining audio signals wherein the audio signals represent spatial sound and can be used to render spatial audio using linear methods;obtaining spatial metadata corresponding to the spatial sound represented by the audio signals; and associating the spatial metadata with the obtained audio signals so that in a first rendering context the obtained audio signals can be rendered without using the spatial metadata and in a second rendering context the obtained audio signals can be rendered with using the spatial metadata.