Depth-Extended DirAC Sound Field Description for 6DoF Reproduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sound field representations, such as Ambisonics and DirAC, do not provide sufficient information for translational shifts in six-degrees-of-freedom (6DoF) applications, as they lack information about object distances and absolute positions, making it impossible to compute sound field descriptions at different listener positions.
Innovation Solution
Generating an enhanced sound field description that includes meta data relating to spatial information, such as depth maps, to facilitate the calculation of modified sound fields at different reference locations, allowing for 6DoF reproduction by integrating distance information into the parametric representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional Ambisonics or DirAC sound field representations are used, then the sound field can be represented with limited information, but translational shifts in 6DoF applications become impossible
Solution Approach 1:
The patent extends the traditional 2D sound field representation by adding a depth dimension. Depth maps are introduced as additional data that specify the distance of sound sources from the reference location, transforming the representation from a 2D angular distribution to a 3D spatial distribution, thereby enabling 6DoF reproduction.
Solution Approach 2:
The sound field representation is segmented into multiple components: traditional Ambisonics channels for angular information and separate depth maps for distance information. This segmentation allows each component to be processed independently and combined to achieve the full 6DoF effect.
2Adaptability or versatility
If depth information is added to the sound field representation, then 6DoF reproduction becomes possible, but the data structure becomes more complex
Solution Approach 1:
Instead of fundamentally redesigning the data structure, the patent adds depth information as an additional dimension. Depth maps are appended to the existing Ambisonics data, creating an extended representation that maintains compatibility with traditional systems while enabling new capabilities.
Solution Approach 2:
The patent introduces intermediate processing steps that convert the extended sound field description with depth maps into the desired 6DoF representation. These intermediaries handle the complexity of integrating multiple data types and transforming them into the appropriate format for different reproduction scenarios.
3Measurement precision
If traditional sound field representations are used, then processing remains simple, but accurate mapping of changed listener perspective is impossible
Solution Approach 1:
The patent performs preliminary analysis of the sound field to extract depth information and create depth maps before the actual rendering process. This preliminary action prepares the necessary spatial data in advance, making the subsequent perspective mapping more accurate and efficient.
Solution Approach 2:
The system uses feedback from depth map analysis to adjust the sound field representation dynamically. By continuously monitoring the depth information and comparing it with the desired listener position, the system can accurately map the changed perspective and maintain precision in spatial audio rendering.
Data Source
AI summary
An apparatus for generating an enhanced sound field description includes: a sound field generator for generating at least one sound field description indicating a sound field with respect to at least one reference location; and a meta data generator for generating meta data relating to spatial information of the sound field, wherein the at least one sound field description and the meta data constitute the enhanced sound field description. The meta data can be a depth map associating a distance information to a direction in a full band or a subband, i.e., a time frequency bin.


