Depth-Extended DirAC Sound Field Description for 6DoF Reproduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current sound field representations, such as Ambisonics and DirAC, do not provide sufficient information for translational shifts in six-degrees-of-freedom (6DoF) applications, as they lack information about object distances and absolute positions, making it impossible to compute sound field descriptions at different listener positions.

Innovation Solution

Generating an enhanced sound field description that includes meta data relating to spatial information, such as depth maps, to facilitate the calculation of modified sound fields at different reference locations, allowing for 6DoF reproduction by integrating distance information into the parametric representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional Ambisonics or DirAC sound field representations are used, then the sound field can be represented with limited information, but translational shifts in 6DoF applications become impossible

Engineering Contradiction:
Improve6DoF reproduction capabilityVSAvoidspatial information completeness
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent extends the traditional 2D sound field representation by adding a depth dimension. Depth maps are introduced as additional data that specify the distance of sound sources from the reference location, transforming the representation from a 2D angular distribution to a 3D spatial distribution, thereby enabling 6DoF reproduction.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The sound field representation is segmented into multiple components: traditional Ambisonics channels for angular information and separate depth maps for distance information. This segmentation allows each component to be processed independently and combined to achieve the full 6DoF effect.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If depth information is added to the sound field representation, then 6DoF reproduction becomes possible, but the data structure becomes more complex

Engineering Contradiction:
Improvetranslational shift capabilityVSAvoiddata structure complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

Instead of fundamentally redesigning the data structure, the patent adds depth information as an additional dimension. Depth maps are appended to the existing Ambisonics data, creating an extended representation that maintains compatibility with traditional systems while enabling new capabilities.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces intermediate processing steps that convert the extended sound field description with depth maps into the desired 6DoF representation. These intermediaries handle the complexity of integrating multiple data types and transforming them into the appropriate format for different reproduction scenarios.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If traditional sound field representations are used, then processing remains simple, but accurate mapping of changed listener perspective is impossible

Engineering Contradiction:
Improvelistener position accuracyVSAvoidprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary analysis of the sound field to extract depth information and create depth maps before the actual rendering process. This preliminary action prepares the necessary spatial data in advance, making the subsequent perspective mapping more accurate and efficient.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses feedback from depth map analysis to adjust the sound field representation dynamically. By continuously monitoring the depth information and comparing it with the desired listener position, the system can accurately map the changed perspective and maintain precision in spatial audio rendering.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11477594B2Concept for generating an enhanced sound-field description or a modified sound field description using a depth-extended DirAC technique or other techniques
Publication Date: 2022.10.18 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • US11477594B2 patent drawing
  • US11477594B2 patent drawing
  • US11477594B2 patent drawing

AI summary

An apparatus for generating an enhanced sound field description includes: a sound field generator for generating at least one sound field description indicating a sound field with respect to at least one reference location; and a meta data generator for generating meta data relating to spatial information of the sound field, wherein the at least one sound field description and the meta data constitute the enhanced sound field description. The meta data can be a depth map associating a distance information to a direction in a full band or a subband, i.e., a time frequency bin.