3D Audio Depth Mapping for Spatial Synchronization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing audio rendering technologies face challenges in achieving flexible and synchronized spatial audio rendering for three-dimensional video, particularly due to variations in display capabilities, leading to a degraded user experience.

Innovation Solution

An audio signal processing apparatus that includes a receiver for audio data with depth position information, a determiner to assess the depth rendering properties of a three-dimensional display, and a mapper that adjusts audio object positions based on the display's depth rendering capabilities, ensuring improved spatial synchronization between audio and video.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio objects are rendered at fixed target positions without adaptation, then audio rendering is simple and straightforward, but spatial synchronization with video degrades when display depth rendering capabilities vary

Engineering Contradiction:
Improvespatial synchronizationVSAvoidaudio processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The audio rendering system dynamically adapts audio object positions based on the detected depth rendering capabilities of the display device. Instead of using fixed target positions, the system modifies audio positions in real-time according to display characteristics, ensuring spatial synchronization across different devices while maintaining processing simplicity through automated detection and adjustment

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system obtains depth rendering capability information from the display device and uses this feedback to adjust audio object positions accordingly. This closed-loop approach ensures that audio rendering adapts to the specific capabilities of each display, improving spatial synchronization without requiring complex manual configuration

Inventive Principle:
Principle #23Feedback

2Reliability

If audio objects are rendered at target positions without depth adaptation, then audio rendering is straightforward, but spatial coherence with visual depth perception degrades

Engineering Contradiction:
Improvespatial coherenceVSAvoidrendering processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system changes the depth parameter of audio objects based on the display's depth rendering capabilities. By detecting the display's ability to render visual depth and adjusting audio depth positions accordingly, the system achieves spatial coherence between audio and visual perceptions while keeping the processing approach systematic and manageable

Inventive Principle:
Principle #35Parameter changes

3Reliability

If audio rendering adapts to each display's depth rendering capabilities, then spatial synchronization improves, but processing complexity and computational resources increase

Engineering Contradiction:
Improvespatial synchronizationVSAvoidrendering efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary detection of depth rendering capabilities and pre-calculates appropriate audio position mappings before actual audio rendering. This preparation work is done once per display device, allowing subsequent audio rendering to proceed efficiently without repeated complex calculations, thus maintaining both high spatial synchronization and rendering efficiency

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10575112B2Method and apparatus for processing an audio signal associated with a video image
Publication Date: 2020.02.25 KONINKLIJKE PHILIPS NV
  • US10575112B2 patent drawing
  • US10575112B2 patent drawing
  • US10575112B2 patent drawing

AI summary

An audio signal processing apparatus includes a receiver which receives an audio signal including audio data for at least a first audio object associated with a three dimensional image. The audio signal also includes depth position data indicative of a target depth position for the first audio object. A determiner determines a visual rendering depth range for a target three dimensional display for presenting the three dimensional image, and a mapper for mapping the target depth position to a rendering depth position for the audio object where the mapping is dependent on the visual rendering depth range. The visual rendering depth range may specifically be a depth range in which the three dimensional display can accurately render objects, and the mapper may amend the positions of audio object sound sources such that these match the depth positions of corresponding visual objects presented by the three dimensional display.