3D Audio Positioning Using Visual Depth Maps
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current digital 3D movie preparation methods often lack efficient conversion of audio components to preserve and include 3D space perception cues, especially in real-time applications and environments like video game rendering and live event broadcasting, and there is a need for an efficient scheme to optimize digital 3D movie preparation and conversion with audio 3D space perception cues for various formats.
Innovation Solution
A system that uses an audio/video encoder to convert content with raw audio tracks and visual object position information into a 2D or 3D format, employing an audio processor to generate encoding coefficients and tracking maps for both visual and audio objects, ensuring accurate positioning of audio objects in 3D space, even when they are not visually represented, using manual or automated methods.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional audio encoding formats (mono to stereo to multi-channel) are used, then audio compression and compatibility are achieved, but 3D space perception cues are lost or insufficient
Solution Approach 1:
The patent transitions from traditional 2D audio planes (stereo/multi-channel) to 3D audio space by introducing depth maps and spatial positioning cues. Audio objects are positioned in three-dimensional space using depth information derived from visual depth maps, enabling listeners to perceive sound sources in front of, behind, or at the screen plane rather than limited to left-right-front channels.
Solution Approach 2:
The patent introduces depth maps as an intermediary element that bridges visual and audio components. The depth map, generated from visual 3D content, serves as a mediator to guide audio positioning and spatial cues, allowing the audio system to leverage visual depth information without requiring complex direct sensing of the acoustic environment.
2Reliability
If audio objects are positioned to match visual objects in 3D space, then audio-visual synchronization and immersion are improved, but tracking accuracy deteriorates when objects are off-screen or occluded
Solution Approach 1:
The patent applies preliminary action by using visual object tracking and depth map generation before audio positioning. The system first identifies and tracks visual objects in the scene, generates depth maps representing the 3D spatial layout, and pre-determines audio object positions based on this visual information. This preliminary visual analysis enables accurate audio positioning even when audio objects are subsequently occluded or move off-screen, as the spatial framework is already established.
3Reliability
If 3D audio positioning is implemented for all audio objects, then spatial immersion is enhanced, but processing time and computational resources increase
Solution Approach 1:
The patent applies local quality by selectively applying 3D audio positioning to specific audio objects rather than uniformly processing all audio. The system identifies audio objects of interest that benefit most from spatial positioning (such as dialogue sources, prominent sound effects) and applies complex 3D positioning cues to these while using simpler positioning or traditional channel assignment for less critical audio elements, thereby reducing overall processing complexity.
Solution Approach 2:
The patent creates a universal audio positioning framework that can adapt to different content types and requirements. The same depth map-based positioning system can be applied to various scenarios including movie playback, video games, and live broadcasting, and can adjust the level of processing detail based on the specific application needs, balancing quality and processing requirements across different use cases.
Data Source
Figure 1
Figure 1A
Figure 2A~2D
AI summary
An apparatus generating audio cues for content indicative of the position of audio objects within the content comprising: an audio processor receiving raw audio tracks for said content and information indicative of the positions of at least some of said audio tracks within frames of said content, said audio processor generating corresponding audio parameters; an authoring tool receiving said audio parameters and generating encoding coefficients, said audio parameters including audio cue of the position of audio objects corresponding to said tracks in at least one spatial dimension; and a first audio/video encoder receiving an input and encoding said input into an audio visual content having visual objects and audio objects, said audio objects being disposed at location corresponding to said one spatial position, said encoder using said encoding coefficients for said encoding.