3D Audio Re-spatialization from Legacy 2D Video

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Legacy audiovisual media lacks full audio spatialization, resulting in a loss of spatial information when recorded in environments with reflective surfaces, as the direct and reflected sounds are mixed into mono or stereo formats, leading to a non-immersive listening experience.

Innovation Solution

A method to convert two-dimensional audio from legacy video into three-dimensional audio by isolating individual sound sources using source separation techniques, removing reverberation, and re-spatializing the direct sound components based on acoustic characteristics of the local area, using visual and audio features, and a mapping server to generate a local area impulse response for immersive audio presentation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If 2-D audio from legacy video is converted to 3-D audio using source separation and re-spatialization, then spatial awareness and immersion are enhanced, but device complexity and processing requirements increase

Engineering Contradiction:
Improvespatial informationVSAvoidaudio processing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The audio signal is segmented into individual sound sources using source separation techniques. Each sound source is processed independently to extract spatial characteristics and apply appropriate 3-D audio rendering, thereby recovering spatial information that was lost in the original 2-D mixing while managing complexity through modular processing of separate sources.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

An audio processing system acts as an intermediary between legacy 2-D audio content and the 3-D audio output. This intermediary performs source separation, spatial analysis, and re-spatialization operations, bridging the gap between simple 2-D recordings and immersive 3-D audio experiences without requiring changes to the original content or playback devices.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If reverberation is removed to obtain direct sound components, then spatial accuracy is improved, but processing complexity increases

Engineering Contradiction:
Improvespatial accuracyVSAvoidsignal processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

Reverberation components are extracted and removed from the audio signal to isolate direct sound components. This extraction process separates the useful direct sound information from the distracting reverberant energy, improving spatial accuracy by focusing on the direct path sound while eliminating the complex reflected sound fields that obscure spatial cues.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If local area impulse response is generated using visual and audio features, then audio realism is improved, but processing time and computational resources increase

Engineering Contradiction:
Improveaudio realismVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

Acoustic characteristics of the local area are determined in advance by analyzing visual features of the environment and audio reverberation properties. This preliminary analysis creates a model of the acoustic space that can be reused for multiple sound sources, reducing processing time while maintaining audio realism through pre-computed spatial and acoustic parameters.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The local area impulse response generation process serves multiple functions: it analyzes visual features of the environment, characterizes acoustic properties from audio reverberation, and creates a reusable acoustic model. This multi-functional approach consolidates processing efforts, reducing overall computational requirements and processing time while improving audio realism through comprehensive environmental characterization.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10721521B1Determination of spatialized virtual acoustic scenes from legacy audiovisual media
Publication Date: 2020.07.21 META PLATFORMS TECHNOLOGIES LLC
  • US10721521B1 patent drawing
  • US10721521B1 patent drawing
  • US10721521B1 patent drawing

AI summary

An audio system generates virtual acoustic environments with three-dimensional (3-D) sound from legacy video with two-dimensional (2-D) sound. The system relocates sound sources within the video from 2-D to into a 3-D geometry to create an immersive 3-D virtual scene of the video that can be viewed using a headset. Accordingly, an audio processing system obtains a video that includes flat mono or stereo audio being generated by one or more sources in the video. The system isolates the audio from each source by segmenting the individual audio sources. Reverberation is removed from the audio from each source to obtain each source's direct sound component. The direct sound component is then re-spatialized to the 3-D local area of the video to generate the 3-D audio based on acoustic characteristics obtained for the local area in the video.