Video-Based 3D Audio Generation With Height-Direction Inference

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing two-dimensional audio signals lack audio information in the height direction, making it difficult to generate three-dimensional audio signals with spatial stereoscopic effects, and current methods for generating three-dimensional audio from two-dimensional signals are time-consuming and challenging.

Innovation Solution

A video processing device and method that utilizes deep neural networks to analyze video signals and two-dimensional audio signals, extracting altitude and planar components to generate a three-dimensional audio signal by synchronizing and correcting audio and video features, using a combination of deep neural networks to enhance audio signals with height direction information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If deep neural networks are used to generate three-dimensional audio signals from two-dimensional audio signals and video information, then the generation process becomes easier and more efficient, but the device complexity increases

Engineering Contradiction:
Improveease of generationVSAvoidsystem complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by pre-processing video signals to extract feature information (movement, position, altitude of objects) before audio generation. The deep neural networks are trained in advance with synchronized audio-video data to learn the mapping relationships, enabling automated three-dimensional audio generation without manual intervention during actual operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces video information as an intermediary element that bridges the gap between two-dimensional audio signals and three-dimensional spatial audio. The video data provides altitude and movement information that mediates the transformation process, allowing the system to infer three-dimensional audio characteristics from two-dimensional audio through video-based feature extraction and correlation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If manual mixing and monitoring methods are used to generate three-dimensional audio signals from two-dimensional audio signals, then the system complexity remains low, but the process becomes time-consuming and difficult

Engineering Contradiction:
Improvegeneration efficiencyVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent replaces manual mechanical mixing and monitoring operations with automated deep neural network-based processing. The DNNs automatically analyze video features, extract relevant parameters (position, movement, altitude), and generate three-dimensional audio signals without requiring human operators to perform time-consuming manual adjustments and monitoring.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system changes parameters by transforming two-dimensional audio signals into three-dimensional audio space using video-derived parameters. The deep neural networks utilize video-based parameters (object altitude, movement velocity, position) to modify audio signal characteristics, enabling automated generation of spatial audio with correct temporal and spatial parameters.

Inventive Principle:
Principle #35Parameter changes

3Loss of information

If two-dimensional audio signals are used without video information, then the system remains simple, but audio information in the height direction is uncertain or absent

Engineering Contradiction:
Improveaudio information completenessVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies dimensionality change by transitioning from two-dimensional audio signals to three-dimensional audio signals. Video information provides the missing height/altitude dimension, and the deep neural networks process this additional dimensional data to reconstruct three-dimensional spatial audio information that is absent in conventional two-dimensional audio.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The system achieves multi-functionality by using video information for multiple purposes: extracting object position, movement, and altitude; synchronizing with audio signals; and providing features for three-dimensional audio generation. This universal use of video data compensates for the limitations of two-dimensional audio across multiple functional requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12425792B2Video processing device and method
Publication Date: 2025.09.23 SAMSUNG ELECTRONICS CO LTD
  • US12425792B2 patent drawing
  • US12425792B2 patent drawing
  • US12425792B2 patent drawing

AI summary

A video processing apparatus includes a memory storing instructions, and at least one processor configured to execute the instructions to generate a plurality of feature information by analyzing a video signal comprising a plurality of images based on a first DNN, extract a first altitude component and a first planar component corresponding to a movement of an object in a video from the video signal based on a second DNN, extract a second planar component corresponding to a movement of a sound source in audio from a first audio signal based on a third DNN, generate a second altitude component based on the first altitude component, the first planar component, and the second planar component, output a second audio signal comprising the second altitude component based on the feature information, and synchronize the second audio signal with the video signal and output the synchronized second audio signal and video signal.