3D Sound Rendering from Mono Audio Using Video Object Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing 3D sound rendering techniques require multi-channel audio inputs, which are not available in many scenarios, limiting the ability to create immersive sound experiences with mono audio content.

Innovation Solution

A method and system that analyzes video content to determine object positions and trajectories, classifies audio sources, and distributes audio streams into multiple channels based on this information to create 3D sound effects using mono audio input.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If multi-channel audio input is used, then 3D sound rendering quality is improved, but device complexity and input requirements worsen

Engineering Contradiction:
Improvesound localization accuracyVSAvoidaudio input requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates virtual multi-channel audio copies from a single mono audio source by synthesizing spatial audio information. Instead of requiring physical multi-channel microphones, the system generates synthetic audio channels that mimic what would be captured by a multi-channel array, using video-derived spatial information to distribute the mono audio signal across multiple virtual channels.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces video content analysis as an intermediary between the mono audio input and the multi-channel output. By extracting object positions and trajectories from video, the system mediates the transformation of spatial information to guide how the mono audio is distributed across multiple channels, enabling 3D sound rendering without direct multi-channel audio capture.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If mono audio input is used, then ease of operation is improved, but sound spatial distribution capability worsens

Engineering Contradiction:
Improveaudio input simplicityVSAvoidspatial audio distribution
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent adds a spatial dimension to mono audio by using video-derived object positions and trajectories. It transforms 2D video spatial information into 3D audio spatial distribution, mapping object locations in the video frame to corresponding audio channel distributions, thereby enabling spatial audio effects from a non-spatial audio source.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent makes the mono audio input universal by enabling it to serve multiple spatial distribution patterns simultaneously. The same mono audio signal can be distributed to different channels based on different object positions and motions, allowing a single audio source to create diverse spatial audio experiences adaptable to various video content scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If audio streams are separated and distributed based on object analysis, then sound rendering quality is improved, but processing complexity worsens

Engineering Contradiction:
Improveaudio source localizationVSAvoidprocessing requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary analysis of video content to extract object positions, trajectories, and classifications before processing the audio distribution. By pre-processing the video information and preparing spatial mapping data in advance, the system reduces the real-time processing burden during audio rendering, as the spatial parameters are already determined from video analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent segments the audio processing task by separating audio stream separation from spatial distribution. It first separates audio streams associated with different objects, then distributes each segment to appropriate channels based on object positions. This segmentation allows independent optimization of each processing stage and reduces overall system complexity.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12425797B2Three-dimensional (3D) sound rendering with multi-channel audio based on mono audio input
Publication Date: 2025.09.23 SAMSUNG ELECTRONICS CO LTD
  • US12425797B2 patent drawing
  • US12425797B2 patent drawing
  • US12425797B2 patent drawing

AI summary

A method includes obtaining video content and associated substantially mono audio content. The method also includes determining at least one of a position or a motion trajectory of each of one or more objects detected in the video content and classifying each of the one or more objects into one of multiple object classes. The method further includes separating audio streams within the audio content based on the video content. Each of the audio streams is associated with one of multiple audio sources. The method also includes classifying each of the audio sources into one of the object classes. In addition, the method includes, for each audio source classified into the same object class as one of the one or more objects, distributing the audio stream associated with that audio source into multiple audio channels based on at least one of the position or the motion trajectory of that object.