Automatic Depth Map Generation for 2D Video Using Saliency and Structure Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for converting 2D videos to 3D require significant manual effort or user interaction, limiting their application in real-time processing and general video sequences, especially for achieving immersive 3D experiences on 3D TVs.

Innovation Solution

An apparatus and method that automatically generates depth maps for 2D images in video sequences using a combination of 3D structure matching, saliency mapping, and spatial-temporal smoothing, allowing for the creation of depth maps without user input, by calculating matching scores and saliency values based on feature distributions and human visual perception models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If manual labeling is used to convert 2D video to 3D, then the conversion quality is satisfying, but too much manpower is required

Engineering Contradiction:
Improveconversion qualityVSAvoidmanpower efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The system performs automatic depth map generation through self-service mechanisms including motion field calculation, occlusion map generation, and depth map synthesis without requiring manual intervention. The computer automatically processes video frames using algorithms that compute motion vectors, generate occlusion maps, and synthesize depth maps, eliminating the need for human operators while maintaining conversion quality

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical manual labeling process with an automated computational system. Instead of human workers manually creating depth maps, the system uses computer-based algorithms including motion compensation, occlusion handling, and depth synthesis to automatically generate depth maps from 2D video sequences, substituting human labor with mechanical computation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If stereo video difference simulation is used based on motion visual difference, then horizontal object movement can be processed, but it is difficult to process general video in real-time

Engineering Contradiction:
Improvevideo type compatibilityVSAvoidreal-time processing capability
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The system performs preliminary actions by pre-calculating and storing motion fields, occlusion maps, and depth information for typical video structures. Motion compensation is performed in advance using reference frames, and occlusion maps are generated beforehand to facilitate rapid depth map synthesis during real-time processing, enabling the system to handle general video sequences efficiently

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent employs parameter changes by adjusting motion compensation parameters, occlusion map thresholds, and depth synthesis coefficients dynamically based on video content characteristics. The system modifies processing parameters such as motion vector search ranges, occlusion detection sensitivity, and depth map resolution to optimize real-time processing performance while maintaining adaptability to different video types

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If a depth display system requiring computer interaction is provided, then depth maps can be generated, but it is difficult to realize unmanned monitoring and real-time operation

Engineering Contradiction:
Improvedepth map accuracyVSAvoidunmanned operation capability
Core Design Contradiction:
Manufacturing precisionVSExtent of automation

Solution Approach 1:

The system achieves complete automation through self-service mechanisms where the computer automatically performs motion field calculation, occlusion map generation, and depth map synthesis without user interaction. The algorithm autonomously processes video frames, calculates motion vectors, handles occlusions, and generates depth maps, enabling unmanned monitoring and real-time operation while maintaining depth map accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system implements feedback mechanisms by using calculated motion fields to adjust occlusion map generation, and using occlusion information to refine depth map synthesis. The algorithm continuously refines its output by feeding back intermediate results (motion vectors, occlusion probabilities) to subsequent processing stages, ensuring accurate depth map generation while maintaining automated operation

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8553972B2Apparatus, method and computer-readable medium generating depth map
Publication Date: 2013.10.08 SAMSUNG ELECTRONICS CO LTD
  • US8553972B2 patent drawing
  • US8553972B2 patent drawing
  • US8553972B2 patent drawing

AI summary

Disclosed are an apparatus, a method and a computer-readable medium automatically generating a depth map corresponding to each two-dimensional (2D) image in a video. The apparatus includes an image acquiring unit to acquire a plurality of 2D images that are temporally consecutive in an input video, a saliency map generator to generate at least one saliency map corresponding to a current 2D image among the plurality of 2D images based on a Human Visual Perception (HVP) model, a saliency-based depth map generator, a three-dimensional (3D) structure matching unit to calculate matching scores between the current 2D image and a plurality of 3D typical structures that are stored in advance, and to determine a 3D typical structure having a highest matching score among the plurality of 3D typical structures to be a 3D structure of the current 2D image, a matching-based depth map generator; a combined depth map generator to combine the saliency-based depth map and the matching-based depth map and to generate a combined depth map, and a spatial and temporal smoothing unit to spatially and temporally smooth the combined depth map.