Automatic Depth Map Generation for 2D Video Using Saliency and Structure Matching
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for converting 2D videos to 3D require significant manual effort or user interaction, limiting their application in real-time processing and general video sequences, especially for achieving immersive 3D experiences on 3D TVs.
Innovation Solution
An apparatus and method that automatically generates depth maps for 2D images in video sequences using a combination of 3D structure matching, saliency mapping, and spatial-temporal smoothing, allowing for the creation of depth maps without user input, by calculating matching scores and saliency values based on feature distributions and human visual perception models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If manual labeling is used to convert 2D video to 3D, then the conversion quality is satisfying, but too much manpower is required
Solution Approach 1:
The system performs automatic depth map generation through self-service mechanisms including motion field calculation, occlusion map generation, and depth map synthesis without requiring manual intervention. The computer automatically processes video frames using algorithms that compute motion vectors, generate occlusion maps, and synthesize depth maps, eliminating the need for human operators while maintaining conversion quality
Solution Approach 2:
The patent replaces the mechanical manual labeling process with an automated computational system. Instead of human workers manually creating depth maps, the system uses computer-based algorithms including motion compensation, occlusion handling, and depth synthesis to automatically generate depth maps from 2D video sequences, substituting human labor with mechanical computation
2Adaptability or versatility
If stereo video difference simulation is used based on motion visual difference, then horizontal object movement can be processed, but it is difficult to process general video in real-time
Solution Approach 1:
The system performs preliminary actions by pre-calculating and storing motion fields, occlusion maps, and depth information for typical video structures. Motion compensation is performed in advance using reference frames, and occlusion maps are generated beforehand to facilitate rapid depth map synthesis during real-time processing, enabling the system to handle general video sequences efficiently
Solution Approach 2:
The patent employs parameter changes by adjusting motion compensation parameters, occlusion map thresholds, and depth synthesis coefficients dynamically based on video content characteristics. The system modifies processing parameters such as motion vector search ranges, occlusion detection sensitivity, and depth map resolution to optimize real-time processing performance while maintaining adaptability to different video types
3Manufacturing precision
If a depth display system requiring computer interaction is provided, then depth maps can be generated, but it is difficult to realize unmanned monitoring and real-time operation
Solution Approach 1:
The system achieves complete automation through self-service mechanisms where the computer automatically performs motion field calculation, occlusion map generation, and depth map synthesis without user interaction. The algorithm autonomously processes video frames, calculates motion vectors, handles occlusions, and generates depth maps, enabling unmanned monitoring and real-time operation while maintaining depth map accuracy
Solution Approach 2:
The system implements feedback mechanisms by using calculated motion fields to adjust occlusion map generation, and using occlusion information to refine depth map synthesis. The algorithm continuously refines its output by feeding back intermediate results (motion vectors, occlusion probabilities) to subsequent processing stages, ensuring accurate depth map generation while maintaining automated operation
Data Source
AI summary
Disclosed are an apparatus, a method and a computer-readable medium automatically generating a depth map corresponding to each two-dimensional (2D) image in a video. The apparatus includes an image acquiring unit to acquire a plurality of 2D images that are temporally consecutive in an input video, a saliency map generator to generate at least one saliency map corresponding to a current 2D image among the plurality of 2D images based on a Human Visual Perception (HVP) model, a saliency-based depth map generator, a three-dimensional (3D) structure matching unit to calculate matching scores between the current 2D image and a plurality of 3D typical structures that are stored in advance, and to determine a 3D typical structure having a highest matching score among the plurality of 3D typical structures to be a 3D structure of the current 2D image, a matching-based depth map generator; a combined depth map generator to combine the saliency-based depth map and the matching-based depth map and to generate a combined depth map, and a spatial and temporal smoothing unit to spatially and temporally smooth the combined depth map.


