2D to 3D Video Conversion via Depth Map Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for converting 2D images to 3D video are time-consuming, costly, and limited, often requiring multiple images from different viewpoints, which restricts their applicability to specific types of content.

Innovation Solution

A 3D video generator determines depth values for each pixel in a 2D image, generates a depth map, and uses view disparity to create modified images, forming a sequence of frames in a 3D video, allowing for fully automatic conversion of 2D content to 3D using a single image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional techniques use multiple images from different viewpoints to create 3D video, then the quality of 3D content is improved, but the complexity of the process and cost increase significantly

Engineering Contradiction:
Improvequality of 3D contentVSAvoidprocess complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the complex multi-image processing task into distinct functional modules: depth map generation from single image, view disparity calculation, and frame synthesis. This modular approach maintains 3D quality while reducing process complexity by handling each aspect separately rather than requiring simultaneous processing of multiple viewpoint images.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces depth maps as an intermediary representation that bridges the gap between single 2D images and multi-view 3D video. By first extracting depth information and then using it to generate synthetic viewpoints, the system achieves 3D quality comparable to multi-image methods while avoiding the complexity of coordinating multiple image captures.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If manual processes are used to create 3D content, then the precision and quality are improved, but the time consumption and cost increase

Engineering Contradiction:
Improve3D content qualityVSAvoidtime consumption
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs automatic depth map generation and view synthesis without requiring manual 3D modeling or post-processing. The algorithm autonomously extracts depth information from the input image and generates the complete 3D video sequence, eliminating time-consuming manual operations while maintaining quality through sophisticated image processing.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical processes (physical camera setup, manual 3D modeling) with automated computational algorithms. The depth map generation and view synthesis are performed through software-based image processing, dramatically reducing time consumption while maintaining or improving precision through consistent algorithmic application.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If conventional conversion techniques are limited to specific image types, then the conversion accuracy for those types is improved, but the versatility of the system deteriorates

Engineering Contradiction:
Improveconversion accuracyVSAvoidapplicability to different content
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent develops a universal depth map generation algorithm that can process various types of 2D images (photographs, illustrations, graphics) using the same color-to-depth mapping approach. The system adapts to different content types by adjusting parameters within the unified framework, maintaining conversion accuracy across diverse image types while achieving broad versatility.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system maintains versatility by adjusting parameters such as color space mappings and depth scale factors to suit different image types. By changing these parameters rather than fundamentally altering the conversion algorithm, the system achieves accurate conversion for photographs, illustrations, and graphics using a single unified approach.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If fully automatic conversion is implemented, then the productivity is improved, but the complexity of the conversion algorithm increases

Engineering Contradiction:
Improveconversion efficiencyVSAvoidalgorithm complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent performs preliminary depth map generation from the input image before view synthesis. This pre-computed depth information serves as a foundation for subsequent frame generation, enabling fully automatic processing without requiring iterative adjustments during view synthesis. The preliminary depth extraction simplifies the overall algorithm structure while maintaining high conversion efficiency.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10154242B1Conversion of 2D image to 3D video
Publication Date: 2018.12.11 GOOGLE LLC
  • US10154242B1 patent drawing
  • US10154242B1 patent drawing
  • US10154242B1 patent drawing

AI summary

A two-dimensional input image to be used in a creation of a three-dimensional video may be received and depth values for pixels in the image may be determined. A depth map may be generated based on the depth values for the pixels and pixel shift values for the pixels may be calculated based on the depth map and a view disparity value. A modified image corresponding to a particular frame of the three-dimensional video may be generated based on the input image and the pixel shift values. An additional modified image corresponding to a next frame of the three-dimensional video may be generated based on the modified image and the pixel shift values used to generate the modified image where the modified image in combination with the input image and the additional modified image are a sequence of frames in the three-dimensional video.