Depth Map Generation Using Motion Cues for 2D to 3D Conversion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for converting 2D visual content to 3D are inefficient, time-consuming, and limited in their ability to handle large volumes or general types of content, requiring costly and technical processes.

Innovation Solution

An image converter identifies subset frames, determines global and dense motion values, calculates a rough depth map, and interpolates depth values for pixels to render 3D video, using a feature-to-depth mapping function for accurate and automatic conversion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional manual processes are used for 2D to 3D conversion, then conversion accuracy can be maintained through expert judgment, but the process becomes extremely time-consuming and expensive

Engineering Contradiction:
Improveconversion accuracyVSAvoidconversion time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical process of expert judgment with an automated computer-based system that uses motion cue analysis and feature-to-depth mapping algorithms to automatically generate depth maps and convert 2D content to 3D, dramatically reducing conversion time while maintaining quality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system enables self-service conversion by automatically analyzing motion cues in video frames and generating depth information without requiring manual intervention, allowing the conversion process to serve itself through automated algorithmic processing

Inventive Principle:
Principle #25Self-service

2Adaptability or versatility

If conventional conversion techniques are used, then specific types of images and video can be converted, but the method is limited and cannot handle general 2D to 3D conversion tasks

Engineering Contradiction:
Improveconversion applicabilityVSAvoidprocess complexity
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent creates a universal conversion system that can handle various types of 2D content (video, images, synthetic imagery) through a single automated platform that analyzes motion cues and applies feature-to-depth mapping, making the process adaptable to different content types without requiring separate specialized methods

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If fully automatic conversion is implemented, then processing efficiency and productivity are greatly improved, but the complexity of the conversion process increases

Engineering Contradiction:
Improveconversion efficiencyVSAvoidconversion process complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent introduces intermediate processing steps including motion cue analysis, optical flow calculation, and feature-to-depth mapping as mediators between the input 2D content and the final 3D output, breaking down the complex conversion process into manageable sequential stages that improve automation while controlling overall complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9661307B1Depth map generation using motion cues for conversion of monoscopic visual content to stereoscopic 3D
Publication Date: 2017.05.23 GOOGLE LLC
  • US9661307B1 patent drawing
  • US9661307B1 patent drawing
  • US9661307B1 patent drawing

AI summary

An image converter identifies a subset of frames in a two-dimensional video and determines a global camera motion value for the subset of frames. The image converter also determines a dense motion value for a plurality of pixels in the subset of frames and compares the global camera motion value and the dense motion value to calculate a rough depth map for the subset of frames. The image converter further interpolates, based on the rough depth map, a depth value for each of the plurality of pixels in the subset of frames and renders a three-dimensional video from the subset of frames using the depth value for each of the plurality of pixels.