Monocular Depth Estimation via Supervised Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for estimating depths from monocular images are limited by assumptions such as noiseless images and dynamic Bayesian models, leading to inaccuracies and require costly equipment like binocular stereo vision systems or laser scanners, which are complex and capital-intensive.

Innovation Solution

A supervised learning approach using a training set of monocular images and their corresponding depth maps, transforming images into intensity and color channels, segmenting them into patches, calculating feature vectors with specific filters, and applying global optimization to estimate depth maps through equations involving matting Laplacian matrices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If binocular stereo vision system or laser scanner is used for depth estimation, then measurement precision is improved, but device complexity and cost increase significantly

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical/optical systems (binocular stereo vision, laser scanners) with a computational approach using monocular images processed through supervised learning algorithms. The depth estimation is achieved by training a neural network on labeled depth data, substituting physical depth-sensing hardware with software-based analysis of standard images.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a computational model that learns to replicate the depth estimation capability of expensive hardware systems. By training on datasets containing both monocular images and corresponding depth maps, the system creates a virtual copy of depth-sensing functionality that can be deployed without physical depth sensors.

Inventive Principle:
Principle #26Copying

2Device complexity

If conventional algorithms with assumptions (noiseless images, dynamic Bayesian model) are used, then device complexity is reduced, but measurement precision deteriorates due to restrictive assumptions

Engineering Contradiction:
Improvesystem complexityVSAvoiddepth estimation accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary training action by collecting and processing large datasets of monocular images with corresponding depth maps before actual depth estimation. This pre-training phase enables the system to learn from diverse real-world conditions including various lighting, textures, and depth configurations, making the model robust without requiring restrictive assumptions during operation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the depth estimation problem from a deterministic algorithmic approach to a probabilistic learning approach. By changing from fixed algorithmic parameters to learned model parameters through supervised training, the system can handle noise and variability in real images while maintaining computational efficiency.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS8284998B2Method of estimating depths from a single image displayed on display
Publication Date: 2012.10.09 ARCSOFT CORP LTD
  • US8284998B2 patent drawing
  • US8284998B2 patent drawing
  • US8284998B2 patent drawing

AI summary

A method of estimating depths on a monocular image displayed on a display is utilized for improving correctness of depths shown on the display. Feature vectors are calculated for each patch on the monocular image for determining an intermediate depth map of the monocular image in advance. For improving the correctness of the intermediate depth map, an energy function in forms of vectors is minimized for calculating a best solution of the depth map of the monocular image. Therefore, the display may display the monocular image according to a calculated output depth map for having an observer of the display to correctly perceive depths on the monocular image.