Monocular Depth Map Generation Using Volumetric Feature Projection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for estimating depth information in autonomous driving systems, such as using infrared rays, ultrasonic waves, or cameras, face challenges like device costs, calibration requirements, and computational complexity, highlighting the need for an efficient and accurate method using a monocular camera.

Innovation Solution

A method and apparatus that generate a depth map by obtaining surround-view images from monocular cameras, encoding multi-scale image features, projecting them into a three-dimensional space, and decoding volumetric features to produce accurate depth maps, while also obtaining pose information and training neural networks for improved accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If depth information is estimated using infrared rays or ultrasonic waves, then depth measurement can be achieved, but the reflected signal is affected by the state of the object

Engineering Contradiction:
Improvedepth measurement capabilityVSAvoidsignal stability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent replaces physical sensing methods (infrared, ultrasonic) with a computational approach using monocular camera images. Instead of relying on physical signals that interact with objects, the system uses image processing and neural networks to estimate depth, eliminating the problem of signal reflection variability.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If depth information is estimated using laser signals, then high accuracy is achieved, but expensive additional devices are required

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoiddevice cost
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts depth information from standard monocular camera images without requiring additional depth-sensing devices like lasers. The system processes existing visual data through multi-scale feature encoding and volumetric feature projection to obtain depth maps, eliminating expensive hardware requirements.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent creates a virtual three-dimensional representation (volumetric feature) from two-dimensional image data. By encoding image features across multiple scales and projecting them into 3D space, the system generates depth information as a computational copy of the physical scene without needing physical 3D sensing devices.

Inventive Principle:
Principle #26Copying

3Measurement precision

If depth information is generated using stereo camera disparity calculation, then depth estimation can be achieved, but precise calibration of two cameras is required and it takes a lot of time to calculate disparity

Engineering Contradiction:
Improvedepth estimation capabilityVSAvoidcalibration complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the image processing into multi-scale feature encoding stages, processing features at different resolutions separately. This segmentation allows efficient feature extraction without requiring complex camera calibration, as each scale level contributes to the final depth estimation independently through volumetric feature projection.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces the mechanical calibration and disparity calculation process of stereo cameras with a neural network-based monocular depth estimation system. The system uses learned features from multi-scale encoding and volumetric projection to directly estimate depth, eliminating the need for precise camera calibration and computationally intensive disparity algorithms.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Measurement precision

If multi-scale image features are encoded and projected into three-dimensional space, then accurate depth maps can be generated, but computational complexity increases

Engineering Contradiction:
Improvedepth map accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the computational process into distinct stages: multi-scale feature encoding at different resolution levels, volumetric feature projection, and depth map decoding. This segmentation allows efficient processing by handling features at appropriate scales at each stage, reducing overall computational complexity while maintaining depth map accuracy.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20240037791A1Apparatus and method for generating depth map by using volumetric feature
Publication Date: 2024.02.01 42DOT INC
  • US20240037791A1 patent drawing
  • US20240037791A1 patent drawing
  • US20240037791A1 patent drawing

AI summary

Provided are an apparatus and method for generating a depth map by using a volumetric feature. The method may generate a single feature map for a base image included in a surround-view image by performing encoding and postprocessing on the base image, and generate a volumetric feature by encoding the single feature map with depth information and then projecting a result of the encoding into a three-dimensional space. In this method, a depth map of a surround-view image may be generated by using a depth decoder to decode a volumetric feature.