Monocular Depth Map Generation Using Volumetric Feature Projection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for estimating depth information in autonomous driving systems, such as using infrared rays, ultrasonic waves, or cameras, face challenges like device costs, calibration requirements, and computational complexity, highlighting the need for an efficient and accurate method using a monocular camera.
Innovation Solution
A method and apparatus that generate a depth map by obtaining surround-view images from monocular cameras, encoding multi-scale image features, projecting them into a three-dimensional space, and decoding volumetric features to produce accurate depth maps, while also obtaining pose information and training neural networks for improved accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If depth information is estimated using infrared rays or ultrasonic waves, then depth measurement can be achieved, but the reflected signal is affected by the state of the object
Solution Approach 1:
The patent replaces physical sensing methods (infrared, ultrasonic) with a computational approach using monocular camera images. Instead of relying on physical signals that interact with objects, the system uses image processing and neural networks to estimate depth, eliminating the problem of signal reflection variability.
2Measurement precision
If depth information is estimated using laser signals, then high accuracy is achieved, but expensive additional devices are required
Solution Approach 1:
The patent extracts depth information from standard monocular camera images without requiring additional depth-sensing devices like lasers. The system processes existing visual data through multi-scale feature encoding and volumetric feature projection to obtain depth maps, eliminating expensive hardware requirements.
Solution Approach 2:
The patent creates a virtual three-dimensional representation (volumetric feature) from two-dimensional image data. By encoding image features across multiple scales and projecting them into 3D space, the system generates depth information as a computational copy of the physical scene without needing physical 3D sensing devices.
3Measurement precision
If depth information is generated using stereo camera disparity calculation, then depth estimation can be achieved, but precise calibration of two cameras is required and it takes a lot of time to calculate disparity
Solution Approach 1:
The patent segments the image processing into multi-scale feature encoding stages, processing features at different resolutions separately. This segmentation allows efficient feature extraction without requiring complex camera calibration, as each scale level contributes to the final depth estimation independently through volumetric feature projection.
Solution Approach 2:
The patent replaces the mechanical calibration and disparity calculation process of stereo cameras with a neural network-based monocular depth estimation system. The system uses learned features from multi-scale encoding and volumetric projection to directly estimate depth, eliminating the need for precise camera calibration and computationally intensive disparity algorithms.
4Measurement precision
If multi-scale image features are encoded and projected into three-dimensional space, then accurate depth maps can be generated, but computational complexity increases
Solution Approach 1:
The patent segments the computational process into distinct stages: multi-scale feature encoding at different resolution levels, volumetric feature projection, and depth map decoding. This segmentation allows efficient processing by handling features at appropriate scales at each stage, reducing overall computational complexity while maintaining depth map accuracy.
Data Source
AI summary
Provided are an apparatus and method for generating a depth map by using a volumetric feature. The method may generate a single feature map for a base image included in a surround-view image by performing encoding and postprocessing on the base image, and generate a volumetric feature by encoding the single feature map with depth information and then projecting a result of the encoding into a three-dimensional space. In this method, a depth map of a surround-view image may be generated by using a depth decoder to decode a volumetric feature.


