Self-Supervised Scale-Aware Monocular Depth Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Monocular cameras face challenges in estimating depth due to limited resolution, image artifacts, and scale ambiguities, which hinder situational awareness and navigation in robotic devices, especially when relying on self-supervised training without additional ground-truth data.
Innovation Solution
A depth system that employs a training architecture using monocular video and incorporates a pose model with a velocity component in the loss function to learn scale-aware depth estimates, eliminating the need for secondary ground-truth data by computing velocity supervision loss from instantaneous camera velocity measurements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If monocular cameras are used for depth estimation, then cost is reduced, but depth accuracy and scale awareness deteriorate
Solution Approach 1:
The system uses the monocular camera itself to provide the velocity information needed for scale-aware depth estimation. The camera's own motion data (velocity) is fed back into the loss function to enable the network to learn metric depth without external sensors or ground truth data, making the system self-sufficient and cost-effective
Solution Approach 2:
The invention changes the training parameter by incorporating velocity information into the loss function. This parameter modification enables the network to learn scale-aware depth estimates from monocular video sequences, transforming an underdetermined problem into a solvable one without adding hardware costs
2Quantity of substance
If self-supervised training is used without ground-truth data, then data availability is improved, but scale awareness deteriorates
Solution Approach 1:
The system implements feedback by using the network's own predictions and the camera's velocity measurements to construct a supervisory signal. The velocity-supervised loss function provides continuous feedback during training, enabling the network to learn scale-aware depth estimates without requiring external ground truth data
Solution Approach 2:
Velocity information acts as an intermediary that bridges the gap between monocular image data and metric depth estimation. This intermediate parameter (camera velocity) enables the training process to infer scale information without directly observing ground truth depth measurements
3Ease of operation
If photometric loss is used for training, then training simplicity is improved, but scale ambiguity worsens
Solution Approach 1:
The invention merges photometric loss with velocity-supervised loss into a combined training objective. This combination maintains the simplicity of photometric loss while adding scale-awareness through velocity constraints, achieving both training ease and scale accuracy simultaneously
Data Source
AI summary
System, methods, and other embodiments described herein relate to self-supervised training of a depth model for monocular depth estimation. In one embodiment, a method includes processing a first image of a pair according to the depth model to generate a depth map. The method includes processing the first image and a second image of the pair according to a pose model to generate a transformation that defines a relationship between the pair. The pair of images are separate frames depicting a scene of a monocular video. The method includes generating a monocular loss and a pose loss, the pose loss including at least a velocity component that accounts for motion of a camera between the training images. The method includes updating the pose model according to the pose loss and the depth model according to the monocular loss to improve scale awareness of the depth model in producing depth estimates.


