Monocular Depth And Ego-Motion Training With GPS Scale Loss
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Monocular self-supervised depth and ego-motion estimation methods suffer from scale inconsistency due to the reliance on appearance-based losses, which limits training to small video sub-sequences without long sequence constraints, leading to inconsistent depth and ego-motion estimates across different video snippets.
Innovation Solution
The introduction of a GPS-to-scale loss (G2S) that synchronizes complementary GPS coordinates with images to enforce scale consistency and awareness, using an exponentially increasing relative weight in the training of neural networks, combined with appearance-based photometric and smoothness losses, to improve scale consistency and awareness in depth and ego-motion estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If appearance-based photometric loss is used for training, then the neural networks can be trained on small video sub-sequences with brightness consistency, but scale inconsistency arises across different video snippets
Solution Approach 1:
GPS coordinates are introduced as an intermediary reference system to mediate between appearance-based photometric loss and scale-consistent depth/ego-motion estimation. The GPS-to-scale loss uses GPS-derived distances as an intermediate metric to constrain the scale of predictions, bridging the gap between appearance-based training and metric-scale accuracy.
Solution Approach 2:
The invention changes the training parameter by introducing a new loss function (GPS-to-scale loss) that operates on a different parameter space (GPS coordinates and distances) than the appearance-based photometric loss. This additional parameter constraint enables scale consistency while maintaining the benefits of appearance-based training.
2Use of energy by stationary object
If monocular vision is used, then the system is compact, low-cost, and energy-efficient, but scale ambiguity and scale inconsistency occur
Solution Approach 1:
GPS coordinates serve as an intermediary metric reference that provides absolute scale information to the monocular vision system. This external reference mediator allows the low-cost monocular camera to achieve metric-scale accuracy without requiring expensive LiDAR sensors or complex multi-view geometry.
3Measurement precision
If self-supervised methods are used, then accurate depth maps can be produced from monocular video, but scale-inconsistency is introduced across video snippets
Solution Approach 1:
The GPS-to-scale loss provides feedback to the neural network training process by comparing predicted depths and ego-motions against GPS-derived distance measurements. This feedback mechanism continuously constrains the scale of predictions during training, ensuring scale consistency across different video snippets while maintaining the self-supervised learning paradigm.
4Productivity
If training is performed without GPS constraints, then appearance-based losses can be effectively utilized, but scale awareness is lost in predictions
Solution Approach 1:
The invention merges two previously separate training approaches: appearance-based photometric loss (which provides training efficiency and brightness consistency) and GPS-to-scale loss (which provides scale awareness). By combining these loss functions, the system achieves both training efficiency and scale awareness simultaneously.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method to improve scale consistency and/or scale awareness in a model of self-supervised depth and ego-motion prediction neural networks processing a video stream of monocular images, wherein complementary GPS coordinates synchronized with the images are used to calculate a GPS to scale loss to enforce the scale-consistency and/or -awareness on the monocular self-supervised ego-motion and depth estimation. A relative weight assigned to the GPS to scale loss exponentially increases as training progresses. The depth and ego-motion prediction neural networks are trained using an appearance-based photometric loss between real and synthesized target images, as well as a smoothness loss on the depth predictions.