Limited-Motion Monocular Video Depth Estimation with Two-Phase Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Accurately determining the depth of objects in videos captured by a monocular camera with minimal camera motion is challenging due to limited parallax, leading to inaccurate results from structure-from-motion techniques.
Innovation Solution
A two-phase training method for neural networks that adjusts camera parameters and network weights, utilizing reprojection and depth prior losses, to estimate camera poses, dense depth maps, and movement maps without requiring known camera poses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If structure-from-motion processing techniques are used to determine depth in monocular video, then depth estimation can be achieved, but the accuracy deteriorates when camera motion is minimal due to limited parallax
Solution Approach 1:
The patent changes the optimization parameters from traditional structure-from-motion approaches to a neural network-based approach that optimizes both camera parameters and depth maps simultaneously. The system uses reprojection loss and depth prior loss to guide the optimization, enabling accurate depth estimation even when camera motion is minimal by leveraging learned patterns rather than relying solely on geometric parallax
Solution Approach 2:
The patent replaces traditional mechanical structure-from-motion algorithms with a neural network-based system. Instead of relying on geometric computations and image processing algorithms, the system uses a trained neural network that processes video frames to directly estimate depth, camera pose, and motion parameters, achieving better accuracy in scenarios with limited camera motion
2Reliability
If neural network training adjusts both camera parameters and network weights simultaneously, then the model can learn from data, but the training complexity and computational cost increase
Solution Approach 1:
The patent segments the training process into two distinct phases: a first phase that optimizes only camera parameters while keeping network weights fixed, and a second phase that optimizes both camera parameters and network weights. This segmentation simplifies the training process by breaking down the complex optimization into manageable steps, improving convergence reliability while reducing computational complexity compared to simultaneous optimization of all parameters
Data Source
AI summary
Methods, systems, and apparatus, including medium-encoded computer program products, for determining video characteristics from videos captured with limited motion cameras. A first set of pairs of images can be selected from a video taken with a limited motion camera. Using the images, a neural network can first be trained for camera parameters while holding the network weights constant. After performing the first training, a second set of pairs of images can be selected from the video. A second training of the neural network can be performed and can include adjusting the camera parameters and the network weights in the neural network. After performing the second training, the camera parameters and the network weights of the neural network can be persisted.


