Limited-Motion Monocular Video Depth Estimation with Two-Phase Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Accurately determining the depth of objects in videos captured by a monocular camera with minimal camera motion is challenging due to limited parallax, leading to inaccurate results from structure-from-motion techniques.

Innovation Solution

A two-phase training method for neural networks that adjusts camera parameters and network weights, utilizing reprojection and depth prior losses, to estimate camera poses, dense depth maps, and movement maps without requiring known camera poses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If structure-from-motion processing techniques are used to determine depth in monocular video, then depth estimation can be achieved, but the accuracy deteriorates when camera motion is minimal due to limited parallax

Engineering Contradiction:
Improvedepth estimation accuracyVSAvoidperformance under limited camera motion
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent changes the optimization parameters from traditional structure-from-motion approaches to a neural network-based approach that optimizes both camera parameters and depth maps simultaneously. The system uses reprojection loss and depth prior loss to guide the optimization, enabling accurate depth estimation even when camera motion is minimal by leveraging learned patterns rather than relying solely on geometric parallax

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional mechanical structure-from-motion algorithms with a neural network-based system. Instead of relying on geometric computations and image processing algorithms, the system uses a trained neural network that processes video frames to directly estimate depth, camera pose, and motion parameters, achieving better accuracy in scenarios with limited camera motion

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If neural network training adjusts both camera parameters and network weights simultaneously, then the model can learn from data, but the training complexity and computational cost increase

Engineering Contradiction:
Improvemodel training convergenceVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the training process into two distinct phases: a first phase that optimizes only camera parameters while keeping network weights fixed, and a second phase that optimizes both camera parameters and network weights. This segmentation simplifies the training process by breaking down the complex optimization into manageable steps, improving convergence reliability while reducing computational complexity compared to simultaneous optimization of all parameters

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS20250259323A1Video characteristic determination from videos captured with limited motion camera
Publication Date: 2025.08.14 GOOGLE LLC
  • US20250259323A1 patent drawing
  • US20250259323A1 patent drawing
  • US20250259323A1 patent drawing

AI summary

Methods, systems, and apparatus, including medium-encoded computer program products, for determining video characteristics from videos captured with limited motion cameras. A first set of pairs of images can be selected from a video taken with a limited motion camera. Using the images, a neural network can first be trained for camera parameters while holding the network weights constant. After performing the first training, a second set of pairs of images can be selected from the video. A second training of the neural network can be performed and can include adjusting the camera parameters and the network weights in the neural network. After performing the second training, the camera parameters and the network weights of the neural network can be persisted.