AR Depth Estimation Using Odometry and Hand Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional depth estimation and object tracking systems in augmented reality environments face challenges such as scale accuracy, occlusion handling, reliance on surface textures, high computational demands, adaptability to environmental changes, and sensitivity to lighting conditions, leading to inaccurate and resource-intensive operations.

Innovation Solution

An interaction system utilizing advanced neural networks for metric depth estimation, intelligent occlusion handling, optimized algorithms, and adaptive depth sensing to ensure accurate and efficient tracking, even in dynamic environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional depth estimation algorithms are used, then the system can operate with simpler hardware, but depth accuracy and scale precision deteriorate

Engineering Contradiction:
Improvedepth accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces hand tracking as an intermediary element that mediates between the camera system and depth estimation algorithms. By tracking hand keypoints and using them as reference points, the system achieves accurate metric depth estimation without requiring complex full-scene depth algorithms, thus improving depth accuracy while managing computational complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system changes the parameter of reference from general scene features to specific hand keypoints. This parameter change enables the depth estimation to be anchored to known physical dimensions of human hands, providing metric scale accuracy without requiring complex environmental understanding or heavy computational resources.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If traditional object tracking systems are used, then the system can handle simple environments, but reliability deteriorates in dynamic environments with occlusions and lighting changes

Engineering Contradiction:
Improvetracking reliabilityVSAvoidenvironmental adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system dynamically adapts to environmental changes by continuously tracking hand movements and adjusting depth estimation in real-time. The hand tracking system remains reliable under occlusions by predicting hand position based on motion patterns, and adapts to lighting changes by focusing on invariant hand features, thus improving tracking reliability while maintaining environmental adaptability.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If computationally intensive depth estimation algorithms are used, then depth precision improves, but processing speed and energy consumption worsen

Engineering Contradiction:
Improvemetric depth precisionVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent extracts the essential information needed for depth estimation by focusing specifically on hand keypoints rather than processing the entire scene. This extraction approach achieves metric depth precision by measuring distances from camera to hand keypoints, while significantly reducing processing speed requirements and energy consumption compared to full-scene depth mapping algorithms.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20260080633A1Depth estimation using odometry and hand tracking
Publication Date: 2026.03.19 SNAP INC
  • US20260080633A1 patent drawing
  • US20260080633A1 patent drawing
  • US20260080633A1 patent drawing

AI summary

A head-worn augmented reality (AR) device system includes cameras, display devices, and processors, along with a memory that stores specific instructions. When these instructions are executed by the processors, they enable the device to perform several operations. First, the device accesses a two-dimensional (2D) camera image taken by its camera. The device then generates a first set of three-dimensional (3D) tracked points using the device's odometry system applied to this 2D image. Optionally, a second set of tracked 3D points is created based on one or more images captured by the camera. These 3D points are projected onto the 2D camera image to create a sparse depth image. Finally, this 2D camera image, along with the newly formed depth image, is fed into a first machine learning model to generate a metric depth estimation.