Monocular Localization via Scale-Aware Neural Depth Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing digital image processors face challenges in accurately determining the scale and depth of objects in images captured by a single camera, leading to scale ambiguity, which is exacerbated by the need for significant processing power and additional hardware in current methods.
Innovation Solution
A neural network-based image processing device and method that estimates the scale of image features by processing multiple images using trained models to identify features, estimate depths, and adjust scaling, allowing for accurate depth and scale inference without additional hardware or knowledge of scene features, using a combination of stereo-based depth and scale estimation techniques.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If monocular vision is used to capture images, then device complexity is reduced, but measurement precision of depth and scale deteriorates
Solution Approach 1:
The system performs preliminary actions by capturing multiple images from different spatial locations before depth estimation. This multi-image capture approach provides sufficient geometric information to resolve scale ambiguity, enabling accurate depth and scale measurement with a simple monocular camera system.
Solution Approach 2:
The patent introduces an intermediary computational process that uses spatial relationships and parallax information from multiple images to infer depth and scale. This intermediary processing step bridges the gap between simple monocular capture and accurate metric measurement, resolving the contradiction without adding hardware complexity.
2Measurement precision
If statistics-based algorithms are used to determine scale, then measurement precision improves, but use of energy increases
Solution Approach 1:
The system extracts only the essential geometric information needed for scale determination from multiple images, rather than performing exhaustive statistical analysis. By focusing on key spatial relationships and parallax effects, the system achieves accurate scale measurement with reduced computational energy consumption.
Solution Approach 2:
The patent applies partial action by using a selective subset of image data and processing steps sufficient for scale determination, rather than applying complete statistical algorithms to all available data. This approach achieves the necessary measurement precision while consuming less processing energy.
3Measurement precision
If additional sensors are added to resolve scale ambiguity, then measurement precision improves, but device complexity increases
Solution Approach 1:
The monocular camera system serves itself by using its own multiple captures from different positions to resolve its inherent scale ambiguity. The system leverages the spatial-temporal information already available from the camera's movement, eliminating the need for additional depth sensors or auxiliary hardware.
Solution Approach 2:
The patent makes the monocular camera system multi-functional by enabling it to perform both image capture and depth/scale measurement using only the camera itself. Through computational processing of multiple images, the system achieves functions traditionally requiring additional sensors, thereby maintaining device simplicity while improving measurement precision.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach enables accurate reconstruction of objects and determination of their distance from the camera, reducing processing requirements and eliminating the need for extra sensors, while providing metrically correct virtual content placement.
Implementation Method 1
the principle of parallax can be used i.e. that the same object appears differently-sized depending how far away it is from the camera. Thus if an image is acquired from two (or more) different spatial locations, points that are seen in both images at different pixel locations can be triangulated.
Implementation Method 2
points that are seen in both images at different pixel locations can be triangulated
Implementation Method 3
estimating the scale of image features by the operations of: processing multiple images of a scene using a first trained model to identify features in the images and to estimate the depths of those features in the images; processing the multiple images by a second trained model to estimate a scaling for the images; and estimating the scales of the features by adjusting the estimated depths in dependence on the estimated scaling.
Data Source
AI summary
Disclosed is an image processing device comprising a processor configured to estimate the scale of image features by the steps of: processing multiple images of a scene by means of a first trained model to identify features in the images and to estimate the depths of those features in the images; processing the multiple images by a second trained model to estimate a scaling for the images; and estimating the scales of the features by adjusting the estimated depths in dependence on the estimated scaling. A method for training an image processing model is also disclosed.


