Multi-view Depth Estimation Using Offline Structure-from-Motion Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing systems for generating 3D representations of environments by autonomous agents, such as vehicles, face reduced accuracy due to the inclusion of images lacking depth information, which affects tasks like scene understanding and obstacle avoidance.

Innovation Solution

A method for estimating depth using a combination of offline structure-from-motion and multi-view depth estimation, where images captured by different agents are filtered based on depth criteria to select the most informative frames for generating a dense map, improving the accuracy of the 3D representation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If all captured images are used for depth estimation, then more data is available for 3D representation, but images lacking depth information reduce the accuracy of the 3D representation

Engineering Contradiction:
Improvenumber of images usedVSAvoiddepth estimation accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent extracts and removes images that lack sufficient depth information from the set of captured images. By identifying and excluding these low-quality images through filtering based on depth criteria, the system ensures that only images with adequate depth information are used for depth estimation, thereby resolving the contradiction between using more images and maintaining accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different quality standards to different images based on their individual depth information quality. Instead of treating all images uniformly, the system evaluates each image's depth information quality and selectively includes or excludes them, ensuring that high-quality images contribute to the 3D representation while low-quality images are filtered out.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If offline structure-from-motion is used to pre-process images, then depth information quality improves, but system complexity increases

Engineering Contradiction:
Improvedepth information qualityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs structure-from-motion processing in advance during an offline phase, before the actual depth estimation task. By pre-processing the images to extract depth information and pre-filtering based on depth quality criteria, the system reduces the computational burden during real-time operation while maintaining high depth estimation accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses its own captured images to perform the structure-from-motion processing, creating a self-contained preprocessing pipeline. The captured images are reused and processed to generate depth information that then serves to filter and improve the same set of images, making the system self-sufficient without requiring external data or complex additional hardware.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12080013B2Multi-view depth estimation leveraging offline structure-from-motion
Publication Date: 2024.09.03 TOYOTA JIDOSHA KK
  • US12080013B2 patent drawing
  • US12080013B2 patent drawing
  • US12080013B2 patent drawing

AI summary

A method for estimating depth of a scene includes selecting an image of the scene from a sequence of images of the scene captured via an in-vehicle sensor of a first agent. The method also includes identifying previously captured images of the scene. The method further includes selecting a set of images from the previously captured images based on each image of the set of images satisfying depth criteria. The method still further includes estimating the depth of the scene based on the selected image and the selected set of images.