Virtual-Augmented Visual SLAM for Large-Baseline Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Visual Simultaneous Localization and Mapping (VSLAM) systems face challenges in constructing accurate 3D models of scenes with a reduced number of images, particularly due to the large-baseline matching problem where features observed from different viewpoints change significantly, leading to incorrect pose localization and increased computational and memory requirements.

Innovation Solution

The implementation of virtually-augmented visual SLAM (VA-VSLAM) generates virtual images from real images captured at different viewpoints, allowing for comparison and registration of features across varying appearances, thereby alleviating the large-baseline matching problem and reducing the number of images needed for 3D model construction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If images are acquired from distant viewpoints to reduce the number of images, then the number of images required for 3D model construction is reduced, but the large-baseline matching problem occurs where features change significantly between viewpoints

Engineering Contradiction:
Improvenumber of imagesVSAvoidfeature matching accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent introduces virtual images as an intermediary between real images captured at distant viewpoints. These virtual images are synthesized to bridge the large baseline gap, providing intermediate visual information that facilitates reliable feature matching while allowing the use of fewer real images for 3D reconstruction.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates virtual copies of real images through image synthesis techniques. These copied virtual images represent viewpoints that were not directly captured, enabling the system to infer scene geometry from fewer real images while maintaining matching reliability through the synthesized intermediates.

Inventive Principle:
Principle #26Copying

2Measurement precision

If multiple images with small pose variation are used for continuous pose tracking, then pose estimation accuracy is improved, but computational cost and memory requirements increase

Engineering Contradiction:
Improvepose estimation accuracyVSAvoidcomputational power
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent applies partial action by using virtual images selectively only where and when they are needed to bridge large baselines, rather than continuously processing all possible image pairs. This reduces computational overhead while maintaining pose estimation accuracy for the critical large-baseline cases.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If motion sensors are combined with VSLAM for pose estimation, then pose tracking reliability is improved, but device complexity increases

Engineering Contradiction:
Improvepose tracking reliabilityVSAvoidsensor system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical sensor system (motion sensors) with a computational image processing system that synthesizes virtual images. This substitution maintains pose tracking reliability through visual information while eliminating the need for additional physical sensors, thereby reducing device complexity.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS10659768B2System and method for virtually-augmented visual simultaneous localization and mapping
Publication Date: 2020.05.19 MITSUBISHI ELECTRIC RESEARCH LABORATORIES INC
  • US10659768B2 patent drawing
  • US10659768B2 patent drawing
  • US10659768B2 patent drawing

AI summary

A system for reconstructing a three-dimensional (3D) model of a scene including a point cloud having points identified by 3D coordinates includes at least one sensor to acquire a set of images of the scene from different poses defining viewpoints of the images and a memory to store the set of images and the 3D model of the scene. The system also includes a processor operatively connected to the memory and coupled with stored instructions to transform the images from the set of images to produce a set of virtual images of the scene viewed from virtual viewpoints; compare at least some features from the images and the virtual images to determine the viewpoint of each image in the set of images; and update 3D coordinates of at least one point in the model of the scene to match coordinates of intersections of ray back-projections from pixels of at least two images corresponding to the point according to the viewpoints of the two images.