Monocular SLAM Object Integration via Bundle Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional monocular SLAM systems face challenges in less textured environments, leading to feature tracking failures and pose estimation inaccuracies due to insufficient tracking of mapped points or motion-induced errors, especially in cases of abrupt motion, and existing edge-based methods suffer from drift due to inaccuracies in optical flow.

Innovation Solution

A method and system for integrating objects in monocular SLAM that performs bundle adjustment and objection detection using edge correspondences, triangulation techniques, and joint optimization to generate an optimized 3D map, incorporating object shape parameters and poses to improve mapping accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If conventional feature-based visual SLAM is used, then the system is lightweight and easy to implement, but it suffers from feature tracking failures in less textured environments and motion-induced errors

Engineering Contradiction:
Improvesystem complexityVSAvoidfeature tracking reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent merges conventional feature-based SLAM with object-based SLAM by integrating object detections (bounding boxes, key points, wireframe models) into the existing SLAM pipeline. This combination allows the system to leverage both feature points and object geometric constraints, improving tracking reliability in less textured environments while maintaining system lightweight characteristics

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent creates a composite approach by combining multiple representation types (corners, edges, object bounding boxes, key points, wireframe models) into a unified SLAM framework. This composite feature space provides redundancy and complementary information, enhancing system robustness against feature tracking failures without significantly increasing computational complexity

Inventive Principle:
Principle #40Composite materials

2Reliability

If edge-based methods are used for SLAM optimization, then the system can work in less textured environments, but it suffers from drift due to inaccuracies in optical flow

Engineering Contradiction:
ImproveSLAM performance in less textured environmentsVSAvoidpose estimation accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent introduces object geometric constraints (wireframe models, key points, bounding boxes) as intermediary elements that mediate between edge correspondences and pose estimation. These object-based features serve as stable reference points that reduce drift by providing additional geometric constraints that are less sensitive to optical flow inaccuracies

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from 2D edge correspondence matching to 3D object pose estimation by fitting wireframe models and detecting key points in three-dimensional space. This dimensional enhancement provides stronger geometric constraints for pose estimation, reducing drift while maintaining effectiveness in less textured environments

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If generic object models (ellipsoids, cuboids) are used in Object SLAM, then the system can identify object categories, but it provides limited information about object pose in the map

Engineering Contradiction:
Improveobject category recognitionVSAvoidobject pose information
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent segments the object representation into multiple components: generic shape models (ellipsoids, cuboids) for category identification, bounding boxes for spatial localization, key points for pose estimation, and wireframe models for detailed geometric structure. This segmentation allows the system to extract both category information and precise pose data from each component

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent enhances generic object models by adding three-dimensional wireframe representations and key point detections that provide explicit pose information. This dimensional enhancement transforms simple category labels into rich spatial representations that include position, orientation, and scale in 3D space

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

4Productivity

If conventional bundle adjustment is performed without object integration, then the processing is computationally efficient, but the accuracy and robustness of the 3D map is limited

Engineering Contradiction:
Improveprocessing speedVSAvoid3D map accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies partial object integration by selectively incorporating object detections (bounding boxes, key points, wireframe models) only for objects that are confidently detected and geometrically consistent with the scene. This partial integration improves 3D map accuracy without processing all possible object data, maintaining computational efficiency

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11670047B2System and method for integrating objects in monocular slam
Publication Date: 2023.06.06 TATA CONSULTANCY SERVICES LTD
  • US11670047B2 patent drawing
  • US11670047B2 patent drawing
  • US11670047B2 patent drawing

AI summary

The embodiments herein provide a system and method for integrating objects in monocular simultaneous localization and mapping (SLAM). State of art object SLAM approach use two popular threads. In first, instance specific models are assumed to be known a priori. In second, a general model for an object such as ellipsoids and cuboids is used. However, these generic models just give the label of the object category and do not give much information about the object pose in the map. The method and system disclosed provide a SLAM framework on a real monocular sequence wherein joint optimization is performed on object localization and edges using category level shape priors and bundle adjustment. The method provides a better visualization incorporating object representations in the scene along with the 3D structure of the base SLAM system, which makes it useful for augmented reality (AR) applications.