Geometric Semantic Robot Navigation for RGB-Only Object Search

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional robot navigation systems struggle with complex real-world scenarios due to the requirement of depth perception and significant policy training, limiting their ability to effectively navigate and detect target objects in unknown environments using only egocentric RGB camera perception.

Innovation Solution

A geometric semantic map-based navigation system that utilizes a hybrid map combining geometric and semantic information, along with ontological knowledge representation, to enable a robot to navigate and detect target objects in indoor environments by iteratively updating its knowledge store and map based on egocentric images, incorporating actuation capabilities and scene relations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional geometric reconstruction and path-planning approaches are used, then navigation can be achieved, but depth perception and significant policy training are required, making deployment difficult

Engineering Contradiction:
Improvenavigation capabilityVSAvoiddepth perception and policy training requirements
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts and removes the requirement for depth perception from the navigation system. By using only 2D RGB camera inputs and processing images through semantic segmentation and feature extraction algorithms, the system achieves navigation without requiring depth sensors or complex 3D reconstruction, thereby eliminating this complexity barrier while maintaining navigation reliability

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces traditional geometric reconstruction methods with a semantic-based approach using deep learning models. Instead of relying on complex geometric algorithms and extensive policy training, the system uses semantic segmentation networks to understand scene content and navigate based on semantic features, significantly reducing the training requirements and deployment complexity

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Loss of information

If semantic visual navigation is used to achieve complex navigation goals, then richer scene understanding is obtained, but depth perception and policy training are still required

Engineering Contradiction:
Improvescene understanding capabilityVSAvoiddepth perception and policy training requirements
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the navigation task into distinct components: semantic segmentation of the scene into meaningful regions, extraction of key features from segmented regions, and decision-making based on these features. This segmentation allows the system to achieve rich scene understanding through 2D image processing alone, without requiring depth perception or extensive policy training for each component

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If complete learning-based approaches are used, then navigation in unknown environments can be achieved, but significant policy training is required

Engineering Contradiction:
Improvenavigation in unknown environmentsVSAvoidpolicy training time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-training semantic segmentation models on large datasets before deployment. Once the semantic understanding capabilities are established through preliminary training, the system can adapt to unknown environments using these pre-learned semantic features without requiring extensive additional policy training, thereby reducing the time loss while maintaining adaptability

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12393197B2Systems and methods for object detection using a geometric semantic map based robot navigation
Publication Date: 2025.08.19 TATA CONSULTANCY SERVICES LTD
  • US12393197B2 patent drawing
  • US12393197B2 patent drawing
  • US12393197B2 patent drawing

AI summary

This disclosure relates generally to systems and methods for object detection using a geometric semantic map based robot navigation using an architecture to empower a robot to navigate an indoor environment with logical decision making at each intermediate stage. The decision making is further enhanced by knowledge on actuation capability of the robots and that of scenes, objects and their relations maintained in an ontological form. The robot navigates based on a Geometric Semantic map which is a relational combination of geometric and semantic map. In comparison to traditional approaches, the robot's primary task here is not to map the environment, but to reach a target object. Thus, a goal given to the robot is to find an object in an unknown environment with no navigational map and only egocentric RGB camera perception.