Shared Cognition for Human-Robot Navigation With Occluded Targets

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional human-robot interaction systems using gestures and speech are limited in navigating indoor environments with non-discrete directions and face challenges in determining the final goal, especially when the target region is occluded, due to the uncorrelated nature of raw input gestures and speech, and require additional hardware sensors.

Innovation Solution

A processor-implemented method and system that captures shared cognition by determining an accurate intermediate goal pose through relative frame transformations based on navigational gestures, transforming it into the robot's coordinate frame, and using language grounding techniques with natural language instructions to determine the final goal, without the need for dedicated sensors, allowing the robot to navigate effectively in indoor environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional techniques use various combinations of hardware sensor modalities for human robot interaction, then communication and interaction between humans and robots is improved, but device complexity increases and additional hardware sensors are required

Engineering Contradiction:
Improvecommunication reliabilityVSAvoidhardware complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies multi-functionality by enabling a monocular camera to perform multiple functions: capturing visual information for gesture recognition, providing spatial context for language grounding, and serving as the primary sensing modality without requiring additional dedicated sensors. This consolidates multiple sensing functions into a single universal device, reducing hardware complexity while maintaining interaction reliability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system applies self-service by using the robot's existing monocular camera for cognition sharing tasks rather than relying on external or additional sensors. The camera serves the dual purpose of navigation and gesture recognition, allowing the robot to leverage its own hardware resources for human-robot interaction without requiring external assistance or additional components

Inventive Principle:
Principle #25Self-service

2Reliability

If conventional techniques are used for navigation in indoor environments with discrete directions, then navigation in structured environments is improved, but adaptability to non-discrete directions and occluded targets deteriorates

Engineering Contradiction:
Improvenavigation reliabilityVSAvoidenvironmental adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by capturing an intermediate goal image at a predicted location before the robot reaches the final destination. This intermediate visual capture allows the system to verify and adjust the navigation path in advance, enabling the robot to adapt to non-discrete directions and handle occluded targets by having visual information available before completing the navigation task

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediate goal image as a mediator between the start position and final destination. This intermediate visual representation serves as a reference point that helps the robot adapt to complex indoor environments with non-discrete directions and occlusions, bridging the gap between discrete navigation commands and continuous environmental adaptation

Inventive Principle:
Principle #24Intermediary (Mediator)

3Ease of operation

If gestures and speech are processed as uncorrelated raw inputs, then input processing simplicity is improved, but cognition sharing accuracy and task performance deteriorate

Engineering Contradiction:
Improveinput processing simplicityVSAvoidcognition sharing accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent applies merging by integrating gesture recognition and language grounding into a unified cognition sharing framework. Instead of processing gestures and speech as separate uncorrelated inputs, the system combines them by using the intermediate goal image from gestures to ground the language instruction, thereby improving cognition sharing accuracy while maintaining processing simplicity through a cohesive integrated approach

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP3916507B1Methods and systems for enabling human robot interaction by sharing cognition
Publication Date: 2023.05.24 TATA CONSULTANCY SERVICES LTD
  • EP3916507B1 patent drawingFigure 1
  • EP3916507B1 patent drawingFigure 2
  • EP3916507B1 patent drawingFigure 3A

AI summary

The disclosure generally relates to methods and systems for enabling human robot interaction by cognition sharing which includes gesture and audio. Conventional techniques that use the gestures and the speech, require extra hardware setup and are limited to navigation in structured outdoor driving environments. The present disclosure herein provides methods and systems that solves the technical problem of enabling the human robot interaction with a two-step approach by transferring the cognitive load from the human to the robot. An accurate shared perspective associated with the task is determined in the first step by computing relative frame transformations based on understanding of navigational gestures of the subject. Then, the shared perspective transformed to the robot in the field view of the robot. The transformed shared perspective is then given to a language grounding technique in the second step, to accurately determine a final goal associated with the task.