Shared Cognition for Human-Robot Navigation With Occluded Targets
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional human-robot interaction systems using gestures and speech are limited in navigating indoor environments with non-discrete directions and face challenges in determining the final goal, especially when the target region is occluded, due to the uncorrelated nature of raw input gestures and speech, and require additional hardware sensors.
Innovation Solution
A processor-implemented method and system that captures shared cognition by determining an accurate intermediate goal pose through relative frame transformations based on navigational gestures, transforming it into the robot's coordinate frame, and using language grounding techniques with natural language instructions to determine the final goal, without the need for dedicated sensors, allowing the robot to navigate effectively in indoor environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional techniques use various combinations of hardware sensor modalities for human robot interaction, then communication and interaction between humans and robots is improved, but device complexity increases and additional hardware sensors are required
Solution Approach 1:
The patent applies multi-functionality by enabling a monocular camera to perform multiple functions: capturing visual information for gesture recognition, providing spatial context for language grounding, and serving as the primary sensing modality without requiring additional dedicated sensors. This consolidates multiple sensing functions into a single universal device, reducing hardware complexity while maintaining interaction reliability
Solution Approach 2:
The system applies self-service by using the robot's existing monocular camera for cognition sharing tasks rather than relying on external or additional sensors. The camera serves the dual purpose of navigation and gesture recognition, allowing the robot to leverage its own hardware resources for human-robot interaction without requiring external assistance or additional components
2Reliability
If conventional techniques are used for navigation in indoor environments with discrete directions, then navigation in structured environments is improved, but adaptability to non-discrete directions and occluded targets deteriorates
Solution Approach 1:
The patent applies preliminary action by capturing an intermediate goal image at a predicted location before the robot reaches the final destination. This intermediate visual capture allows the system to verify and adjust the navigation path in advance, enabling the robot to adapt to non-discrete directions and handle occluded targets by having visual information available before completing the navigation task
Solution Approach 2:
The patent introduces an intermediate goal image as a mediator between the start position and final destination. This intermediate visual representation serves as a reference point that helps the robot adapt to complex indoor environments with non-discrete directions and occlusions, bridging the gap between discrete navigation commands and continuous environmental adaptation
3Ease of operation
If gestures and speech are processed as uncorrelated raw inputs, then input processing simplicity is improved, but cognition sharing accuracy and task performance deteriorate
Solution Approach 1:
The patent applies merging by integrating gesture recognition and language grounding into a unified cognition sharing framework. Instead of processing gestures and speech as separate uncorrelated inputs, the system combines them by using the intermediate goal image from gestures to ground the language instruction, thereby improving cognition sharing accuracy while maintaining processing simplicity through a cohesive integrated approach
Data Source
Figure 1
Figure 2
Figure 3A
AI summary
The disclosure generally relates to methods and systems for enabling human robot interaction by cognition sharing which includes gesture and audio. Conventional techniques that use the gestures and the speech, require extra hardware setup and are limited to navigation in structured outdoor driving environments. The present disclosure herein provides methods and systems that solves the technical problem of enabling the human robot interaction with a two-step approach by transferring the cognitive load from the human to the robot. An accurate shared perspective associated with the task is determined in the first step by computing relative frame transformations based on understanding of navigational gestures of the subject. Then, the shared perspective transformed to the robot in the field view of the robot. The transformed shared perspective is then given to a language grounding technique in the second step, to accurately determine a final goal associated with the task.