Human-Robot Cognition Sharing for Gesture-Guided Indoor Navigation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional human-robot interaction methods using gestures and speech are limited in navigating indoor environments with non-discrete directions and face challenges in occluded perspectives, requiring additional hardware and being inefficient in transferring cognitive load from humans to robots.
Innovation Solution
A processor-implemented method and system that acquires visual and audio feeds to estimate a pointing direction, generates a trajectory using a 2-D occupancy map and ROS movebase planner, and employs a zero-shot single-stage network for language grounding to navigate robots to a final goal point, enabling effective cognition sharing between humans and robots.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional techniques use gestures and speech for human robot interaction, then communication between humans and robots is improved, but the system is limited to structured outdoor driving environments or discrete set of directions in indoor environments
Solution Approach 1:
The system changes the parameter of direction representation from discrete predefined directions to continuous angular measurements. The processor determines angular positions of gestures relative to the robot's current orientation and generates corresponding continuous navigation commands, enabling adaptation to any direction in indoor environments rather than being limited to discrete direction sets.
2Measurement precision
If additional hardware sensors are added to improve gesture and speech recognition accuracy, then measurement precision is improved, but device complexity and cost increase
Solution Approach 1:
The system makes the existing camera multi-functional by using it not only for visual scene capture but also for gesture recognition and directional estimation. The processor analyzes video feed from the camera to detect hand gestures, determine their angular positions, and translate them into navigation commands, eliminating the need for separate gesture sensors while improving measurement precision through software-based analysis.
3Measurement precision
If the robot waits for complete directional instructions before moving, then navigation accuracy is improved, but productivity and response time deteriorate
Solution Approach 1:
The system performs preliminary navigation actions by generating and executing trajectory commands as soon as a gesture is detected, without waiting for complete verbal instructions. The processor determines the angular position of the gesture, generates a trajectory to the corresponding location, and initiates movement immediately, allowing the robot to start navigating before the human speaker finishes providing all directional information.
Solution Approach 2:
The system implements feedback by continuously monitoring the robot's navigation progress and comparing it with the intended destination derived from gestures. The processor can adjust the trajectory in real-time based on the robot's current position, obstacles detected by sensors, and any additional gestures or verbal corrections provided by the human, thereby maintaining navigation accuracy while enabling continuous movement.
Data Source
AI summary
The disclosure generally relates to methods and systems for enabling human robot interaction by cognition sharing which includes gesture and audio. Conventional techniques that use the gestures and the speech, require extra hardware setup and are limited to navigation in structured outdoor driving environments. The present disclosure herein provides methods and systems that solves the technical problem of enabling the human robot interaction with a two-step approach by transferring the cognitive load from the human to the robot. An accurate shared perspective associated with the task is determined in the first step by computing relative frame transformations based on understanding of navigational gestures of the subject. Then, the shared perspective transformed to the robot in the field view of the robot. The transformed shared perspective is then given to a language grounding technique in the second step, to accurately determine a final goal associated with the task.


