Robot Gaze Mapping for Off-Camera Object Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for identifying an object gazed at by a user in an image captured by a robot fail when the object is not included within the image's viewing angle.
Innovation Solution
A robot system that utilizes a camera, processor, and artificial intelligence model to identify an object gazed at by a user by mapping user position and gaze direction onto an environment map, even when the object is outside the camera's viewing angle, through simultaneous localization and mapping (SLAM) and a trained AI model.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the robot uses traditional gaze identification methods based on head pose or eye gaze direction in captured images, then the identification process is simple, but the object cannot be identified when it is not included in the image
Solution Approach 1:
The patent merges multiple information sources including head pose estimation, eye gaze direction analysis, SLAM-based environment mapping, and AI object recognition models into a unified gaze-based object identification system. This integration allows the robot to identify objects even when they are not directly visible in the camera frame by combining visual cues from the user's gaze with spatial information from the environment map.
Solution Approach 2:
The patent introduces an environment map generated through SLAM as an intermediary data structure that bridges the gap between the user's gaze direction and objects not visible in the current image. The map serves as a mediator that allows the system to infer object locations based on gaze direction alone, without requiring the object to be present in the captured image.
2Adaptability or versatility
If the robot integrates SLAM and AI model to identify objects beyond camera view, then object identification capability is improved, but the system complexity increases
Solution Approach 1:
The patent performs preliminary actions by pre-training an AI object recognition model with extensive object data before deployment. The environment map is also pre-generated through SLAM as the robot explores the space. These preliminary preparations enable the system to quickly and accurately identify gazed objects without requiring complex real-time processing when the actual identification is needed.
Solution Approach 2:
The patent transitions from two-dimensional image-based object identification to three-dimensional spatial awareness by integrating SLAM-generated environment maps. This dimensional expansion allows the system to identify objects in 3D space based on gaze direction alone, even when those objects are not visible in the 2D camera image, thereby enhancing adaptability without proportionally increasing complexity.
3Measurement precision
If the robot uses only image-based gaze analysis, then the processing speed is fast, but the object identification fails when object is outside viewing angle
Solution Approach 1:
The patent creates a universal gaze interpretation system that functions both when objects are visible in the image and when they are not. The same head pose and eye gaze analysis algorithms serve dual purposes: identifying objects within the camera frame and inferring object locations beyond the frame using the environment map, thereby expanding identification scope without sacrificing existing functionality.
Data Source
AI summary
A robot and a method for controlling a robot are provided. The method includes: acquiring an image of a user; acquiring, by analyzing the image, a first information regarding a position of the user and a gaze direction of the user; acquiring, based on an image capturing position associated with the image and an image capturing direction associated with the image, matching information for matching the first information with a map corresponding to an environment in which the robot is operated; acquiring, based on the matching information and the first information, second information regarding the position of the user on the map and the gaze direction of the user on the map; and identifying an object corresponding to the gaze direction of the user on the map by inputting the second information into an artificial intelligence model trained to identify an object on the map.


