Gaze-Tracked Robot Interface for Faster Workspace Information Access
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In robot systems, users face inefficiencies when accessing desired information through hierarchical selection menus on teach pendants, requiring sequential selection processes that can be time-consuming and cumbersome.
Innovation Solution
A robot system incorporating a robot controller, video acquisition device, and head-mounted video display with visual line tracking, which identifies the user's gaze target and displays associated information alongside the target in a single image, allowing users to access information by simply looking at the desired object or feature.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If hierarchical selection menus are used on teach pendant display section, then information can be organized systematically, but user operation time and complexity increase
Solution Approach 1:
The patent transitions from traditional 2D hierarchical menus on a teach pendant screen to a 3D spatial interface where information is overlaid directly onto the physical workspace. The head-mounted display presents virtual information panels that float in three-dimensional space, allowing users to access information by looking at different spatial locations rather than navigating through menu hierarchies.
Solution Approach 2:
The patent introduces a camera as an intermediary device that captures the real workspace and overlays virtual information panels onto the captured image. This intermediary layer bridges the physical workspace and the digital information system, allowing users to see both the actual environment and relevant information simultaneously through the head-mounted display.
2Loss of information
If hierarchical selection menus are used on teach pendant, then information can be structured, but ease of operation deteriorates
Solution Approach 1:
The patent transitions from traditional 2D hierarchical menus on a teach pendant screen to a 3D spatial interface where information is overlaid directly onto the physical workspace. The head-mounted display presents virtual information panels that float in three-dimensional space, allowing users to access information by looking at different spatial locations rather than navigating through menu hierarchies.
Solution Approach 2:
The system automatically detects which information panel the user is looking at through camera tracking and visual line detection, and automatically presents the relevant information without requiring the user to manually navigate menus or make selections. The interface serves itself by responding to the user's gaze direction.
3Productivity
If visual line tracking and gaze-based selection are implemented, then operation speed improves, but device complexity increases
Solution Approach 1:
The head-mounted display device performs multiple functions: it captures video of the workspace, tracks the user's visual line, identifies gaze targets, and displays overlaid information panels. By consolidating these functions into a single wearable device, the system reduces the need for multiple separate components (camera, tracker, display) that would otherwise be distributed throughout the workspace.
Solution Approach 2:
The patent replaces manual mechanical operations (pressing buttons, navigating menus with hands) with optical and computational methods (visual line tracking, gaze detection, automated information presentation). The mechanical teach pendant interface is substituted with an optical gaze-based interface through the head-mounted display.
Data Source
AI summary
A robot system includes a robot, a robot controller, a video acquisition device configured to acquire a real video of a work space, and a head-mounted type video display device provided with a visual line tracking section configured to acquire visual line information. A robot controller includes an information storage section configured to store information used for controlling the robot while associating the information with a type of an object, a gaze target identification section configured to identify, in the video, a gaze target viewed by a wearer based on the visual line information, and a display processing section configured to cause the video display device to display the information associated with the object corresponding to the identified gaze target, side by side with the gaze target in the form of one image through which the wearer can visually grasp, select, or set contents of the information.


