Gaze Annotation System for Wearable Devices
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems for distributed collaboration lack a precise and natural mechanism for annotating images, especially in wearable computing devices, as they often capture peripheral objects and have limited input methods for adding markers, leading to potential misinterpretations during communication.
Innovation Solution
A gaze annotation system that uses deictic gaze gestures, where users look at points of interest and speak annotations, with the system recording gaze positions and speech to create visual anchors, allowing remote users to access and view these annotations, utilizing a head-mounted device with a scene camera, display, eye tracking, and automatic speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Area of stationary object
If a front-facing camera is used to capture the user's field of view, then the captured images include peripheral objects, but this causes distraction from the main message and reduces communication precision
Solution Approach 1:
The system extracts only the relevant portion of the field of view by using gaze tracking to identify the user's point of regard. The annotation is then placed precisely at this location, effectively extracting the important information from the peripheral clutter and presenting it in a focused manner to remote collaborators.
Solution Approach 2:
The system applies local quality by making the annotation placement location-specific based on gaze data. Instead of uniform annotation placement, the system adapts the annotation position to match the user's visual attention, ensuring that annotations appear exactly where the user is looking at the scene or in the captured image.
2Ease of manufacture
If conventional input methods are used for adding annotations, then the system has basic annotation capability, but the limited input methods make precise marker placement difficult
Solution Approach 1:
The system replaces manual mechanical input methods (buttons, touchscreens, keyboards) with a gaze-based pointing mechanism. The eye tracking camera detects the user's point of regard, and this gaze position directly determines the annotation placement location, eliminating the need for precise manual positioning while maintaining high precision.
Solution Approach 2:
The system uses the user's own natural gaze behavior as the input mechanism. Instead of requiring the user to learn and operate complex input devices, the system leverages the user's existing visual attention patterns, making precise annotation placement as natural as looking at the object of interest.
3Ease of operation
If manual annotation methods are used in wearable devices, then the system has annotation functionality, but the process becomes intrusive and complex
Solution Approach 1:
The system makes the wearable device self-serve the annotation process by automatically tracking the user's gaze and determining annotation placement without requiring manual intervention. The eye tracking camera continuously monitors the user's eye position, and the system automatically translates this into precise annotation locations, making the process as simple as looking and speaking.
Solution Approach 2:
The system replaces complex mechanical input operations with a gaze-based control mechanism. Instead of requiring the user to manipulate buttons, touchscreens, or other physical interfaces, the system uses the eye tracking camera to detect gaze position and automatically places annotations, significantly reducing operational complexity.
4Area of stationary object
If wide-angle camera views are used to match the user's field of view, then the captured images show the complete scene, but peripheral objects distract from the main message
Solution Approach 1:
The system extracts the relevant information from the wide-angle scene by using gaze tracking to identify the user's point of regard. Annotations are placed precisely at this location, effectively extracting the important message from the peripheral clutter and presenting it in a focused manner to remote collaborators.
Solution Approach 2:
The system applies local quality by making the annotation placement location-specific based on gaze data. The annotation is placed exactly where the user is looking, ensuring that the message is delivered with clarity and precision, while the surrounding peripheral objects remain in the background without distracting from the main communication.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Facilitates precise and non-intrusive image annotation, enhancing communication by allowing users to point out specific areas in images or real-world scenes, reducing misinterpretations and improving the accuracy of shared information.
Implementation Method 1
an eye tracking camera to track the users' gaze... From the gaze data received from the eye-tracking camera, a point-of-regard is calculated
Implementation Method 2
a front-facing scene camera to record video and photos
Implementation Method 3
a display to show captured photos or recorded videos
Data Source
AI summary
A gaze annotation method for an image includes: receiving a user command to capture and display a captured image; receiving another user command to create an annotation for the displayed image; in response to the second user command, receiving from the gaze tracking device a point-of-regard estimating a user's gaze in the displayed image; displaying an annotation anchor on the image proximate to the point-of-regard; and receiving a spoken annotation from the user and associating the spoken annotation with the annotation anchor. A gaze annotation method for a real-world scene includes: receiving a field of view and location information; receiving from the gaze tracking device a point-of-regard from the user located within the field of view; capturing and displaying a captured image of the field of view; while capturing the image, receiving a spoken annotation from the user; and displaying an annotation anchor on the image.


