Gaze-Assisted Input with Speech Confirmation for Precise Targeting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing gaze estimation systems struggle with accurately determining when a user intends to trigger a functionality, leading to responsiveness issues due to inappropriate dwell time settings, and are limited in applications beyond simple gaze detection on common hardware devices.
Innovation Solution
A method that utilizes eye input in the form of x,y Point-of-Gaze coordinates to trigger functionality on electronic devices, combined with speech inputs, by identifying a subject display element through spatial and temporal overlap, and performing actions based on this identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If dwell time is set short for responsive gaze-triggered functionality, then system responsiveness improves, but false triggers increase causing overly responsive behavior
Solution Approach 1:
The patent combines multiple input modalities (gaze estimation and speech recognition) to trigger functionality. The speech input acts as a confirmation signal that merges with the gaze direction data, allowing the system to respond quickly when both signals align while avoiding false triggers from accidental or unintended gaze movements alone.
Solution Approach 2:
The speech recognition system serves as an intermediary confirmation mechanism. Instead of relying solely on dwell time as the trigger condition, the system uses speech input as an intermediate step that validates the user's intent to trigger the functionality associated with the gazed-at element.
2Reliability
If dwell time is set long to avoid false triggers, then false trigger rate decreases, but system appears unresponsive
Solution Approach 1:
By merging gaze estimation with speech recognition, the system eliminates the need for long dwell times. The speech input provides rapid confirmation of user intent, allowing the system to respond immediately without requiring extended gaze periods, thus maintaining both reliability and responsiveness.
3Measurement precision
If gaze estimation is used for precise region identification, then input precision improves, but device complexity increases
Solution Approach 1:
The patent applies gaze estimation technology across multiple device types and contexts (mobile devices, tablets, AR/VR headsets, IoT devices). By creating a universal solution that works on common hardware with standard cameras and processors, the system achieves high precision without requiring specialized or complex hardware for each device type.
Solution Approach 2:
The system uses the device's existing camera and processor to perform gaze estimation, allowing the device to serve itself for eye tracking functionality. This self-service approach eliminates the need for separate, complex eye-tracking hardware while maintaining accurate gaze detection capabilities.
4Ease of operation
If speech input alone is used for command recognition, then ease of operation improves, but command disambiguation becomes difficult
Solution Approach 1:
The patent merges speech input with gaze direction data to resolve command ambiguity. When a user speaks a command while looking at a specific display element, the gaze information disambiguates which element the command applies to, eliminating the information loss that would occur with speech input alone.
Solution Approach 2:
The system provides visual feedback by highlighting or indicating the display element that has been identified as the target of the speech command. This feedback loop allows users to confirm that their gaze and speech are aligned with the intended target, reducing misinterpretation of command intent.
Data Source
AI summary
A computer implemented method and system for gaze-assisted input that includes displaying a plurality of display elements in a display space, tracking a user's point of gaze within the display space, receiving a speech input, identifying one of the plurality of display elements as a subject display element for the speech input based on the tracking, and automatically performing an action based on the subject display element and the speech input.


