Eye-Gaze Detection Using Single-Axis Calibration on Low-Res Cameras
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current eye-gaze data collection methods require expensive, high-resolution, high-frame-rate cameras and complex calibration processes, limiting their accessibility and accuracy on standard devices.
Innovation Solution
Methods for accurate eye-gaze detection using low-resolution, low-frame-rate cameras by constraining the area of interest to a single axis, such as the x-axis, and employing machine learning models trained on single-axis calibration points, allowing for word-level accuracy and reduced calibration time.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If high-resolution, high-frame-rate cameras are used for eye-gaze detection, then measurement precision is improved, but device cost and complexity increase
Solution Approach 1:
The patent segments the eye-gaze detection task into two independent one-dimensional problems (x-axis and y-axis) instead of solving the full two-dimensional problem simultaneously. This allows using simpler, lower-cost cameras capable of capturing only one dimension at a time, while still achieving accurate word-level eye-gaze detection through the combination of results from both axes.
Solution Approach 2:
The patent transforms the two-dimensional eye-gaze detection problem into two separate one-dimensional problems by constraining the area of interest to single axes. This dimensional reduction enables the use of low-resolution, low-frame-rate cameras while maintaining detection accuracy, as each axis can be captured independently with simpler hardware.
2Measurement precision
If advanced machine learning models with high frame rate cameras are used, then measurement precision is improved, but calibration time and setup complexity increase
Solution Approach 1:
The calibration process is segmented into two independent one-dimensional calibration procedures (x-axis and y-axis) rather than a single complex two-dimensional calibration. This segmentation reduces the complexity of each calibration step and allows for faster, more efficient calibration using standard cameras, while still achieving word-level accuracy through the combination of calibrated axes.
3Adaptability or versatility
If standard low-resolution cameras are used for eye-gaze detection, then device accessibility is improved, but measurement precision deteriorates
Solution Approach 1:
By segmenting the detection task into independent x-axis and y-axis components, the patent enables standard low-resolution cameras to achieve word-level detection accuracy. Each axis can be captured with simpler hardware, and the combination of these one-dimensional results provides sufficient precision for identifying word-level eye-gaze locations, making the system accessible to devices with standard cameras.
Solution Approach 2:
The patent reduces the dimensional requirements for camera capability by transforming the two-dimensional detection problem into two separate one-dimensional problems. This allows standard cameras with lower resolution and frame rate to be sufficient, as they only need to capture one dimension at a time, while the combination of both axes achieves the required measurement precision for word-level accuracy.
Data Source
AI summary
Disclosed herein are methods and system for eye-gaze location detection and accurate collection of eye-gaze data from low resolution, low frame rate cameras. The method includes calibrating a machine learning model by constraining an area of interest to a single axis. Presenting to the user one or more lines of words and/or images on a screen, capturing by a camera one or more images of the user's eye-gaze, looking at said lines of words and/or images, applying the calibrated machine learning model on the one or more images of the user's eye-gaze, constraining the area of interest to a single-axis, and detecting an eye-gaze location on the screen with a word-level accuracy. The method further includes combining eye-gaze and audio data collection by measuring a time difference between a first timestamp of a word eye-gaze location detection and a second timestamp of the word first phoneme pronunciation by a user.


