Eye Tracking with Keypoint Velocity Mapping for Low-Noise Precision
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current feature-based eye tracking systems suffer from significant noise due to the detection of first-order features like the pupil and corneal reflection, leading to inaccuracies and temporal smoothing methods that sacrifice precision and accuracy.
Innovation Solution
A method that identifies keypoints in a sequence of images, maps corresponding keypoints between images, calculates individual velocities, and extracts distributions to obtain eye-in-head velocity, reducing noise by considering spatial domain distributions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If feature-based eye tracking systems use first-order features (pupil and corneal reflection) for gaze detection, then the system can determine eye position, but significant noise is introduced leading to measurement inaccuracies
Solution Approach 1:
The patent segments the eye region into multiple keypoints (pupil center, corneal reflection, iris boundary points) and tracks each independently. By dividing the eye into multiple feature points rather than relying on a single pupil-CR feature pair, the system reduces the impact of noise on any single point while maintaining accurate gaze detection through aggregation of multiple measurements.
Solution Approach 2:
The patent transitions from tracking only 2D image plane positions to incorporating 3D spatial information by calculating velocity vectors in three dimensions. By deriving velocity signals from 3D keypoint positions and integrating them to obtain gaze angles, the system adds a temporal-velocity dimension that helps filter noise while preserving accurate eye movement tracking.
2Reliability
If temporal smoothing (averaging or median filtering) is applied to reduce noise in eye tracking signals, then measurement noise is masked, but temporal precision is sacrificed
Solution Approach 1:
The patent replaces traditional mechanical/temporal filtering methods (averaging, median filtering) with a spatial-domain approach. Instead of smoothing signals over time, the system uses spatial distribution analysis of multiple keypoints and their velocity vectors to naturally reduce noise. This substitution of the noise-reduction mechanism preserves temporal precision while achieving reliable noise filtering through geometric and statistical properties of the keypoint distribution.
3Measurement precision
If multiple cameras and illuminators are used to build a 3D model of the eye, then calibration-free tracking is achieved, but device complexity increases
Solution Approach 1:
The patent makes a single camera multi-functional by using it to capture both the eye region and surrounding facial features simultaneously. The same camera that tracks eye keypoints also captures non-eye region keypoints for motion compensation, eliminating the need for separate cameras. This universal use of a single imaging device maintains 3D tracking accuracy while significantly reducing system complexity.
Solution Approach 2:
The system uses the subject's own facial features (non-eye region keypoints) to compensate for head and camera motion. By leveraging naturally present facial landmarks that move with the head, the system achieves self-compensation without requiring external reference cameras or complex calibration procedures. The facial features themselves provide the reference frame needed for accurate eye tracking.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method and system for monitoring the motion of one or both eyes, includes capturing a sequence of overlapping images of a subject's face including an eye and the corresponding non-eye region; identifying a plurality of keypoints in each image; mapping corresponding keypoints in two or more images of the sequence; assigning the keypoints to the eye and to the corresponding non-eye region; calculating individual velocities of the corresponding keypoints in the eye and the corresponding non-eye region to obtain a distribution of velocities; extracting at least one velocity measured for the eye and at least one velocity measured for the corresponding non-eye region; calculating the eye-in-head velocity for the eye based upon the measured velocity for the eye and the measured velocity for the corresponding non-eye region; and calculating the eye-in-head position based upon the eye- in-head velocity.