Unconstrained Gaze Estimation Using CNN and Synthetic Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current gaze tracking technologies require extensive calibration and high-quality image sensors, making them unsuitable for consumer-grade devices and environments with varying lighting conditions.
Innovation Solution
An unconstrained appearance-based gaze estimation method using a convolutional neural network (CNN) that identifies eye and head orientations from images captured with consumer-grade sensors, without the need for calibration, by analyzing images within the CNN and utilizing synthetic and real data for training to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If current gaze tracking technology is used, then accurate gaze estimation can be achieved, but extensive calibration and high-quality image sensors are required
Solution Approach 1:
The patent replaces traditional mechanical calibration procedures with a deep learning-based gaze estimation system. The CNN model is trained on synthetic and real eye images to learn gaze patterns, eliminating the need for manual calibration procedures while maintaining high accuracy in gaze estimation.
Solution Approach 2:
The patent uses synthetic eye images generated from 3D models as training data for the CNN. These synthetic copies of real eye appearances under various conditions allow the model to learn robust gaze estimation without requiring extensive real-world calibration data, reducing both calibration complexity and data collection requirements.
2Measurement precision
If current gaze tracking technology is used, then accurate gaze estimation can be achieved, but high-quality image sensors are required
Solution Approach 1:
The patent transforms the gaze estimation problem from a direct image analysis task to a learned parameter estimation task. The CNN model learns to map various image quality parameters to gaze directions, enabling accurate estimation even with lower-quality consumer-grade sensors by compensating for image quality variations through training.
Solution Approach 2:
By training the model on synthetic images that replicate various sensor qualities and lighting conditions, the system learns to generalize across different sensor types. This allows consumer-grade sensors to achieve comparable performance to high-quality sensors without requiring hardware upgrades.
3Measurement precision
If current gaze tracking technology is used, then accurate gaze estimation can be achieved, but proper lighting conditions are required
Solution Approach 1:
The patent creates a dynamic training dataset that includes eye images under varying lighting conditions, head poses, and eye appearances. The CNN model learns to adapt to different lighting scenarios by experiencing diverse conditions during training, enabling robust gaze estimation in unconstrained environments without requiring controlled lighting.
Solution Approach 2:
The system changes the approach from requiring controlled lighting parameters to learning invariant features that work across varying lighting conditions. The model learns to extract gaze-relevant features that remain stable despite lighting variations, effectively decoupling gaze estimation accuracy from specific lighting requirements.
4Ease of manufacture
If consumer-grade sensors are used, then device accessibility is improved, but gaze estimation accuracy deteriorates
Solution Approach 1:
The patent substitutes hardware-based accuracy requirements with software-based deep learning compensation. The CNN model compensates for the lower quality of consumer-grade sensors through learned features and patterns, achieving high accuracy without requiring expensive high-quality sensors, thus improving device accessibility.
Data Source
AI summary
A method, computer readable medium, and system are disclosed for performing unconstrained appearance-based gaze estimation. The method includes the steps of identifying an image of an eye and a head orientation associated with the image of the eye, determining an orientation for the eye by analyzing, within a convolutional neural network (CNN), the image of the eye and the head orientation associated with the image of the eye, and returning the orientation of the eye.


