Gaze Estimation Training Data via Multi-Camera Feedback
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current gaze estimation systems face challenges in accurately tracking gaze direction, especially when the subject's position is obstructed or outside the camera's view, and existing training data methods do not adequately prepare for difficult estimation scenarios.
Innovation Solution
A system and method that involves displaying a target image at a known location, using multiple image sensors and an eye-tracker to determine a reference gaze vector, calculating uncertainty and error, and providing feedback to collect robust training data, encouraging the subject to intentionally make it difficult for the system to estimate gaze direction to improve model performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single camera is used for gaze tracking, then the system is simple, but the subject's position must remain in view of the camera to produce accurate tracking results
Solution Approach 1:
The system divides the tracking task into multiple segments by using multiple cameras positioned at different locations. Each camera captures images from its own perspective, and the system integrates data from all cameras to achieve comprehensive gaze tracking coverage that exceeds what a single camera could provide.
Solution Approach 2:
The multiple cameras serve universal functions by each being capable of capturing gaze information independently. Any single camera can potentially track the subject, but collectively they provide redundant and complementary views that ensure tracking accuracy across a wider range of subject positions and orientations.
2Ease of manufacture
If training data is gathered by positioning a person in front of a large screen, then the data collection process is simple, but the method does not provide training data for situations when it is difficult for the gaze estimation system to estimate the gaze
Solution Approach 1:
The system dynamically adjusts the difficulty of training scenarios by varying subject orientations, occlusions, and head positions. Instead of static front-facing views, the system creates diverse dynamic conditions that simulate real-world challenging situations, making the training data more robust while maintaining automated collection processes.
Solution Approach 2:
The system pre-configures multiple cameras and occlusion elements before data collection to automatically create challenging training scenarios. This preliminary setup enables the system to generate diverse training data without manual intervention during the actual data collection process, maintaining ease of manufacture while improving reliability.
3Adaptability or versatility
If the subject's face is obstructed or turned away, then the gaze estimation system cannot determine the gaze, but the system lacks capability to handle such situations
Solution Approach 1:
The system transitions from a single 2D camera view to multiple 2D camera views that collectively provide 3D spatial information. By capturing images from different angles and combining them, the system can infer gaze direction even when the face is partially obscured or turned away, as other cameras may capture visible features from their perspectives.
Data Source
AI summary
A method of training a gaze estimation model includes displaying a target image at a known location on a display in front of a subject and receiving images captured from a plurality of image sensors surrounding the subject, wherein each image sensor has a known location relative to the display. The method includes determining a reference gaze vector for one or more eyes of the subject based on the images and the known location of the target image and then determining, with the model, a gaze direction vector of each of the one or more eyes of the subject from data captured by an eye-tracker. The method further includes determining, with the model, an uncertainty in measurement of the gaze direction vector and an error between the reference gaze vector and the gaze direction vector and providing feedback based on at least one of the uncertainty and the error.


