Gaze Estimation Training Data via Multi-Camera Feedback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current gaze estimation systems face challenges in accurately tracking gaze direction, especially when the subject's position is obstructed or outside the camera's view, and existing training data methods do not adequately prepare for difficult estimation scenarios.

Innovation Solution

A system and method that involves displaying a target image at a known location, using multiple image sensors and an eye-tracker to determine a reference gaze vector, calculating uncertainty and error, and providing feedback to collect robust training data, encouraging the subject to intentionally make it difficult for the system to estimate gaze direction to improve model performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If a single camera is used for gaze tracking, then the system is simple, but the subject's position must remain in view of the camera to produce accurate tracking results

Engineering Contradiction:
Improvesystem complexityVSAvoidtracking envelope
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The system divides the tracking task into multiple segments by using multiple cameras positioned at different locations. Each camera captures images from its own perspective, and the system integrates data from all cameras to achieve comprehensive gaze tracking coverage that exceeds what a single camera could provide.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The multiple cameras serve universal functions by each being capable of capturing gaze information independently. Any single camera can potentially track the subject, but collectively they provide redundant and complementary views that ensure tracking accuracy across a wider range of subject positions and orientations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of manufacture

If training data is gathered by positioning a person in front of a large screen, then the data collection process is simple, but the method does not provide training data for situations when it is difficult for the gaze estimation system to estimate the gaze

Engineering Contradiction:
Improvedata collection processVSAvoidtraining data robustness
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The system dynamically adjusts the difficulty of training scenarios by varying subject orientations, occlusions, and head positions. Instead of static front-facing views, the system creates diverse dynamic conditions that simulate real-world challenging situations, making the training data more robust while maintaining automated collection processes.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system pre-configures multiple cameras and occlusion elements before data collection to automatically create challenging training scenarios. This preliminary setup enables the system to generate diverse training data without manual intervention during the actual data collection process, maintaining ease of manufacture while improving reliability.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the subject's face is obstructed or turned away, then the gaze estimation system cannot determine the gaze, but the system lacks capability to handle such situations

Engineering Contradiction:
Improvehandling difficult scenariosVSAvoidgaze estimation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system transitions from a single 2D camera view to multiple 2D camera views that collectively provide 3D spatial information. By capturing images from different angles and combining them, the system can infer gaze direction even when the face is partially obscured or turned away, as other cameras may capture visible features from their perspectives.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS10866635B2Systems and methods for capturing training data for a gaze estimation model
Publication Date: 2020.12.15 TOYOTA JIDOSHA KK
  • US10866635B2 patent drawing
  • US10866635B2 patent drawing
  • US10866635B2 patent drawing

AI summary

A method of training a gaze estimation model includes displaying a target image at a known location on a display in front of a subject and receiving images captured from a plurality of image sensors surrounding the subject, wherein each image sensor has a known location relative to the display. The method includes determining a reference gaze vector for one or more eyes of the subject based on the images and the known location of the target image and then determining, with the model, a gaze direction vector of each of the one or more eyes of the subject from data captured by an eye-tracker. The method further includes determining, with the model, an uncertainty in measurement of the gaze direction vector and an error between the reference gaze vector and the gaze direction vector and providing feedback based on at least one of the uncertainty and the error.