Gaze Tracking and Bystander Exclusion in Video Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video collaboration systems face inaccuracies and inconsistencies in tracking user gaze and distinguishing between intended users and bystanders, especially in multi-party communications, which affects the reliability of input commands and user interactions.
Innovation Solution
A system that includes processors, a display, and an image sensor to determine the user's gaze direction, perform facial recognition, and track the user's face and body, using confirmation inputs to generate input commands while excluding bystanders, and calculating confidence scores to ensure accurate user identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If gaze tracking is used to generate input commands in multi-party video systems, then user interaction capability is improved, but accuracy in distinguishing intended users from bystanders deteriorates
Solution Approach 1:
The system segments the field of view into multiple regions of interest, each associated with a detected face. By dividing the monitoring space into distinct zones and tracking which region the gaze is directed toward, the system can differentiate between multiple users and bystanders, resolving the contradiction between supporting multi-user interaction and maintaining accurate user identification.
Solution Approach 2:
The system introduces an intermediary confirmation mechanism where gaze direction alone is not sufficient to generate input commands. Instead, the system requires additional confirmation (such as facial recognition verification or prolonged gaze duration) to validate that the gazed-at region corresponds to the intended user, thereby improving identification accuracy while preserving interaction capability.
2Measurement precision
If facial recognition is performed periodically to confirm user identity, then user identification accuracy is improved, but system processing time increases
Solution Approach 1:
The system implements periodic facial recognition analysis at strategically chosen intervals rather than continuously. By triggering recognition checks based on specific conditions (such as when gaze enters a region of interest or when interaction is anticipated), the system maintains high identification accuracy while minimizing unnecessary processing time and computational overhead.
Solution Approach 2:
The system uses the user's own gaze behavior and facial data already being captured for video collaboration purposes to perform self-verification. By leveraging existing observational data from the video system rather than requiring separate authentication inputs, the system achieves accurate user confirmation without adding significant processing time or user burden.
3Reliability
If the system tracks face and body location to distinguish users from bystanders, then user differentiation capability is improved, but computational complexity increases
Solution Approach 1:
The system uses a single image sensor and processing pipeline to simultaneously perform multiple functions: video collaboration, gaze tracking, face detection, and body location tracking. By making the observational system multi-functional rather than adding separate specialized sensors for each task, the system improves user differentiation capability while minimizing the increase in computational complexity through resource sharing and integrated processing.
Data Source
AI summary
Some embodiments include a method comprising receiving gaze data from an image sensor that indicates where a user is looking, determining a location that the user is directing their gaze on a display based on the gaze data, receiving a confirmation input from an input device, and generating and effectuating an input command based on the location on the display that the user is directing their gaze when the confirmation input is received. When a bystander is in a field-of-view of the image sensor, the method may further include limiting the input command to be generated and effectuated based solely on the location on the display that the user is directing their gaze and the confirmation input and actively excluding detected bystander gaze data from the generation of the input command.


