Eye Gaze Correction via Machine Learning Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conference technologies require additional equipment like semitransparent mirrors or stereocameras to correct gaze offset, and existing methods necessitate prerecording and result in unnatural gaze direction, while also being resource-intensive.

Innovation Solution

A method using machine learning to predict displacement vectors for correcting gaze orientation in images using a single webcam, employing neural networks or decision trees to adjust pixel color components, allowing for accurate eye image correction with reduced resource usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If additional equipment like semitransparent mirrors or stereocameras is used, then gaze correction accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvegaze correction accuracyVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent extracts and isolates the eye region from the full face image, focusing computational resources only on the relevant area for gaze correction. This allows accurate gaze correction using a single webcam by concentrating processing on eye pixels rather than the entire image, resolving the contradiction between accuracy and device complexity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent replaces mechanical/optical correction systems (semiransparent mirrors, stereocameras) with a computational approach using machine learning models. The neural network predicts gaze direction from eye image features, substituting physical correction mechanisms with algorithmic processing, thereby achieving accurate gaze correction with simpler hardware.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If prerecording of imagery data is performed, then correction accuracy is improved, but loss of time increases

Engineering Contradiction:
Improvecorrection accuracyVSAvoidprerecording time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary training of the neural network model using prerecorded imagery data, but this training is done once offline. During actual video conferencing, the pre-trained model processes images in real-time without requiring additional prerecording, thus eliminating the time loss during live sessions while maintaining correction accuracy through the pre-learned model.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements a dynamic approach where the system adapts to different users and conditions. The model can be retrained or fine-tuned when needed, but operates dynamically in real-time during video calls, adjusting to varying gaze directions and lighting conditions without requiring prerecording of each session, thereby resolving the time loss issue.

Inventive Principle:
Principle #15Dynamics

3Measurement precision

If complex geometric modeling and texture projection are used, then correction accuracy is improved, but use of energy increases

Engineering Contradiction:
Improvecorrection accuracyVSAvoiduse of energy
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The patent extracts only the eye region from the full face image and processes only this extracted portion through the neural network. This selective extraction and processing of relevant pixels significantly reduces computational energy requirements compared to processing entire face images with complex geometric modeling, while maintaining correction accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses a lightweight neural network model that can be deployed on standard consumer hardware without requiring powerful graphic accelerators. The model processes images efficiently using simplified computations rather than expensive complex geometric modeling and texture projection, reducing energy consumption while achieving practical correction accuracy.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

4Measurement precision

If global face proportion deformation is applied, then gaze correction is achieved, but manufacturing precision worsens

Engineering Contradiction:
Improvegaze correctionVSAvoidface proportion accuracy
Core Design Contradiction:
Measurement precisionVSManufacturing precision

Solution Approach 1:

The patent extracts and processes only the eye region independently, rather than deforming the entire face. By focusing computational processing on eye pixels separately from the rest of the face, the system achieves gaze correction without distorting global face proportions, thereby maintaining both correction accuracy and face structure integrity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies local processing to the eye region with different transformation rules than the rest of the face. The neural network predicts gaze correction specifically for eye pixels based on local eye features, while leaving other facial regions unchanged, thus achieving gaze correction without compromising overall face proportion accuracy.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11908241B2Method for correction of the eyes image using machine learning and method for machine learning
Publication Date: 2024.02.20 AVTONOMNAYA NEKOMMERCHESKAYA OBRAZOVATELNAYA ORGANIZATSIYA VYSSHEGO OBRAZOVANIYA SKOLKOVSKIJ INST NAUKI I TEKHNOLOGIJ
  • US11908241B2 patent drawing
  • US11908241B2 patent drawing
  • US11908241B2 patent drawing

AI summary

The present invention refers to automatics and computing technology, namely to the field of processing images and video data, namely to correction the eyes image of interlocutors in course of video chats, video conferences with the purpose of gaze redirection. A method of correction of the image of eyes wherein the method obtains, at least, one frame with a face of a person, whereupon determines positions of eyes of the person in the image and forms two rectangular areas closely circumscribing the eyes, and finally replaces color components of each pixel in the eye areas for color components of a pixel shifted according to prediction of the predictor of machine learning. Technical effect of the present invention is rising of correction accuracy of the image of eyes with the purpose of gaze redirection, with decrease of resources required for the process of handling a video image.