Eye Gaze Correction via Machine Learning Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current video conference technologies require additional equipment like semitransparent mirrors or stereocameras to correct gaze offset, and existing methods necessitate prerecording and result in unnatural gaze direction, while also being resource-intensive.
Innovation Solution
A method using machine learning to predict displacement vectors for correcting gaze orientation in images using a single webcam, employing neural networks or decision trees to adjust pixel color components, allowing for accurate eye image correction with reduced resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If additional equipment like semitransparent mirrors or stereocameras is used, then gaze correction accuracy is improved, but device complexity increases
Solution Approach 1:
The patent extracts and isolates the eye region from the full face image, focusing computational resources only on the relevant area for gaze correction. This allows accurate gaze correction using a single webcam by concentrating processing on eye pixels rather than the entire image, resolving the contradiction between accuracy and device complexity.
Solution Approach 2:
The patent replaces mechanical/optical correction systems (semiransparent mirrors, stereocameras) with a computational approach using machine learning models. The neural network predicts gaze direction from eye image features, substituting physical correction mechanisms with algorithmic processing, thereby achieving accurate gaze correction with simpler hardware.
2Measurement precision
If prerecording of imagery data is performed, then correction accuracy is improved, but loss of time increases
Solution Approach 1:
The patent performs preliminary training of the neural network model using prerecorded imagery data, but this training is done once offline. During actual video conferencing, the pre-trained model processes images in real-time without requiring additional prerecording, thus eliminating the time loss during live sessions while maintaining correction accuracy through the pre-learned model.
Solution Approach 2:
The patent implements a dynamic approach where the system adapts to different users and conditions. The model can be retrained or fine-tuned when needed, but operates dynamically in real-time during video calls, adjusting to varying gaze directions and lighting conditions without requiring prerecording of each session, thereby resolving the time loss issue.
3Measurement precision
If complex geometric modeling and texture projection are used, then correction accuracy is improved, but use of energy increases
Solution Approach 1:
The patent extracts only the eye region from the full face image and processes only this extracted portion through the neural network. This selective extraction and processing of relevant pixels significantly reduces computational energy requirements compared to processing entire face images with complex geometric modeling, while maintaining correction accuracy.
Solution Approach 2:
The patent uses a lightweight neural network model that can be deployed on standard consumer hardware without requiring powerful graphic accelerators. The model processes images efficiently using simplified computations rather than expensive complex geometric modeling and texture projection, reducing energy consumption while achieving practical correction accuracy.
4Measurement precision
If global face proportion deformation is applied, then gaze correction is achieved, but manufacturing precision worsens
Solution Approach 1:
The patent extracts and processes only the eye region independently, rather than deforming the entire face. By focusing computational processing on eye pixels separately from the rest of the face, the system achieves gaze correction without distorting global face proportions, thereby maintaining both correction accuracy and face structure integrity.
Solution Approach 2:
The patent applies local processing to the eye region with different transformation rules than the rest of the face. The neural network predicts gaze correction specifically for eye pixels based on local eye features, while leaving other facial regions unchanged, thus achieving gaze correction without compromising overall face proportion accuracy.
Data Source
AI summary
The present invention refers to automatics and computing technology, namely to the field of processing images and video data, namely to correction the eyes image of interlocutors in course of video chats, video conferences with the purpose of gaze redirection. A method of correction of the image of eyes wherein the method obtains, at least, one frame with a face of a person, whereupon determines positions of eyes of the person in the image and forms two rectangular areas closely circumscribing the eyes, and finally replaces color components of each pixel in the eye areas for color components of a pixel shifted according to prediction of the predictor of machine learning. Technical effect of the present invention is rising of correction accuracy of the image of eyes with the purpose of gaze redirection, with decrease of resources required for the process of handling a video image.


