Eye Image Gaze Correction via Style Transfer
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current gaze correction methods in video processing, especially those using 3D devices or deep neural networks, face challenges such as high hardware costs, low training efficiency, and the need for large numbers of training samples, resulting in inefficient model training and resource consumption.
Innovation Solution
A method for training an image processing model that involves acquiring a set of to-be-corrected eye images, performing style transfer using a style transfer network, and generating a training sample with eye images at different gaze positions to train a correction network, thereby improving training efficiency and reducing resource usage while maintaining recognition precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If gaze correction is performed using 3D devices, then the gaze direction can be corrected, but the hardware costs become high
Solution Approach 1:
The patent replaces 3D hardware devices with a deep learning-based image processing system. Instead of using complex 3D imaging equipment to capture and correct gaze, the system uses 2D images processed through neural networks to achieve gaze correction, thereby eliminating high hardware costs while maintaining correction capability
Solution Approach 2:
The patent creates synthetic training data by copying and transforming existing eye images to simulate various gaze directions and lighting conditions. This allows the model to be trained without requiring expensive 3D capture equipment, as the synthetic data replicates the information that would otherwise require specialized hardware
2Reliability
If gaze correction is performed through deep neural network, then the correction can be achieved, but the training efficiency is low and computing resources are consumed
Solution Approach 1:
The patent performs preliminary actions by pre-processing images to extract eye regions and pre-aligning them to standardized positions before training. This preliminary preparation organizes the data in advance, allowing the neural network to focus on learning gaze patterns rather than basic image processing, thereby improving training efficiency
Solution Approach 2:
The patent segments the training process into distinct stages: data collection, synthetic data generation, model training, and evaluation. This segmentation allows each stage to be optimized independently, with synthetic data generation creating targeted training samples that reduce the overall computational burden
3Measurement precision
If a large number of training samples are collected and labeled, then the recognition precision can be improved, but the sample collection and labeling process becomes time-consuming
Solution Approach 1:
The patent creates synthetic copies of eye images by applying transformations such as rotation, scaling, and lighting changes to a limited set of real eye images. These synthetic copies serve as training samples, providing diverse gaze scenarios without requiring manual collection and labeling of numerous real images
Solution Approach 2:
The system performs self-service by automatically generating training data and labels through algorithmic processes. The synthetic data generation pipeline automatically creates labeled training samples without human intervention, eliminating the time-consuming manual labeling process while maintaining data quality
Data Source
AI summary
A method for training an image processing model includes: acquiring a to-be-corrected eye image set matching a usage environment of the image processing model; performing style transfer on to-be-corrected eye images in the to-be-corrected eye image set through a style transfer network in the image processing model, to obtain a target eye image; acquiring a training sample matching the usage environment of the image processing model based on the to-be-corrected eye images and the target eye image, the training sample including object eye images matching different gaze positions; and training a correction network in the image processing model through the training sample matching the usage environment of the image processing model, to obtain a model update parameter matching the correction network, and generating a trained image processing model based on the model update parameter.


