Multi-View Eye Patch Transformation for Accurate Gaze Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In systems where stereoscopic or multi-view images of a head are captured and displayed, the perceived gaze of the subject may not be directed at the observer due to camera positioning relative to the display, leading to disconcerting errors in human interactions.
Innovation Solution
An image processing technique that identifies and transforms image patches containing the eyes, using derived feature vectors to correct gaze by adjusting displacement vector fields based on reference data, ensuring consistent transformations across images.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If camera system is offset above the display, then the display can be positioned for optimal viewing, but the perceived gaze direction becomes incorrect (downwards)
Solution Approach 1:
The patent applies preliminary action by pre-computing displacement vector fields and storing them in reference data structures before actual gaze correction is needed. The system prepares correction mappings in advance that can be quickly applied during teleconferencing operations, allowing the camera to remain in its optimal offset position while gaze accuracy is maintained through pre-prepared correction data.
Solution Approach 2:
The patent changes parameters by transforming the image data through displacement vector fields that adjust pixel positions. By modifying the positional parameters of image features through learned displacement fields, the system corrects the perceived gaze direction without physically moving the camera or display, thus maintaining both optimal display positioning and accurate gaze perception.
2Ease of manufacture
If displacement vector fields are derived independently for each image, then processing is simpler, but the transformations become inconsistent across stereoscopic images
Solution Approach 1:
The patent applies segmentation by dividing the correction process into distinct components: extracting local image descriptors from each image, computing feature vectors, and applying displacement vector fields. By segmenting the processing into these manageable steps while maintaining coordination through the feature vector lookup, the system achieves both processing simplicity and transformation consistency across stereoscopic image pairs.
Solution Approach 2:
The patent implements feedback by using the extracted feature vectors to look up corresponding displacement vector fields from reference data. This feedback mechanism ensures that the transformation applied to each image is coordinated with the other, maintaining consistency across the stereoscopic pair while allowing independent processing of each image's correction.
3Measurement precision
If feature vectors are derived from local image descriptors, then gaze correction accuracy improves, but computational complexity increases
Solution Approach 1:
The patent reduces computational complexity through preliminary action by pre-computing and storing displacement vector fields in reference data structures during an offline training phase. During actual teleconferencing operations, the system only needs to extract local image descriptors, compute feature vectors, and perform a lookup operation, significantly reducing the computational burden while maintaining high gaze correction accuracy.
Data Source
AI summary
Gaze is corrected by adjusting multi-view images of a head. Image patches containing the left and right eyes of the head are identified and a feature vector is derived from plural local image descriptors of the image patch in at least one image of the multi-view images. A displacement vector field representing a transformation of an image patch is derived, using the derived feature vector to look up reference data comprising reference displacement vector fields associated with possible values of the feature vector produced by machine learning. The multi-view images are adjusted by transforming the image patches containing the left and right eyes of the head in accordance with the derived displacement vector field.


