Deep Learning Face Object Removal Using Gaze Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer imaging methods face challenges in accurately and realistically adding or removing objects from face images, particularly due to difficulties in capturing and utilizing head pose and gaze-tracking data for precise object placement or removal.

Innovation Solution

The integration of deep learning networks that incorporate head pose analysis, gaze tracking, and object characteristics for training neural networks to encode and decode facial object characteristics, enabling the generation of realistic augmented face images by accurately adding or removing objects based on user input.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If deep learning networks are trained with object characteristics and gaze tracking data, then the accuracy and realism of object placement or removal is improved, but the device complexity and computational demands increase

Engineering Contradiction:
Improveaccuracy of object placementVSAvoidcomplexity of neural network system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments the face image into multiple regions including eye regions, facial landmarks, and object regions. The neural network processes these segmented regions independently, integrating head pose data and gaze tracking information to achieve accurate object placement while managing computational complexity through localized processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent incorporates head pose data and gaze tracking information as additional dimensional inputs to the neural network. By adding these dimensional parameters (head orientation angles, gaze direction vectors) to the traditional 2D face image data, the system achieves more accurate object placement without fundamentally changing the network architecture.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If gaze tracking and head pose analysis are integrated into the deep learning network, then the realism of augmented face images is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improverealism of augmented imagesVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary processing by extracting head pose data and gaze tracking information before feeding the data into the neural network. Pre-computing these parameters from the input images allows the main neural network to focus on object placement tasks, reducing overall processing time while maintaining realism.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses intermediate representations that combine face image data with head pose and gaze information. These intermediate features serve as mediators that bridge the gap between raw input data and final augmented images, enabling more efficient processing by pre-organizing information in a compressed representation space.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If the neural network encodes and decodes facial object characteristics using gaze tracking data, then the precision of object removal is improved, but the difficulty of detecting and measuring head pose and gaze parameters increases

Engineering Contradiction:
Improveprecision of object removalVSAvoiddifficulty of detecting head pose and gaze
Core Design Contradiction:
Manufacturing precisionVSDifficulty of detecting and measuring

Solution Approach 1:

The system creates a virtual copy of the face with adjusted object characteristics by encoding and decoding through the neural network. The gaze tracking data is used to determine the virtual camera position and viewing angle, allowing precise object removal and replacement while simplifying the detection of head pose and gaze parameters through standard image processing techniques.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20240303889A1Passive and continuous deep learning methods and systems for removal of objects relative to a face
Publication Date: 2024.09.12 BLINK O G LTD
  • US20240303889A1 patent drawing
  • US20240303889A1 patent drawing
  • US20240303889A1 patent drawing

AI summary

Methods, systems, and computer-readable media for generating realistic augmented face images are disclosed. Implementations include a) receiving head pose data, segmented eye region image data, eye position data, gaze direction data, and face and eye landmark data from an individual; b) receiving a user selection of a facial object to be added or removed from an image of the individual; and c) using a deep learning model, i) generating one or more images of the individual with the facial object in place on the one or more images, or ii) generating one or more images of the individual with the facial object removed from the one or more images.