Autoencoder Face Swapping Reducing Computational Resource Consumption

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional techniques for digital image and voice processing, such as face-swapping, require significant computer resources and time, limiting their efficiency and ability to produce high-resolution outputs.

Innovation Solution

The use of autoencoders trained with CGI faces and real faces to swap facial expressions efficiently, along with neural networks that minimize error and adjust weights for improved performance, allowing for reduced resource usage and higher resolution outputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If conventional techniques are used for digital image processing such as face-swapping, then processing capability is achieved, but computer resource consumption increases and processing time increases

Engineering Contradiction:
Improveface-swapping processing speedVSAvoidcomputer resource consumption
Core Design Contradiction:
ProductivityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary training of the autoencoder model using a comprehensive dataset of facial images and expressions before actual face-swapping operations. This pre-training phase allows the model to learn optimal feature representations and transformations, enabling fast and accurate face-swapping during inference without requiring excessive computational resources for each individual processing task

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses autoencoders to create latent representations and reconstructed images that copy essential facial features and expressions. The encoder compresses input facial images into latent codes, and the decoder reconstructs these codes into output images with swapped faces, preserving expressions through the copying of latent facial characteristics rather than direct pixel manipulation

Inventive Principle:
Principle #26Copying

2Productivity

If conventional techniques are used for digital image processing such as face-swapping, then processing capability is achieved, but processing time increases

Engineering Contradiction:
Improveface-swapping processing speedVSAvoidprocessing time
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The system performs preliminary training of the autoencoder model using a comprehensive dataset of facial images and expressions before actual face-swapping operations. This pre-training phase allows the model to learn optimal feature representations and transformations, enabling fast and accurate face-swapping during inference without requiring excessive computational resources for each individual processing task

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces conventional mechanical image processing techniques with neural network-based deep learning methods. The autoencoder uses neural networks to automatically learn and transform facial features, substituting traditional algorithmic approaches with biologically-inspired neural processing that achieves superior speed and accuracy

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Quantity of substance

If one-shot architecture is used to reduce the number of images needed, then training data requirement decreases, but model complexity increases

Engineering Contradiction:
Improvenumber of training imagesVSAvoidmodel architecture complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent implements a universal one-shot architecture where a single trained autoencoder model can perform multiple face-swapping tasks across different individuals and expressions. The model learns general facial feature representations during initial training and can then handle diverse face-swapping scenarios without requiring separate models for each individual, reducing the total number of training images needed while maintaining versatility

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS10671838B1Methods and systems for image and voice processing
Publication Date: 2020.06.02 NEON EVOLUTION INC
  • US10671838B1 patent drawing
  • US10671838B1 patent drawing
  • US10671838B1 patent drawing

AI summary

Systems and methods are disclosed configured to train an autoencoder using images that include faces, wherein the autoencoder comprises an input layer, an encoder configured to output a latent image from a corresponding input image, and a decoder configured to attempt to reconstruct the input image from the latent image. An image sequence of a face exhibiting a plurality of facial expressions and transitions between facial expressions is generated and accessed. Images of the plurality of facial expressions and transitions between facial expressions are captured from a plurality of different angles and using different lighting. An autoencoder is trained using source images that include the face with different facial expressions captured at different angles with different lighting, and using destination images that include a destination face. The trained autoencoder is used to generate an output where the likeness of the face in the destination images is swapped with the likeness of the source face, while preserving expressions of the destination face.