Tunable Neural Network for Facial Identity Interpolation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques for changing facial identities in high-resolution images or video frames are limited, requiring re-training of neural network models for each change and needing large amounts of labeled data, making them ineffective for blending different facial identities, especially in high-resolution scenarios.

Innovation Solution

A computer-implemented method using a machine learning model with an encoder and decoder, along with dense layers, to generate latent representations and output images with changed facial identities, allowing for interpolation between identities without separate training for each combination and without requiring labeled data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing neural network models are used to change facial identities, then facial identity can be changed, but the models need to be re-trained for each change and require large amounts of labeled data

Engineering Contradiction:
Improvefacial identity change capabilityVSAvoidmodel re-training time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent pre-trains a single neural network model on a diverse dataset of multiple facial identities before deployment. This preliminary training enables the model to adapt to different facial identities without requiring re-training when identity changes are needed, thus resolving the contradiction between adaptability and training time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a universal neural network model that can handle multiple facial identities simultaneously through a single trained model. This multi-functional model replaces the need for separate trained models for each identity, eliminating re-training requirements while maintaining versatility across different identities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Adaptability or versatility

If existing techniques are used to change facial identities, then identity transformation is possible, but the techniques are limited to low-resolution images

Engineering Contradiction:
Improvefacial identity transformationVSAvoidimage resolution quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent modifies the neural network model's parameters and architecture to process and preserve high-resolution image data. By adjusting parameters such as network depth, filter sizes, and processing stages, the model achieves both high-resolution output quality and effective facial identity transformation capability.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If multiple facial identities need to be blended together, then realistic CG images can be produced, but separate training is required for each combination of identities

Engineering Contradiction:
ImproveCG image realismVSAvoidtraining process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a universal blending mechanism within a single trained model that can combine multiple facial identities through parameter adjustment rather than separate training processes. This reduces training complexity while maintaining the ability to produce realistic blended CG images of multiple identities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11568524B2Tunable models for changing faces in images
Publication Date: 2023.01.31 DISNEY ENTERPRISES INC
  • US11568524B2 patent drawing
  • US11568524B2 patent drawing
  • US11568524B2 patent drawing

AI summary

Techniques are disclosed for changing the identities of faces in images. In embodiments, a tunable model for changing facial identities in images includes an encoder, a decoder, and dense layers that generate either adaptive instance normalization (AdaIN) coefficients that control the operation of convolution layers in the decoder or the values of weights within such convolution layers, allowing the model to change the identity of a face in an image based on a user selection. A separate set of dense layers may be trained to generate AdaIN coefficients for each of a number of facial identities, and the AdaIN coefficients output by different sets of dense layers can be combined to interpolate between facial identities. Alternatively, a single set of dense layers may be trained to take as input an identity vector and output AdaIN coefficients or values of weighs within convolution layers of the decoder.