3D Face Model Texture and Shape Disentanglement for Realistic Expression Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for changing facial expressions in images result in blurry and non-realistic faces due to their inability to generate facial features that do not exist in the target image, such as teeth and tongue, which detracts from the user experience.

Innovation Solution

The approach involves separately processing the texture and shape of a 3D face model using machine learning techniques, specifically a conditional generative neural network for texture and a fully connected neural network for shape, to generate a new facial expression, allowing for the inclusion of features that do not exist in the input image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If conventional methods are used to change facial expressions in images, then the process is simple, but the output quality becomes blurry and non-realistic

Engineering Contradiction:
Improveface rendering qualityVSAvoidprocessing method complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent segments the face into multiple components (skin, eyes, eyebrows, mouth, teeth, tongue) and processes each component separately using specialized neural networks. This segmentation allows each component to be rendered with appropriate detail and realism, resolving the contradiction between simple processing and high-quality output.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transitions from 2D image processing to 3D face model manipulation. By creating and manipulating a 3D face model with depth information and spatial relationships, the system achieves realistic facial expressions while maintaining computational efficiency through structured processing.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If facial features that do not exist in the target image are generated, then expression realism improves, but the complexity of generating accurate features increases

Engineering Contradiction:
Improveexpression realismVSAvoidfeature generation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary actions by pre-defining the structure and possible configurations of facial features (teeth, tongue, eyes, etc.) in the 3D face model before expression transformation. This preliminary setup allows the system to generate realistic features that may not be visible in the original image by simply activating pre-modeled structures, rather than generating them from scratch.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes parameters of the 3D face model (vertex positions, texture coordinates, material properties) to transform expressions while maintaining feature consistency. By adjusting geometric and textural parameters rather than regenerating entire features, the system achieves realistic expressions with controlled complexity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11798261B2Image face manipulation
Publication Date: 2023.10.24 SNAP INC
  • US11798261B2 patent drawing
  • US11798261B2 patent drawing
  • US11798261B2 patent drawing

AI summary

Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and a method for synthesizing a realistic image with a new expression of a face in an input image by receiving an input image comprising a face having a first expression; obtaining a target expression for the face; and extracting a texture of the face and a shape of the face. The program and method for generating, based on the extracted texture of the face, a target texture corresponding to the obtained target expression using a first machine learning technique; generating, based on the extracted shape of the face, a target shape corresponding to the obtained target expression using a second machine learning technique; and combining the generated target texture and generated target shape into an output image comprising the face having a second expression corresponding to the obtained target expression.