Normalized 3D Face Synthesis from Single Image

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for creating high-quality normalized three-dimensional avatars from unconstrained two-dimensional images face challenges in producing consistent, neutralized facial expressions and lighting conditions, requiring large datasets and sophisticated equipment, which limits their applicability in consumer applications.

Innovation Solution

A two-stage deep learning framework using a generative adversarial network (GAN) with a non-linear morphable face model and iterative perceptual refinement, trained on a normalized face dataset with neutral expressions and diffuse lighting, enables the generation of high-quality textured three-dimensional face models from a single unconstrained image.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional linear 3DMM methods are used for avatar digitization, then the process is simpler and requires fewer training samples, but the generated models lack details and likeness of the original subject

Engineering Contradiction:
Improvemodel quality and likenessVSAvoidtraining data volume
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent transitions from linear 3DMM parameterization to non-linear 3DMM parameterization, fundamentally changing the mathematical model's parameters and structure. This enables the system to capture complex facial variations and achieve high-fidelity reconstructions with fewer training samples by learning non-linear mappings from 2D images to 3D face parameters.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent replaces traditional manual or semi-automatic 3D scanning equipment with an automated deep learning-based system. The mechanical/optical scanning process is substituted with a computational approach using GANs and non-linear 3DMM, enabling avatar creation from simple 2D photographs without requiring specialized capture equipment.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Manufacturing precision

If GAN-based non-linear 3DMM methods are used, then high-quality detailed face models can be recovered, but hundreds of thousands of training subjects are required

Engineering Contradiction:
Improveface model detail qualityVSAvoiddata collection infrastructure
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent pre-trains the GAN model on a large dataset of facial images to learn general facial structures and variations. This preliminary training phase enables the model to capture essential facial features and lighting conditions, which are then transferred to the non-linear 3DMM parameterization, reducing the need for extensive subject-specific training data.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces a non-linear 3DMM as an intermediary representation between the GAN-generated images and the final 3D avatar. This intermediary model serves as a bridge that connects the powerful generative capabilities of GANs with the structured parameterization needed for 3D face analysis, enabling high-quality reconstruction with fewer training samples.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Adaptability or versatility

If images are taken under various lighting conditions and expressions, then more diverse training data is obtained, but the generated textures integrate environmental lighting making unshaded albedo extraction difficult

Engineering Contradiction:
Improvetraining data diversityVSAvoidalbedo texture accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent extracts and separates the lighting information from the albedo (intrinsic color) information in the input images. By using the non-linear 3DMM framework, the system decomposes the observed image into components including lighting conditions, facial expressions, and the underlying albedo texture, enabling accurate recovery of unshaded face textures even from images taken under varying environmental lighting.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of trying to capture faces under controlled lighting conditions, the patent inverts the approach by training the model on diverse lighting conditions and then learning to remove or normalize these lighting effects. The non-linear 3DMM is trained to infer the true albedo by inverting the lighting transformations present in the training data, enabling robust texture extraction from unconstrained images.

Inventive Principle:
Principle #13The other way round (Inversion)

4Manufacturing precision

If controlled capture environments with neutral expressions are used, then consistent training data is obtained, but the process requires sophisticated equipment and professional production studios

Engineering Contradiction:
Improvetraining data consistencyVSAvoidcapture equipment sophistication
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical and optical capture systems with a purely computational approach. Instead of using controlled environments, specialized cameras, and professional scanning equipment, the system uses standard 2D digital photographs combined with deep learning algorithms to achieve high-quality 3D avatar reconstruction, making the technology accessible for consumer applications.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12033277B2Normalized three-dimensional avatar synthesis and perceptual refinement
Publication Date: 2024.07.09 PINSCREEN INC
  • US12033277B2 patent drawing
  • US12033277B2 patent drawing
  • US12033277B2 patent drawing

AI summary

A system, method, and apparatus for generating a normalized three-dimensional model of a human face from a single unconstrained two-dimensional image of the human face. The system includes a processor that executes instructions including receiving the single unconstrained two-dimensional image of the human face, using an inference network to determine an inferred normalized three-dimensional model of the human face based on the single unconstrained two-dimensional image of the human face, and using a refinement network to iteratively determine the normalized three-dimensional model of the human face with a neutral expression and unshaded albedo textures under diffuse lighting conditions based on the inferred normalized three-dimensional model of the human face.