3D Facial Reconstruction Using Multi-View Neural Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional 3D facial reconstruction technologies suffer from low reconstruction accuracy and poor expression discrimination, limiting their effectiveness in facial recognition and expression tracking applications.

Innovation Solution

A deep learning-based three-dimensional facial reconstruction system employing a main color range camera, auxiliary color cameras, and a processor to generate 3D ground truth models and train artificial neural networks, enhancing accuracy by capturing images from multiple angles and using 3D morphable models to predict face shape and expression coefficients.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional 3D facial reconstruction technology is used, then the system is simpler, but reconstruction accuracy is low

Engineering Contradiction:
Improvereconstruction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transitions from 2D facial images to 3D facial reconstruction by introducing depth information through auxiliary cameras positioned at multiple angles. This dimensional transformation enables accurate 3D face shape and expression coefficient prediction, directly resolving the low reconstruction accuracy problem while managing system complexity through structured multi-view geometry.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces a 3D morphable model as an intermediary between the captured multi-view images and the final 3D facial reconstruction. This model serves as a bridge that integrates color and depth information from multiple cameras, enabling accurate prediction of face shape and expression coefficients while maintaining system coherence.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If conventional expression tracking technology is used, then the device is simpler, but expression discrimination is poor

Engineering Contradiction:
Improveexpression discriminationVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the facial reconstruction process into distinct components: face shape coefficient prediction and expression coefficient prediction. By separately modeling these aspects using the 3D morphable model and training them independently with ground truth data from multiple camera views, the system achieves high expression discrimination capability while maintaining manageable complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent replaces conventional mechanical or signal-processing-based expression tracking with a deep learning-based artificial neural network system. The trained model automatically learns expression patterns from multi-view training images, substituting complex signal processing pipelines with a unified neural network approach that achieves superior expression discrimination.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If additional signal processing is used to improve accuracy, then reconstruction accuracy improves, but processing time increases

Engineering Contradiction:
Improvereconstruction accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by pre-training the artificial neural network model with ground truth 3D facial data captured from multiple cameras during a training phase. Once trained, the model can directly predict face shape and expression coefficients from single or multiple input images without requiring additional iterative signal processing, thus achieving high accuracy while reducing real-time processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a learned copy of the complex relationship between multi-view images and 3D facial parameters through the trained neural network model. Instead of repeatedly performing complex signal processing operations, the system uses the trained model to directly map input images to output coefficients, significantly reducing processing time while maintaining high reconstruction accuracy.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11321960B2Deep learning-based three-dimensional facial reconstruction system
Publication Date: 2022.05.03 ARCSOFT CORP LTD
  • US11321960B2 patent drawing
  • US11321960B2 patent drawing
  • US11321960B2 patent drawing

AI summary

A 3D facial reconstruction system includes a main color range camera, a plurality of auxiliary color cameras, a processor and a memory. The main color range camera is arranged at a front angle of a reference user to capture a main color image and a main depth map of the reference user. The plurality of auxiliary color cameras are arranged at a plurality of side angles of the reference user to capture a plurality of auxiliary color images of the reference user. The processor executes instructions stored in the memory to generate a 3D front angle image according to the main color image and the main depth map, generate 3D side angle images according to the 3D front angle image and the plurality of auxiliary color images, and train an artificial neural network model according to a training image, the 3D front angle image and 3D side angle images.