3D Facial Expression Modeling Using Disparity Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for capturing and processing human geometry, appearance, and motion in computer animation fail to accurately predict the weights representing facial expression details, leading to inaccurate digital actor modeling.

Innovation Solution

A facial expression modeling method that uses two-dimensional images processed through a deep learning process to generate a disparity map, which is then concatenated with the images to create a three-channel feature map, subsequently processed by a weighting calculation neural network to generate blend-shape weightings for accurate three-dimensional facial expression modeling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing methods are used for capturing and processing human geometry and appearance, then the process is simpler, but the accuracy of predicting facial expression weights deteriorates

Engineering Contradiction:
Improveaccuracy of predicting facial expression weightsVSAvoidcomplexity of modeling process
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent transitions from traditional 2D facial image processing to 3D disparity map generation, adding depth information as a new dimension. This dimensional enhancement enables more accurate prediction of facial expression weights by incorporating spatial depth relationships, directly resolving the contradiction between measurement precision and device complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent combines multiple data sources (2D facial images from different views, disparity maps, and expression weight predictions) into a composite modeling framework. This composite approach integrates various information types to achieve higher prediction accuracy while managing system complexity through unified processing.

Inventive Principle:
Principle #40Composite materials

2Manufacturing precision

If traditional 3D scanning and appearance acquisition methods are used, then the workflow is more straightforward, but the detail accuracy of facial expressions deteriorates

Engineering Contradiction:
Improvedetail accuracy of facial expressionVSAvoidease of digital actor creation
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent replaces traditional mechanical 3D scanning methods with a computational approach using deep learning networks. The system substitutes physical scanning equipment with image processing algorithms that generate disparity maps and predict expression weights, achieving higher detail accuracy while simplifying the manufacturing workflow.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates accurate digital copies of facial expressions by generating synthetic disparity maps and expression weight predictions from 2D images. This copying mechanism reproduces detailed facial geometry and expression characteristics without requiring complex physical scanning, improving both accuracy and ease of creation.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10885702B2Facial expression modeling method, apparatus and non-transitory computer readable medium of the same
Publication Date: 2021.01.05 HTC CORP
  • US10885702B2 patent drawing
  • US10885702B2 patent drawing

AI summary

A facial expression modeling method used in a facial expression modeling apparatus is provided that includes the steps outlined below. Two two-dimensional images of a facial expression retrieved by two image retrieving modules respectively are received. A deep learning process is performed on the two two-dimensional images to generate a disparity map. The two two-dimensional images and the disparity map are concatenated to generate a three-channel feature map. The three-channel feature map is processed by a weighting calculation neural network to generate a plurality of blend-shape weightings. A three-dimensional facial expression is modeled according to the blend-shape weightings.