Neural Network Post-Processing for Avatar Facial Expression Fidelity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing 3-D rendering methods for avatars result in poor-quality facial expressions due to limited resources, and traditional methods like Blendshapes require manual rebuilding for each person, leading to unrealistic face movements and high computational costs.

Innovation Solution

A system utilizing a neural network (NN) to post-process classically rendered avatars, allowing for high-fidelity facial expressions that preserve individual identity without requiring high computational power or manual rebuilding of Blendshapes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional 3-D rendering methods are used for avatars, then computational resources are limited and processing is efficient, but facial expression quality is poor

Engineering Contradiction:
Improvefacial expression qualityVSAvoidcomputational resources
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The system separates the avatar rendering process into two independent stages: (1) classical 3-D rendering to generate base avatar images, and (2) neural network post-processing to enhance facial expressions. This segmentation allows each stage to optimize for its specific function, with the NN stage focusing solely on expression quality using computational resources efficiently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A neural network serves as an intermediary between the classical rendering pipeline and the final output. The NN takes classically rendered avatars as input and produces enhanced versions with high-fidelity facial expressions, acting as a mediator that bridges the gap between computational efficiency and expression quality.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If Blendshapes are used for avatar animation, then facial expressions can be generated, but manual rebuilding is required for each person and face movements appear unrealistic

Engineering Contradiction:
Improveavatar customizationVSAvoidfacial expression realism
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

Instead of manually creating Blendshapes for each person, the system uses a neural network trained on diverse facial data to automatically generate and copy realistic facial expressions. The NN learns from training data and can replicate natural facial movements and expressions for any avatar, eliminating the need for manual reconstruction while maintaining high realism.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If manual shader crafting is performed for each avatar, then facial expression quality can be improved, but computational power requirements increase and time consumption increases

Engineering Contradiction:
Improvefacial expression accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system changes the approach from manually adjusting shader parameters for each avatar to using a pre-trained neural network that automatically adjusts expression parameters. The NN has learned optimal parameter transformations from training data, allowing it to rapidly generate accurate facial expressions without manual shader crafting or high computational power requirements during inference.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12322020B1System apparatus and method for providing facial expression to avatars
Publication Date: 2025.06.03 GOODSIZE INC
  • US12322020B1 patent drawing
  • US12322020B1 patent drawing
  • US12322020B1 patent drawing

AI summary

A system and method for providing a facial expression to a virtual avatar. The system includes a training system to train a neural network system to replace a face of the virtual avatar with a source face and to provide a facial expression of the source face to the face of the avatar in a real-time and an inference system configured to use the trained neural network system to provide one or more facial expressions of the source face to the face of the avatar in real-time to cause the one or more facial expressions of the avatar to imitate approximately in an exact manner the one or more facial expressions of a source face media which is represented by the virtual avatar.