Pretrained Visual Encoder for Identity-Invariant 3D Facial Animation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems for generating 3D facial models and animations struggle with efficiency and scalability, requiring expensive equipment and multiple cameras in controlled environments, and often fail to generalize across different identities, poses, and lighting conditions.

Innovation Solution

The development of a machine learning model that uses a pretrained visual encoder to extract and encode facial expressions into a latent code, which remains constant across different view angles, identities, and color styles, allowing for efficient generation and animation of 3D facial models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional systems use multiple cameras and controlled environments to generate 3D facial models, then manufacturing precision and reliability are improved, but device complexity and cost increase

Engineering Contradiction:
Improve3D facial model qualityVSAvoidcamera system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent replaces the mechanical camera-based 3D scanning system with a machine learning model that processes 2D images. Instead of using multiple physical cameras to capture depth information, the system uses a pretrained visual encoder to extract facial features and generate 3D models from standard 2D images, eliminating the need for complex multi-camera setups while maintaining model quality

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a virtual copy of the facial expression through latent code representation. The machine learning model encodes facial expressions into a compressed latent space that captures essential features, allowing the system to reconstruct and animate 3D facial models from this encoded representation rather than requiring direct physical capture of expressions

Inventive Principle:
Principle #26Copying

2Measurement precision

If traditional systems require controlled environments and extensive training data, then measurement precision is improved, but loss of time and productivity decrease

Engineering Contradiction:
Improvefacial expression accuracyVSAvoidmodel generation speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies preliminary action by pretraining the visual encoder on large datasets of facial images and expressions before deployment. The model is pre-trained to recognize and encode various facial expressions, view angles, and lighting conditions, allowing it to generalize well to new inputs without requiring controlled environments or extensive additional training data during actual use

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent transforms the facial expression data into a different parameter space through latent code encoding. Instead of working directly with raw pixel data or complex 3D coordinates, the system encodes expressions into a compressed latent representation that captures essential features, enabling faster processing and generation while maintaining accuracy

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If the system uses identity-specific training data, then manufacturing precision for specific identities is improved, but adaptability to new identities decreases

Engineering Contradiction:
Improveidentity-specific model accuracyVSAvoidgeneralization across identities
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements universality by designing a machine learning model that processes facial images in a view-angle and identity-invariant manner. The pretrained visual encoder learns to extract facial expression features that are consistent across different identities, view angles, and lighting conditions, allowing the same model to accurately generate and animate 3D facial models for any identity without requiring identity-specific training or adjustment

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250095259A1Avatar animation with general pretrained facial movement encoding
Publication Date: 2025.03.20 QUALCOMM INC
  • US20250095259A1 patent drawing
  • US20250095259A1 patent drawing
  • US20250095259A1 patent drawing

AI summary

Techniques and systems are provided for generating a representation of a face. For instance, a process can include obtaining one or more images of a face. The process can further include generating an encoded expression representing an expression of the face, wherein predetermined characteristics of the face remain constant relative to the encoded expression. The process can further include mapping the encoded expression to a corresponding expression of a facial model. The process can further include generating the representation of the facial model based on the encoded expression.