Avatar Training Instances for Hard Facial Expression Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing extended reality (XR) technologies struggle to accurately render avatars with lifelike facial expressions that correspond to the wearer's real-time facial movements, particularly due to variations in facial features and expressions among users, leading to disjointed and unnatural representations.
Innovation Solution
A machine learning model is trained and refined using diverse avatars to predict facial action units from captured images, with preprocessing and postprocessing to enhance accuracy, and iteratively improved through generating new avatars to address specific prediction challenges.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a machine learning model is trained using diverse avatars with different facial features, then the accuracy of facial expression prediction is improved, but the complexity of the training process and data requirements increase
Solution Approach 1:
The patent creates synthetic avatar images that replicate diverse facial features and expressions. These synthetic avatars serve as copies of real human faces, allowing the model to learn from a controlled set of diverse examples without requiring actual photos of every possible face type, thus reducing data collection complexity while maintaining prediction accuracy across different demographics
Solution Approach 2:
The patent pre-processes and curates a comprehensive set of diverse avatar images before training the machine learning model. By preparing the training data in advance with varied facial features, expressions, and demographics, the system avoids the complexity of handling unprocessed real-time data during the training phase, streamlining the overall training process
2Reliability
If real-time facial expression rendering is implemented, then user immersion in XR environments is enhanced, but the computational resources and processing time required increase
Solution Approach 1:
The patent replaces complex real-time 3D facial rendering mechanisms with a machine learning-based image processing system. Instead of using computationally intensive 3D avatar manipulation, the system uses pre-trained neural networks to predict facial expressions from 2D images and generate corresponding avatar frames, significantly reducing computational energy requirements while maintaining rendering realism
Solution Approach 2:
The patent pre-trains machine learning models on diverse facial expressions and avatars before actual use. This preliminary training allows the system to make rapid predictions during real-time operation without requiring heavy computational resources, as the complex pattern recognition has already been performed during the training phase
3Manufacturing precision
If facial action units are predicted and smoothed to enhance expression accuracy, then the lifelike quality of avatars is improved, but the processing time and computational complexity increase
Solution Approach 1:
The patent performs facial action unit prediction and smoothing operations during the training phase rather than only during real-time inference. By pre-computing and pre-smoothing facial action units for diverse expressions and storing them in the training data, the system reduces the processing time required during actual avatar rendering, as the complex computations have already been performed in advance
Data Source
AI summary
For each avatar, testing images are rendered for different facial expressions that each have ground truth facial action units. An instance of a machine learning model is applied to the testing images to generate predicted facial action units for each testing image. A predictive performance of the instance is calculated for each avatar based on the predicted and ground truth facial action units for the testing images of the avatar. A first set of features common to the avatars for which the predictive performance was better than a first threshold, and a second set of features common to the avatars for which the predictive performance was worse than a second threshold, are identified. The features present only in the second set are identified, as difference features. New avatars having the difference features are generated.WO


