Blendshape Weight Prediction Using Rendered Avatar Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The existing methods for training machine learning models to predict blendshape weights for facial expressions in HMD wearers are labor-intensive and prone to inaccuracies due to the need for manual labeling of large numbers of HMD-captured training images, which are time-consuming and subjective.
Innovation Solution
Training a machine learning model using rendered avatar training images with specified blendshape weights, eliminating the need for manual labeling and allowing for faster and more accurate prediction of blendshape weights for facial expressions, which are then applied to render a corresponding facial avatar.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If manual labeling of HMD-captured training images is used, then the model can be trained to predict blendshape weights, but the process becomes labor-intensive and time-consuming
Solution Approach 1:
The patent uses rendered avatar images as synthetic training data instead of manually labeling real HMD-captured images. The rendering process generates images with known ground truth blendshape weights, eliminating manual labeling while providing sufficient training data for accurate prediction models
Solution Approach 2:
The patent pre-renders a large dataset of avatar images with known blendshape weights before training begins. This preliminary generation of training data with accurate labels eliminates the need for time-consuming manual annotation during the actual model development process
2Quantity of substance
If manual labeling of training images is performed, then training data can be obtained, but the process becomes subjective and prone to inaccuracies
Solution Approach 1:
Instead of manually labeling real images subjectively, the system creates synthetic copies through rendering. The rendered avatar images come with automatically generated ground truth labels from the rendering process itself, eliminating human subjectivity and improving labeling accuracy
Solution Approach 2:
The rendering system automatically generates both the training images and their corresponding ground truth labels without human intervention. The blendshape weights are inherently known from the rendering parameters, making the system self-labeling and eliminating manual annotation errors
3Reliability
If large numbers of HMD-captured training images are collected and labeled, then the model can achieve good prediction performance, but the process becomes labor-intensive
Solution Approach 1:
The patent generates synthetic training data through avatar rendering, creating large datasets without the labor of collecting and labeling real images. The rendered images maintain sufficient realism for training while eliminating the productivity bottleneck of manual data preparation
Solution Approach 2:
The system varies rendering parameters (blendshape weights, lighting, camera angles) to generate diverse training images automatically. This parameter-driven approach creates large datasets efficiently by systematically exploring the feature space without manual intervention
Data Source
AI summary
Avatar training images of facial avatars having facial expressions corresponding to specified blendshape weights are rendered. A two-stage machine learning model is trained based on the rendered avatar training images and the specified blendshape weights. The machine learning model has a first stage extracting image features from the rendered avatar training images and a second stage predicting blendshape weights from the extracted image features. The trained machine learning model is applied to predict the blendshape weights for a facial expression of a wearer of a head-mountable display (HMD) from a set of images captured by the HMD of a face of the wearer when exhibiting the facial expression.


