Avatar Facial Expression Generation via Audio Phoneme Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing virtual reality and augmented reality techniques for training sessions lack dynamic and compatible facial expressions across different sessions, as they are specific to the scanned session and not adaptable in real-time.
Innovation Solution
A method and system that generate facial expressions for avatars in virtual environments by extracting voice and text features, identifying phonemes, and determining facial features using pre-trained learning models, allowing for real-time generation of expressions based on speech and previously generated features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If facial movements and expressions are ported to the avatar through scanning during a training session, then the avatar can display realistic facial expressions for that specific session, but the facial expressions cannot be used for other training sessions of the same trainer
Solution Approach 1:
The system performs preliminary action by training the deep learning model on the trainer's facial characteristics in advance, before any training sessions occur. This pre-training enables the model to generate accurate facial expressions for the avatar across multiple sessions without requiring repeated scanning or retraining, thus resolving the contradiction between session compatibility and time loss.
2Adaptability or versatility
If facial expressions are generated using traditional scanning techniques, then the expressions are specific to the scanned session, but the system lacks dynamic adaptation for real-time use in different sessions
Solution Approach 1:
The system replaces the mechanical scanning process with a deep learning-based computational approach. Instead of using complex scanning hardware and manual porting procedures, the system uses a trained neural network that processes audio input to generate facial expressions dynamically, reducing device complexity while improving adaptability across sessions.
Solution Approach 2:
The deep learning model acts as an intermediary between the audio input and the avatar's facial expression output. This intermediary component translates speech audio directly into corresponding facial expressions without requiring direct scanning or manual intervention, enabling dynamic adaptation while maintaining system simplicity.
3Ease of operation
If pre-recorded digital sessions are used as an alternative to online training sessions, then availability issues are resolved, but the immersive experience and real-time interaction are lost
Solution Approach 1:
The system introduces dynamics by enabling real-time generation of facial expressions based on audio input, transforming static pre-recorded sessions into dynamic, interactive experiences. The avatar's facial expressions change dynamically in response to the trainer's speech, maintaining immersive quality while preserving the availability benefits of pre-recorded content.
Data Source
AI summary
The present invention relates to a method of generating a facial expression of a user for a virtual environment. The method comprises obtaining a video and an associated speech of the user. Further, extracting in real-time at least one of one or more voice features and one or more text features based on the speech. Furthermore, identifying one or more phonemes in the speech. Thereafter, determining one or more facial features relating to the speech of the user using a pre-trained second learning model based on the one or more voice features, the one or more phonemes, the video and one or more previously generated facial features of the user. Finally, generating the facial expression of the user corresponding to the speech for an avatar representing the user in the virtual environment.


