Virtual Facial Model Expression Simulation via Acoustic Sentiment and Image Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for simulating facial expressions in virtual reality face challenges due to low identification accuracy and high hardware costs associated with using multiple sensors in head-mounted displays, which also disrupt the integration of upper and lower half facial expressions.
Innovation Solution
A system and method that uses a processor and storage to identify user sentiment through acoustic signals, select a corresponding three-dimensional facial model, predict the upper half face image from the lower half face image, and combine them to generate feature relationships for accurate facial expression simulation without additional sensors.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple sensors (three-dimensional sensors, infrared sensors, EMG sensors, EOG sensors) are disposed in head-mounted display to detect upper half face muscle changes, then facial expression detection capability is improved, but hardware cost increases and device complexity increases
Solution Approach 1:
The patent extracts the facial expression detection function from the head-mounted display sensors and relocates it to a mobile device. The mobile device captures lower half face images and acoustic signals independently, separating the detection function from the VR hardware to reduce device complexity and cost.
Solution Approach 2:
The patent creates a virtual copy of the upper half face through image prediction algorithms. Instead of using physical sensors to detect upper half face expressions, the system generates a predicted upper half face image that mirrors the lower half face expressions, providing a cost-free solution.
2Measurement precision
If multiple sensors are used to detect upper half face expressions, then facial expression simulation accuracy is improved, but integration with lower half face expression becomes difficult and system complexity increases
Solution Approach 1:
The patent merges the upper half face prediction results with the lower half face image analysis in a unified processing framework. Both expressions are processed together through the same neural network model, ensuring consistent integration without requiring separate processing pipelines.
Solution Approach 2:
Instead of using sensors to detect upper half face expressions and then integrating with lower half face images, the patent inverts the approach by using lower half face images to predict and generate upper half face expressions, simplifying the integration process.
3Adaptability or versatility
If head-mounted display covers upper half face to enable virtual reality application, then virtual reality immersion is improved, but facial expression identification accuracy deteriorates
Solution Approach 1:
The patent introduces acoustic signals as an intermediary to bridge the gap caused by HMD coverage. The voice-based sentiment analysis serves as a mediator that provides additional information about upper half face expressions, compensating for the blocked visual field.
Solution Approach 2:
The patent transitions from purely visual 2D image analysis to a multi-dimensional approach by incorporating acoustic signal analysis. This adds a temporal and spectral dimension to expression detection, enabling accurate identification even when the visual field is partially blocked.
Data Source
AI summary
A system and method for simulating facial expression of a virtual facial model are provided. The system stores a plurality of three-dimensional facial models corresponding to a plurality of preset sentiments one-to-one. The system identifies a present sentiment according to an acoustic signal and selects a selected model from the three-dimensional facial models according to the present sentiment, wherein the preset sentiment corresponding to the selected model is same as the present sentiment. The system predicts an upper half face image according to a lower half face image, combines the lower half face image and the upper half face image to form a whole face image, and generates a plurality of feature relationships by matching the facial features of the whole face image with the facial features of the selected model so that a virtual facial model can simulate an expression based on the feature relationships.


