Head-Mounted Display Facial Expression Detection Using Segmented Cameras
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current virtual reality and augmented reality systems lack effective real-time facial expression detection capabilities, limiting the ability to convey emotional states and enhance user interaction through 3D representations.
Innovation Solution
A head-mounted display unit equipped with multiple image capturing devices, such as infrared and depth cameras, processes images to extract facial expression parameters, generating a graphical representation of the user's face by detecting landmark locations and applying a blendshape model, enabling real-time facial expression tracking and display in VR/AR environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple image capturing devices are used to detect facial expressions, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The face detection task is segmented into multiple regions (upper face with eyes/eyebrows captured by first image capturing device, lower face captured by second image capturing device). Each device focuses on specific facial landmarks, improving measurement precision while distributing the complexity across specialized components rather than requiring one device to capture the entire face.
2Speed
If real-time processing of facial images is performed, then speed is improved, but use of energy increases
Solution Approach 1:
The system extracts only the essential facial expression parameters from the captured images rather than processing the entire image data. By identifying and tracking specific landmark locations (eyes, eyebrows, lower face features) and applying blendshape models, the system achieves real-time performance with reduced computational energy consumption compared to full-face analysis.
3Measurement precision
If personalized calibration is performed for each user, then measurement precision is improved, but loss of time increases
Solution Approach 1:
The system performs preliminary calibration by capturing images of the user's neutral face expression before actual facial expression tracking begins. This preliminary action establishes a baseline for personalized blendshape model fitting, which improves subsequent measurement precision. The calibration process is streamlined by focusing only on neutral expression capture rather than requiring extensive multi-expression training data.
Data Source
AI summary
Embodiments relate to detecting a user's facial expressions in real-time using a head-mounted display unit that includes a 2D camera (e.g., infrared camera) that capture a user's eye region, and a depth camera or another 2D camera that captures the user's lower facial features including lips, chin and cheek. The images captured by the first and second camera are processed to extract parameters associated with facial expressions. The parameters can be sent or processed so that the user's digital representation including the facial expression can be obtained.


