Monocular Avatar Generation Using Morphable Model Vertex Displacement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing videoconference systems face challenges in generating realistic avatars with facial expressions using monocular cameras, as they struggle with generating three-dimensional models efficiently and cost-effectively, often requiring expensive depth cameras and multiple camera setups.
Innovation Solution
A computing system generates avatars based on monocular images using a three-dimensional morphable model with vertices in UV space, displacing these vertices to represent facial expressions, allowing for the use of inexpensive cameras like webcams and enabling photorealistic three-dimensional avatar representation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If multiple cameras and depth cameras are used to generate realistic avatars with facial expressions, then the realism and accuracy of the avatar is improved, but the hardware expense and system complexity increases
Solution Approach 1:
The patent uses a pre-built three-dimensional morphable model as a template or copy of a human face, which can be deformed and adjusted to match the user's facial expressions. Instead of capturing complete 3D geometry from multiple cameras, the system copies the structure of the morphable model and modifies it based on monocular image inputs, significantly reducing hardware requirements while maintaining avatar realism
Solution Approach 2:
The patent replaces the mechanical/optical system of multiple physical cameras and depth sensors with a computational approach using a morphable model and image processing algorithms. The system substitutes physical 3D capture hardware with a mathematical model that can be manipulated computationally to achieve the same avatar generation goal
2Manufacturing precision
If multiple cameras and depth cameras are used to generate realistic avatars with facial expressions, then the realism and accuracy of the avatar is improved, but the hardware cost increases
Solution Approach 1:
The patent replaces expensive, sophisticated camera systems with inexpensive monocular cameras (such as standard webcams). The system accepts that the input images may be of lower quality but compensates through sophisticated processing of the morphable model, achieving high-quality avatar output from low-cost input hardware
Solution Approach 2:
The system uses a pre-existing morphable model as a reusable template that can be applied to multiple users and sessions. This model serves as a persistent asset that eliminates the need for expensive repeated 3D scanning hardware, allowing the same computational model to generate avatars for different users using only their monocular camera inputs
3Ease of manufacture
If monocular cameras are used to generate three-dimensional avatar models, then the hardware cost and complexity is reduced, but the ability to accurately capture three-dimensional facial geometry deteriorates
Solution Approach 1:
The patent introduces the morphable model as an intermediary between the monocular camera input and the final avatar output. The model acts as a bridge that translates 2D image data into 3D facial geometry by mapping observed facial features onto the pre-defined morphable model structure, enabling accurate 3D reconstruction from limited 2D inputs
Solution Approach 2:
The system changes the parameters of the morphable model (such as shape coefficients, expression coefficients, and texture parameters) based on analysis of the monocular images. By adjusting these parameters rather than directly reconstructing raw 3D geometry from images, the system can accurately represent facial expressions and features using only single-camera input
Data Source
AI summary
A method comprises receiving a first sequence of images of a portion of a user, the first sequence of images being monocular images; generating an avatar based on the first sequence of images, the avatar being based on a model including a feature vector associated with a vertex; receiving a second sequence of images of the portion of the user; and based on the second sequence of images, modifying the avatar with a displacement of the vertex to represent a gesture of the avatar.


