Face Reconstruction via Mesh Convolution Network
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing facial capture systems require controlled settings and the physical presence of individuals, limiting their ability to perform facial reconstruction under uncontrolled conditions, such as arbitrary human identities and facial expressions, and lack the capability to handle images with undetermined lighting environments.
Innovation Solution
The techniques involve generating an identity mesh and an expression mesh based on respective encodings that represent the identity and expression of a face in images, and using a machine learning model to combine these meshes and produce an output mesh that accurately reconstructs the face, even from limited or uncontrolled image data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional facial capture systems are used, then high-quality 3D face reconstruction is achieved, but controlled settings and physical presence of individuals are required
Solution Approach 1:
The patent uses 2D images as copies of the target face to create 3D reconstructions, eliminating the need for physical presence. The system learns to map 2D image features to 3D geometry through training on paired datasets, enabling reconstruction from arbitrary images without requiring the subject to be physically present in a controlled environment.
Solution Approach 2:
The patent replaces the mechanical optical system (light stage, multiple cameras, controlled lighting) with a data-driven machine learning approach. A neural network is trained to directly predict 3D face geometry from 2D images, substituting complex physical capture infrastructure with an intelligent algorithm that can process images from any source.
2Measurement precision
If multiple calibrated camera views and controlled lighting are used, then accurate 3D geometry is captured, but device complexity and setup requirements increase
Solution Approach 1:
The patent extracts only the essential information needed for 3D reconstruction from images, rather than requiring a complex array of sensors and cameras. The neural network learns to extract depth, geometry, and shape information directly from standard 2D images, eliminating the need for specialized capture devices, multiple cameras, or calibrated setups.
Solution Approach 2:
The patent creates a universal reconstruction system that can process images from any camera or image source, not just specialized capture equipment. The trained model works with images from mobile phones, webcams, legacy footage, or any other image source, making the system universally applicable without requiring specific hardware configurations.
3Adaptability or versatility
If legacy footage with single viewpoint and undetermined lighting is used, then reconstruction of historical figures is possible, but conventional techniques fail
Solution Approach 1:
The patent performs preliminary training on large datasets of paired 2D-3D face images to equip the neural network with robust knowledge of face geometry under various conditions. This pre-training enables the model to handle challenging inputs like legacy footage, single-viewpoint images, and undetermined lighting, achieving accurate reconstruction without requiring the input images to meet specific quality criteria.
Data Source
AI summary
Embodiment of the present invention sets forth techniques for performing face reconstruction. The techniques include generating an identity mesh based on an identity encoding that represents an identity associated with a face in one or more images. The techniques also include generating an expression mesh based on an expression encoding that represents an expression associated with the face in the one or more images. The techniques also include generating, by a machine learning model, an output mesh of the face based on the identity mesh and the expression mesh.


