Anatomically-Constrained Local Model for Facial Performance Capture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current facial performance capture techniques face challenges in achieving high-quality reconstructions with less-constrained acquisition setups, often requiring extensive hardware and manual stabilization, and struggle to accurately capture facial expressions outside the pre-defined blendshape rig, leading to unstable results and reduced robustness.
Innovation Solution
The development of an anatomically-constrained local subspace model that combines a local shape subspace and an anatomical subspace, using anatomical bone structures to constrain facial deformations, allowing for high-quality facial performance capture from a single camera view with automatic rigid stabilization and reduced need for pre-acquired expressions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional global blendshape rig is used for facial performance capture, then high-quality reconstruction can be achieved, but extensive hardware and manual stabilization are required, and the setup is highly constrained
Solution Approach 1:
The face is divided into multiple local patches, each with its own shape subspace. This segmentation allows the system to capture facial expressions locally without requiring a complex global rig, reducing hardware constraints while maintaining reconstruction quality.
Solution Approach 2:
The patent changes from using a fixed global blendshape rig to using learnable local shape subspaces with anatomical constraints. This parameter change enables the system to adapt to different facial geometries and expressions dynamically, reducing the need for extensive pre-acquired expressions and manual stabilization.
2Ease of operation
If less-constrained acquisition setup is used to give subject more freedom, then ease of operation improves, but measurement precision and reconstruction quality deteriorate
Solution Approach 1:
The system uses iterative optimization with energy minimization to refine patch positions and shape coefficients. This feedback mechanism ensures that even with less-constrained acquisition setups, the reconstructed facial expressions remain accurate by continuously adjusting parameters to match observed motion data.
Solution Approach 2:
The patent employs dynamic optimization that adapts to each frame's motion data. The system dynamically adjusts patch configurations and shape subspaces based on observed facial movements, allowing subjects greater freedom while maintaining capture accuracy through real-time parameter optimization.
3Adaptability or versatility
If local models are used to increase flexibility, then adaptability improves, but robustness and stability deteriorate
Solution Approach 1:
Each local patch has its own shape subspace tailored to specific facial regions, providing local adaptability. Meanwhile, anatomical constraints ensure that these local variations remain physically plausible, maintaining overall reconstruction stability and robustness.
Solution Approach 2:
The patent combines multiple components (local shape subspaces, anatomical constraints, rigid transformations) into a composite model. This composite structure provides both the flexibility of local adaptation and the stability of anatomical grounding, resolving the contradiction between adaptability and robustness.
4Manufacturing precision
If extensive pre-acquired expressions are used to build rigid, then manufacturing precision improves, but loss of time and productivity worsen
Solution Approach 1:
The system performs preliminary learning of shape subspaces from a small set of training expressions, rather than requiring extensive pre-acquired expressions. This preliminary action creates a foundation that can be efficiently optimized during performance capture, reducing acquisition time while maintaining rig quality.
Solution Approach 2:
The optimization process uses the observed motion data itself to refine the model parameters during capture. This self-service approach eliminates the need for extensive pre-acquisition by allowing the system to learn and adapt during the actual performance capture session.
Data Source
AI summary
Techniques and systems are described for generating an anatomically-constrained local model and for performing performance capture using the model. The local model includes a local shape subspace and an anatomical subspace. In one example, the local shape subspace constrains local deformation of various patches that represent the geometry of a subject's face. In the same example, the anatomical subspace includes an anatomical bone structure, and can be used to constrain movement and deformation of the patches globally on the subject's face. The anatomically-constrained local face model and performance capture technique can be used to track three-dimensional faces or other parts of a subject from motion data in a high-quality manner. Local model parameters that best describe the observed motion of the subject's physical deformations (e.g., facial expressions) under the given constraints are estimated through optimization. The optimization can solve for rigid local patch motion, local patch deformation, and the rigid motion of the anatomical bones. The solution can be formulated as an energy minimization problem for each frame that is obtained for performance capture.


