3D Facial Capture via Coupled Photometric and Multi-View Stereo

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer vision techniques for reconstructing high-detail 3D geometry and appearances of real objects face limitations, such as integration drift in photometric stereo, matching errors in multi-view stereo, and residual misalignment due to object motion and illumination changes, which affect the accuracy and detail of 3D facial expressions capture.

Innovation Solution

A method that simultaneously computes photometric stereo, multi-view stereo, and optical flow in a coupled manner, using a Lambertian model with five degrees of freedom and color- and time-multiplexed illumination, with RGB lights positioned in an array around the target object, and iteratively reconstructs the 3D surface by aligning images using motion estimation until convergence.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If photometric stereo and multi-view stereo are computed independently as separate processing stages, then the reconstruction process is simpler to implement, but errors are exacerbated due to early regularization and residual misalignment

Engineering Contradiction:
Improve3D reconstruction accuracyVSAvoidprocessing system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges photometric stereo and multi-view stereo computations into a unified coupled processing framework. Instead of computing them as independent stages, the system simultaneously solves for surface normals, depth maps, and motion fields by combining constraints from both methods, thereby reducing error propagation while maintaining implementation feasibility through integrated optimization

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified processing system performs multiple functions simultaneously: it computes photometric stereo for surface normals, multi-view stereo for depth information, and optical flow for motion estimation all within a single coupled framework. This multi-functional approach eliminates the need for separate processing stages while improving overall reconstruction accuracy

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If images are processed frame-by-frame independently, then the processing is computationally simpler, but integration drift and misalignment errors accumulate over time

Engineering Contradiction:
Improvetemporal consistencyVSAvoidprocessing algorithm complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements continuous temporal integration by processing images in overlapping three-frame intervals rather than independent frame-by-frame processing. This continuous approach maintains temporal consistency by carrying forward surface and motion estimates across frames, preventing integration drift while managing complexity through efficient interval-based processing

Inventive Principle:
Principle #20Continuity of useful action

Solution Approach 2:

The system applies different processing strategies to different temporal scales: local three-frame intervals are processed with detailed optimization for short-term accuracy, while longer-term temporal consistency is maintained through cumulative integration. This local-global quality differentiation resolves the contradiction between temporal reliability and processing complexity

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If smoothness constraints are applied early in the reconstruction process, then the computation is more stable, but fine geometric details are lost

Engineering Contradiction:
Improvegeometric detail preservationVSAvoidoptimization process complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies preliminary rough reconstruction without strong smoothness constraints to establish an initial geometric framework, then progressively refines the solution with increasing detail preservation. This preliminary action approach maintains computational stability while avoiding early over-smoothing that would lose fine geometric details

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The optimization process dynamically adjusts the balance between smoothness constraints and detail preservation across different processing stages. Early iterations use stronger regularization for stability, while later iterations reduce constraints to preserve fine geometric details, creating a dynamic adaptation that resolves the contradiction between stability and detail

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS9883167B2Photometric three-dimensional facial capture and relighting
Publication Date: 2018.01.30 DISNEY ENTERPRISES INC
  • US9883167B2 patent drawing
  • US9883167B2 patent drawing
  • US9883167B2 patent drawing

AI summary

Embodiments disclosed herein relate to a method and apparatus for generating a three-dimensional surface. In one embodiment, there is a method for generating a three-dimensional surface. The method includes capturing a plurality of images of a target object with at least two cameras, the target object illuminated by at least two sets of red-green-blue (RGB) lights positioned in an array about the target object and generating a three-dimensional surface of the target object by iteratively reconstructing a surface estimate of the target object and aligning images of the target object using motion estimation until the images converge, wherein the images are processed in n-frame intervals.