3D Face Model Fitting via Tensor Decoupling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video capture technologies face challenges in capturing high-quality video in harsh environments and efficiently processing human face images across varying expressions and conditions, particularly in video streams where equipment limitations and environmental factors hinder effective video capture and editing.

Innovation Solution

The development of a face-fitting module that identifies two-dimensional local feature points and generates a three-dimensional face model by combining predefined models, reducing error and constraining facial expression changes smoothly across frames, using a 3-mode tensor model to decouple identity and expression, and employing energy minimization algorithms for robust fitting and editing applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional video capture equipment is used in harsh environments, then video capture capability is limited, but equipment portability and robustness are improved

Engineering Contradiction:
Improvevideo capture capabilityVSAvoidenvironmental adaptability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent replaces traditional mechanical video capture systems with a digital face-fitting system that uses computer vision algorithms and 3D modeling. This substitution allows the system to operate in harsh environments where traditional equipment would fail, while maintaining high video capture quality through software-based image processing and reconstruction techniques.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If complex face model processing is performed, then face fitting accuracy is improved, but computational complexity and processing time increase

Engineering Contradiction:
Improveface fitting accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the complex face fitting process into distinct modules: feature point detection, 3D model selection, parameter optimization, and expression analysis. Each module handles a specific aspect of the problem, reducing overall computational complexity while maintaining high accuracy through specialized processing at each stage.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary actions by pre-defining a library of 3D face models and feature point configurations before actual video processing. This pre-computation reduces real-time computational requirements, allowing complex face fitting to be performed efficiently during video playback without excessive processing delays.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If facial expression changes are allowed to vary freely, then natural expression representation is improved, but temporal coherence and smoothness deteriorate

Engineering Contradiction:
Improveexpression representationVSAvoidtemporal coherence
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The patent implements feedback mechanisms that monitor temporal coherence of facial expressions across video frames. When expression changes become too abrupt or inconsistent, the system adjusts the transformation parameters to maintain smooth transitions, while still preserving the essential characteristics of natural facial expressions through controlled adaptation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8923392B2Methods and apparatus for face fitting and editing applications
Publication Date: 2014.12.30 ADOBE INC
  • US8923392B2 patent drawing
  • US8923392B2 patent drawing
  • US8923392B2 patent drawing

AI summary

Various embodiments of methods and apparatus for face fitting are disclosed. In one embodiment, sets of two-dimensional local feature points on a face in each image of a set of images are identified. The set of images includes a sequence of frames a video stream. A three-dimensional face model for the face in the each image is generated as a combination of a set of predefined three-dimensional face models. In some embodiments, the generating includes reducing an error between a projection of vertices of the set of predefined three-dimensional face models and the two-dimensional local feature points of the each image, and constraining facial expression of the three-dimensional face model to change smoothly from image to image in the sequence of video frames.