Facial Marker Tracking via Dual-View Triangulation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current motion-capture systems face challenges in accurately capturing facial performance due to the small movement of facial markers compared to body markers, leading to unrealistic computer-animated characters and discomfort for actors wearing heavy rigging systems.

Innovation Solution

The system captures images from two different perspectives simultaneously using HD video cameras, employing image processing techniques and graph matching algorithms to triangulate marker positions in 3D space, allowing for more accurate facial performance capture without the need for heavy rigging.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional motion-capture systems use multiple cameras to track facial markers, then measurement precision of facial movement is improved, but device complexity and weight increase due to heavy rigging requirements

Engineering Contradiction:
Improvefacial marker position accuracyVSAvoidrigging system complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system divides the capture task into two distinct parts: a stationary camera system for capturing high-resolution reference images of the face, and a lightweight head-mounted camera for capturing motion. This segmentation allows each component to be optimized independently, eliminating the need for complex heavy rigging while maintaining measurement precision.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a photograph (reference image) as an intermediary element that mediates between the stationary camera system and the moving head-mounted camera. By matching features in the reference photograph with features captured by the moving camera, the system achieves accurate facial marker tracking without requiring the moving camera to directly observe all markers at all times.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If statistical models are used for facial animation, then productivity of animation production is improved, but loss of information occurs regarding nuanced facial performance

Engineering Contradiction:
Improveanimation production efficiencyVSAvoidfacial performance nuances
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The system uses the reference photograph as feedback to continuously correct and refine the positioning of facial markers during animation. By constantly comparing the head-mounted camera's view against the reference image and adjusting marker positions accordingly, the system preserves nuanced facial performance information while maintaining efficient automated animation production.

Inventive Principle:
Principle #23Feedback

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This method improves the accuracy of facial performance capture, reducing the 'uncanny valley' effect in animations and enhancing the realism of computer-generated characters while reducing the burden on actors with lighter and more comfortable equipment.

Implementation Method 1

Based upon the matched points, the markers on the surface of the object can be triangulated to a 3D position

Methodology Applied
Scientific EffectTriangulation:

Data Source

PatentEP2342676B1Methods and apparatus for dot marker matching
Publication Date: 2014.06.18 TWO PIC MC LLC
  • EP2342676B1 patent drawingFigure 1A~1B
  • EP2342676B1 patent drawingFigure 2A~2B
  • EP2342676B1 patent drawingFigure 3A

AI summary

A method for a computer system includes receiving a first camera image of a 3D object having sensor markers, captured from a first location, at a first instance, receiving a second camera image of the 3D object from a second location, at a different instance, determining points from the first camera image representing sensor markers of the 3D object, determining points from the second camera image representing sensor markers of the 3D object, determining approximate correspondence between points from the first camera image and points from the second camera image, determining approximate 3D locations some sensor markers of the 3D object, and rendering an image including the 3D object in response to the approximate 3D locations.