Facial Landmark Localization Using Pose-Specific Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing facial landmark localization methods are inadequate for handling variations in pose, illumination, expression, occlusion, and image resolution, particularly failing to accurately localize a dense set of landmarks in non-frontal faces with large yaw variation and partial occlusions, and often require extensive manual annotation and are sensitive to facial detection errors.

Innovation Solution

A unified framework for dense facial landmark localization that addresses pose variation (including yaw and roll), expressions, and partial occlusion, using a novel l1-regularized least squares approach and local texture classifiers to constrain shape coefficients and refine facial shapes, allowing for the automatic localization of a dense set of landmarks even in low-resolution images.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional deformable template models (ASMs, AAMs) are used for facial landmark localization, then the method can handle basic facial shape variations, but it fails to accurately localize landmarks in faces with large pose variation, occlusion, and expression changes

Engineering Contradiction:
Improveadaptability to pose variation and occlusionVSAvoidlandmark localization accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent segments the facial landmark localization problem into multiple pose-specific models (frontal, left profile, right profile, and intermediate poses). Each model is trained on images with specific pose ranges, allowing the system to handle large pose variations by selecting the appropriate model. This segmentation resolves the contradiction by making the system adaptable to different poses while maintaining high localization accuracy within each pose category.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by using pose-specific training data and models tailored to each facial pose category. Instead of using a single global model, the system employs localized models that are optimized for specific pose ranges, thereby achieving high accuracy for each local pose condition while maintaining overall versatility across all poses.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If dense landmark localization is performed to capture detailed facial features, then the system can provide comprehensive facial analysis, but it becomes more sensitive to detection errors and computational complexity increases

Engineering Contradiction:
Improvenumber of landmarks localizedVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent segments the dense landmark localization task into pose-specific subtasks, where each pose-specific model handles localization for its designated pose range. This segmentation reduces the complexity of each individual model while collectively achieving comprehensive dense landmark localization across all poses, thereby resolving the contradiction between quantity of landmarks and system complexity.

Inventive Principle:
Principle #1Segmentation

3Measurement precision

If manual annotation is used to provide ground truth for training, then the training data quality is high, but the annotation process is arduous and time-consuming

Engineering Contradiction:
Improvetraining data accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary action by carefully curating and annotating training images for each pose category before model training. The training sets are pre-prepared with pose-specific manual annotations, which are then used to train the respective pose-specific models. This preliminary annotation work, though time-consuming, is done once and enables the models to achieve high accuracy without requiring additional annotation during operation.

Inventive Principle:
Principle #10Preliminary action

4Loss of time

If the system is trained on limited training data to reduce annotation effort, then the annotation time is reduced, but the model's ability to generalize to unseen variations deteriorates

Engineering Contradiction:
Improvetraining preparation timeVSAvoidgeneralization ability
Core Design Contradiction:
Loss of timeVSAdaptability or versatility

Solution Approach 1:

The patent segments the training data into multiple pose-specific subsets, where each subset contains images with specific pose characteristics. By training separate models on these segmented subsets, the system achieves better generalization within each pose category compared to training a single model on all data, while the total annotation effort is distributed across manageable segments.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10121055B1Method and system for facial landmark localization
Publication Date: 2018.11.06 CARNEGIE MELLON UNIV
  • US10121055B1 patent drawing
  • US10121055B1 patent drawing
  • US10121055B1 patent drawing

AI summary

This invention describes methods and systems for the automated facial landmark localization. Our approach proceeds from sparse to dense landmarking steps using a set of models to best account for the shape and texture variation manifested by facial landmarks across pose and expression. We also describe the use of an l1-regularized least squares approach that we incorporate into our shape model, which is an improvement over the shape model used by several prior Active Shape Model (ASM) based facial landmark localization algorithms.