FACS-Guided Facial Image Generation for Identity-Preserving Expressions

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional image generation models struggle to accurately control facial expressions while maintaining the identity of a person in the output image, often entangling facial expression with identity information, leading to inconsistent results.

Innovation Solution

The use of a Facial Action Coding System (FACS) representation to disentangle identity information from facial expressions, allowing an image generation model to generate synthetic images with targeted expressions while preserving the identity of a different person.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional image generation models are used to generate synthetic images with facial expressions, then the model can produce varied facial expressions, but the identity information becomes entangled with expression information leading to inconsistent results

Engineering Contradiction:
Improvefacial expression controlVSAvoididentity preservation accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the facial image into two distinct components: identity information and expression information. By using FACS (Facial Action Coding System) representation, the expression is encoded separately from the identity features. This segmentation allows independent control of expression while preserving identity, resolving the entanglement problem in conventional models.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces FACS representation as an intermediary between the input image and the generated output. The FACS encoding acts as a mediator that captures expression information in a structured format independent of identity, allowing the generation model to control expressions without compromising identity preservation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If the image generation model tries to maintain both identity and expression control, then more features are preserved, but the complexity of the model increases

Engineering Contradiction:
Improveexpression consistencyVSAvoidmodel structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent changes the parameter representation of facial expressions from raw pixel variations to FACS action unit codes. This parameter transformation simplifies the expression control space into discrete, interpretable units that are easier for the model to process, reducing complexity while improving expression consistency.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If the model entangles identity with expression information for simpler processing, then the model structure is simpler, but the expression generation becomes inaccurate

Engineering Contradiction:
Improveprocessing simplicityVSAvoidexpression accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent extracts expression information from the input image and represents it separately using FACS encoding. This extraction removes the expression component from the identity information, allowing the model to process them independently. The result is improved expression accuracy without requiring complex entangled representations.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS20260080598A1Animatable facial image generation with facial action coding system
Publication Date: 2026.03.19 ADOBE INC
  • US20260080598A1 patent drawing
  • US20260080598A1 patent drawing
  • US20260080598A1 patent drawing

AI summary

A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining an expression input indicating a facial expression, generating a guidance feature based on the expression input, where the guidance feature comprises a facial action coding system (FACS) representation of the facial expression, and generating a synthetic image based on the guidance feature, where the synthetic image depicts the facial expression