FACS-Guided Facial Image Generation for Identity-Preserving Expressions
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional image generation models struggle to accurately control facial expressions while maintaining the identity of a person in the output image, often entangling facial expression with identity information, leading to inconsistent results.
Innovation Solution
The use of a Facial Action Coding System (FACS) representation to disentangle identity information from facial expressions, allowing an image generation model to generate synthetic images with targeted expressions while preserving the identity of a different person.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional image generation models are used to generate synthetic images with facial expressions, then the model can produce varied facial expressions, but the identity information becomes entangled with expression information leading to inconsistent results
Solution Approach 1:
The patent segments the facial image into two distinct components: identity information and expression information. By using FACS (Facial Action Coding System) representation, the expression is encoded separately from the identity features. This segmentation allows independent control of expression while preserving identity, resolving the entanglement problem in conventional models.
Solution Approach 2:
The patent introduces FACS representation as an intermediary between the input image and the generated output. The FACS encoding acts as a mediator that captures expression information in a structured format independent of identity, allowing the generation model to control expressions without compromising identity preservation.
2Reliability
If the image generation model tries to maintain both identity and expression control, then more features are preserved, but the complexity of the model increases
Solution Approach 1:
The patent changes the parameter representation of facial expressions from raw pixel variations to FACS action unit codes. This parameter transformation simplifies the expression control space into discrete, interpretable units that are easier for the model to process, reducing complexity while improving expression consistency.
3Device complexity
If the model entangles identity with expression information for simpler processing, then the model structure is simpler, but the expression generation becomes inaccurate
Solution Approach 1:
The patent extracts expression information from the input image and represents it separately using FACS encoding. This extraction removes the expression component from the identity information, allowing the model to process them independently. The result is improved expression accuracy without requiring complex entangled representations.
Data Source
AI summary
A method, apparatus, non-transitory computer readable medium, and system for image processing include obtaining an expression input indicating a facial expression, generating a guidance feature based on the expression input, where the guidance feature comprises a facial action coding system (FACS) representation of the facial expression, and generating a synthetic image based on the guidance feature, where the synthetic image depicts the facial expression


