Identity-Preserving Face Synthesis From Unlabeled Face Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional face synthesis models require training data that includes face images of the identity to be preserved and labeled attributes, limiting their ability to preserve identity and reflect diverse attributes, especially for faces not in the training data, and they struggle with labeling numerous attributes in large datasets.

Innovation Solution

A learning network architecture comprising sub-networks for identity and attribute extraction, with unsupervised training using unlabeled data to generate identity-preserving face images, allowing synthesis for any identity without requiring labeled attributes beyond identity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional face synthesis models are trained with labeled attributes, then they can preserve identity for faces in the training data, but they cannot synthesize faces for identities not present in the training data and require extensive labeling effort

Engineering Contradiction:
Improveability to synthesize any identityVSAvoidlabeling effort
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system performs self-service by automatically extracting identity features and attributes from unlabeled face images using deep learning models. The attribute extraction network and identity extraction network work autonomously to learn from raw data without requiring manual annotation, enabling the system to handle any identity while minimizing human labeling effort

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The face synthesis model achieves universality by being trained on diverse unlabeled face images from multiple identities. The generated face image network learns to generalize across different identities and attributes, enabling it to synthesize faces for any identity rather than being limited to specific trained identities

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Manufacturing precision

If face synthesis models use labeled training data with multiple attributes, then they can achieve precise attribute control, but the labeling process becomes extremely complex and time-consuming for large datasets

Engineering Contradiction:
Improveattribute synthesis precisionVSAvoidtraining data preparation time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent replaces the manual mechanical labeling process with automated computational systems. Deep learning-based attribute extraction networks automatically identify and extract attributes such as pose, expression, and illumination from face images, substituting human annotators with algorithms that can process large datasets efficiently and consistently

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The system introduces an intermediary attribute extraction network that acts as a bridge between raw face images and the synthesis process. This intermediary automatically derives attribute information from images, providing structured attribute data without requiring direct manual labeling, thus reducing preparation time while maintaining precision

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If conventional models are trained only with labeled data, then they achieve good performance on labeled attributes, but they fail to capture diverse attributes present in unlabeled real-world data

Engineering Contradiction:
Improveperformance on labeled attributesVSAvoidattribute diversity
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent merges labeled and unlabeled data training approaches into a unified framework. The model is trained on labeled data for supervised attribute learning, then further trained or fine-tuned on unlabeled real-world data to capture additional attribute variations and distributions, combining the reliability of supervised learning with the versatility of unsupervised learning

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP3746934B1Face synthesis
Publication Date: 2025.11.05 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3746934B1 patent drawingFigure 1
  • EP3746934B1 patent drawingFigure 2
  • EP3746934B1 patent drawingFigure 3

AI summary

In accordance with implementations of the subject matter described herein, there is provided a solution for face synthesis. In this solution, a first image about a face of a first user and a second image about a face of a second user are obtained. A first feature characterizing an identity of the first user is extracted from the first image, and a second feature characterizing a plurality of attributes of the second image is extracted from the second image, where the plurality of attributes do not include the identity of the second user. Then, a third image about a face of the first user is generated based on the first and second features, the third image reflecting the identity of the first user and the plurality of attributes of the second image.