Autoencoder Face Normalization for Cross-Group Expression Classification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing facial expression recognition technologies struggle to generalize across different groups of people due to variations in facial appearances, demographics, and data collection settings, leading to performance gaps and biases.

Innovation Solution

A self-supervised denoising autoencoder is used to transfer facial expressions onto a common facial template, separating the learning process into phases that reduce individual differences while preserving dynamic characteristics, allowing for improved model generalization and classification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional facial expression recognition models are trained on diverse datasets, then they can handle more variations in facial appearances, but they still fail to generalize across different groups due to individual differences and biases

Engineering Contradiction:
Improvegeneralization across different groupsVSAvoidperformance consistency
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces a template face as an intermediary representation that mediates between individual facial appearances and expression recognition. The normalization process uses this template face to transform individual faces into a common representation space, enabling the model to generalize across different groups while maintaining reliable performance. This intermediary template face serves as a universal reference that bridges the gap between diverse individual characteristics and consistent recognition outcomes.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If face normalization is performed using traditional methods, then processing speed is maintained, but performance gaps and biases related to gender, skin type, and demographics persist

Engineering Contradiction:
Improveclassification accuracyVSAvoidnormalization process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary face normalization using a template face before the main expression recognition process. By pre-processing the facial images to align them with a common template representation, the system eliminates individual appearance variations in advance. This preliminary action ensures that subsequent classification operates on normalized data, improving accuracy while managing complexity through a structured two-stage process.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If models are trained to recognize individual facial features, then they achieve high accuracy for specific individuals, but they fail to generalize to other groups

Engineering Contradiction:
Improvefacial expression recognition accuracyVSAvoidcross-group generalization
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent extracts and separates the expression-related features from identity-specific features through template-based normalization. By removing individual facial characteristics and retaining only the expression-related variations in the normalized representation, the system achieves high recognition accuracy for expressions while enabling generalization across different individuals and groups. This extraction process isolates the relevant information needed for expression recognition.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP4238073B1Classification of a human characteristic data sample after normalization with an autoencoder
Publication Date: 2026.03.18 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP4238073B1 patent drawingFigure 1
  • EP4238073B1 patent drawingFigure 2
  • EP4238073B1 patent drawingFigure 3

AI summary

Generally discussed herein are devices, systems, and methods for. A method can include obtaining a normalizing autoencoder, the normalizing autoencoder trained based on first data samples of a template person and second data samples of a variety of people, normalizing, by the normalizing autoencoder, an input data sample by combining dynamic characteristics of a person in the input data sample with static characteristics in the first data samples, to generate normalized data, and providing the normalized data as input to a classifier model to classify the input data based on the dynamic characteristics of the input data and the static characteristics of the first data samples.