Spatial Transformer Network for Face Recognition Data Efficiency

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing face recognition technologies require a large number of image data to recognize faces efficiently, which hampers speed and data efficiency, especially when dealing with target objects different from those in the training data.

Innovation Solution

A method and apparatus that transform images into a standard structure based on labeled feature points, specifically using a spatial transformer network (STN) to determine a standard structure formed by eyes and a nose, and then use AI models to extract and verify faces, reducing the need for extensive training data by applying few-shot adaptation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a large number of image data are used for face recognition training, then recognition accuracy is improved, but processing speed and data efficiency deteriorate

Engineering Contradiction:
Improveface recognition accuracyVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies preliminary action by pre-transforming input images into a standardized structure using a spatial transformer network before face recognition. The STN model pre-aligns facial features (eyes, nose, mouth) to canonical positions, creating a normalized representation that enables faster and more accurate recognition without requiring extensive training data. This preprocessing step prepares the data in advance, allowing the recognition system to operate efficiently on standardized inputs.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If a large number of image data are used for face recognition training, then recognition accuracy is improved, but data efficiency deteriorates

Engineering Contradiction:
Improveface recognition accuracyVSAvoidtraining data volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies parameter changes by transforming the input image parameters (position, scale, orientation) into a standardized canonical form using the spatial transformer network. By changing the parameter space representation of facial features to a normalized coordinate system, the system achieves high recognition accuracy with fewer training samples. The STN learns to adjust transformation parameters (translation, rotation, scaling) to align faces consistently, enabling efficient learning from limited data.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If traditional face recognition methods are used, then recognition of trained target objects is achieved, but recognition of different target objects deteriorates

Engineering Contradiction:
Improverecognition reliabilityVSAvoidcross-object recognition capability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by designing a spatial transformer network that learns generalizable facial feature transformations applicable across different target objects. The STN model identifies and transforms canonical facial landmarks (eyes, nose, mouth) that are universal across various objects, enabling the system to recognize faces of different target objects reliably. The learned transformation parameters and feature extraction methodology transfer across object types, providing both reliability for trained objects and adaptability for new objects.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20240330688A1Method for processing image, and apparatus therefor
Publication Date: 2024.10.03 SAMSUNG ELECTRONICS CO LTD
  • US20240330688A1 patent drawing
  • US20240330688A1 patent drawing
  • US20240330688A1 patent drawing

AI summary

Provided are an artificial intelligence (AI) system using a machine learning algorithm such as deep learning, and applications thereof. A method, performed by an electronic apparatus, of processing images includes obtaining a plurality of training image sets corresponding to a plurality of types of target objects, wherein training images in the training image sets are labeled with feature points forming a preset structure, generating a first artificial intelligence (AI) model for determining a standard structure based on the labeled feature points, by using the training images in the training image sets, identifying a face in a training image transformed based on the standard structure, and training a second AI model for verifying the first AI model, based on an image regressed from the transformed training image, and the training image before being transformed.