Synthetic Training Data Generation for Biometric Identity Verification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional user identification systems face challenges such as susceptibility to fraud, slow speed, inaccuracy, and operational limitations, particularly when using machine learning systems for biometric identification, as they often require large sets of unique training data that are costly and impractical to acquire, especially when dealing with variations like replicas, dirty hands, or anatomical features.

Innovation Solution

The techniques involve processing input data to generate enhanced training data using synthetic or augmented image processing methods, such as Laplacian pyramid decomposition and parameter modulation, to create images that depict artifacts or varied hand conditions, allowing for the rapid generation of a large set of training data that includes replicas, dirty hands, and other real-world scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If large sets of unique training data are acquired to improve machine learning system performance, then accuracy is improved, but cost and practicality worsen

Engineering Contradiction:
ImproveaccuracyVSAvoidcost
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent uses synthetic image generation to create copies of training data that simulate various real-world conditions (replicas, dirty hands, different lighting). Instead of acquiring millions of unique real images, the system generates synthetic copies from a smaller set of source images, applying transformations to create diverse training samples that improve accuracy without proportional increases in acquisition cost

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system applies parameter changes to source images by modifying various attributes such as adding artificial dirt, changing lighting conditions, adjusting colors, and applying filters. These parameter modifications create varied training samples from limited source material, enabling the system to generate large volumes of diverse training data without proportional increases in acquisition cost

Inventive Principle:
Principle #35Parameter changes

2Reliability

If diverse training scenarios are included to improve robustness, then reliability is improved, but data acquisition complexity worsens

Engineering Contradiction:
ImproverobustnessVSAvoiddata acquisition complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system generates synthetic copies of training images that represent diverse real-world scenarios including replicas, dirty hands, and various lighting conditions. By copying and transforming a limited set of source images, the system creates comprehensive training coverage without the complexity of acquiring each scenario separately

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The synthetic image generation system serves multiple functions: it creates replicas, simulates dirty conditions, adjusts lighting, and generates various hand positions from a single source image. This multi-functional approach allows one data acquisition process to produce diverse training scenarios that would otherwise require multiple separate acquisition efforts

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Productivity

If synthetic image processing methods are used to generate enhanced training data, then productivity is improved, but manufacturing precision worsens

Engineering Contradiction:
Improvedata generation speedVSAvoidimage quality
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The system carefully controls parameter changes during synthetic image generation, applying transformations such as dirt addition, lighting adjustments, and color modifications with precision. By systematically varying parameters within realistic ranges, the system maintains image quality and realism while enabling rapid generation of diverse training samples

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The synthetic image generation process creates high-quality copies of source images with realistic variations. The copying process incorporates sophisticated rendering techniques that preserve anatomical accuracy and visual realism, ensuring that generated images maintain sufficient quality for training purposes while enabling high-speed data generation

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12190566B1System for generating enhanced training data
Publication Date: 2025.01.07 AMAZON TECH INC
  • US12190566B1 patent drawing
  • US12190566B1 patent drawing
  • US12190566B1 patent drawing

AI summary

Enhanced training data representative of possible inputs is used to train a machine learning system. For example, a machine learning system to determine identity based on an image of a human palm may be trained using enhanced training data comprising images. The enhanced training data may comprise source images that have been modified to appear to depict synthetic artifacts that attempt to simulate human palms, augmented images of dirty hands, and so forth. A synthetic artifact image may be produced by selectively removing some data from a source image. An augmented image may be produced by selectively blending the source image with features extracted from sample images. These images may then be used as training data to train the machine learning system.