Synthetic Training Data Generation for Biometric Identity Verification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional user identification systems face challenges such as susceptibility to fraud, slow speed, inaccuracy, and operational limitations, particularly when using machine learning systems for biometric identification, as they often require large sets of unique training data that are costly and impractical to acquire, especially when dealing with variations like replicas, dirty hands, or anatomical features.
Innovation Solution
The techniques involve processing input data to generate enhanced training data using synthetic or augmented image processing methods, such as Laplacian pyramid decomposition and parameter modulation, to create images that depict artifacts or varied hand conditions, allowing for the rapid generation of a large set of training data that includes replicas, dirty hands, and other real-world scenarios.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If large sets of unique training data are acquired to improve machine learning system performance, then accuracy is improved, but cost and practicality worsen
Solution Approach 1:
The patent uses synthetic image generation to create copies of training data that simulate various real-world conditions (replicas, dirty hands, different lighting). Instead of acquiring millions of unique real images, the system generates synthetic copies from a smaller set of source images, applying transformations to create diverse training samples that improve accuracy without proportional increases in acquisition cost
Solution Approach 2:
The system applies parameter changes to source images by modifying various attributes such as adding artificial dirt, changing lighting conditions, adjusting colors, and applying filters. These parameter modifications create varied training samples from limited source material, enabling the system to generate large volumes of diverse training data without proportional increases in acquisition cost
2Reliability
If diverse training scenarios are included to improve robustness, then reliability is improved, but data acquisition complexity worsens
Solution Approach 1:
The system generates synthetic copies of training images that represent diverse real-world scenarios including replicas, dirty hands, and various lighting conditions. By copying and transforming a limited set of source images, the system creates comprehensive training coverage without the complexity of acquiring each scenario separately
Solution Approach 2:
The synthetic image generation system serves multiple functions: it creates replicas, simulates dirty conditions, adjusts lighting, and generates various hand positions from a single source image. This multi-functional approach allows one data acquisition process to produce diverse training scenarios that would otherwise require multiple separate acquisition efforts
3Productivity
If synthetic image processing methods are used to generate enhanced training data, then productivity is improved, but manufacturing precision worsens
Solution Approach 1:
The system carefully controls parameter changes during synthetic image generation, applying transformations such as dirt addition, lighting adjustments, and color modifications with precision. By systematically varying parameters within realistic ranges, the system maintains image quality and realism while enabling rapid generation of diverse training samples
Solution Approach 2:
The synthetic image generation process creates high-quality copies of source images with realistic variations. The copying process incorporates sophisticated rendering techniques that preserve anatomical accuracy and visual realism, ensuring that generated images maintain sufficient quality for training purposes while enabling high-speed data generation
Data Source
AI summary
Enhanced training data representative of possible inputs is used to train a machine learning system. For example, a machine learning system to determine identity based on an image of a human palm may be trained using enhanced training data comprising images. The enhanced training data may comprise source images that have been modified to appear to depict synthetic artifacts that attempt to simulate human palms, augmented images of dirty hands, and so forth. A synthetic artifact image may be produced by selectively removing some data from a source image. An augmented image may be produced by selectively blending the source image with features extracted from sample images. These images may then be used as training data to train the machine learning system.


