GAN-Based Auxiliary Training Data Generation for ML Boundary Accuracy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for generating training data for machine learning systems are inefficient and inadequate in ensuring sufficient representation along the true input feature space class boundaries, leading to errors in object classification and increased vulnerability to adversarial attacks.
Innovation Solution
The method involves generating auxiliary training data by augmenting 3D models and their 2D projections, using generative adversarial networks (GANs) to create additional input data and labels that are closer to the true input feature space class boundaries, thereby improving the quality and efficiency of training data collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional training data generation methods are used, then the training process is simple, but the representation along true input feature space class boundaries is insufficient
Solution Approach 1:
The patent introduces an intermediary system (GAN-based data generation apparatus) that mediates between the training process and the training data. This intermediary generates synthetic training data that specifically targets regions near true class boundaries in the input feature space, thereby improving representation accuracy without requiring manual intervention for each data point.
Solution Approach 2:
The patent changes the parameters of the training data generation process by using GANs to transform the distribution of generated data. The generator and discriminator networks adjust data parameters iteratively to produce samples that closely match the true data distribution, particularly in critical boundary regions, thereby improving measurement precision.
2Measurement precision
If more training data is collected to improve boundary representation, then the accuracy improves, but the data collection efficiency decreases
Solution Approach 1:
The patent uses copying by generating synthetic replicas of training data through GANs. Instead of collecting additional real data, the system copies and transforms existing data patterns to create synthetic samples that specifically target boundary regions, thereby improving accuracy without the time-consuming process of manual data collection.
Solution Approach 2:
The patent performs preliminary action by pre-processing and augmenting training data before the main training process. The GAN-based system prepares synthetic data in advance, focusing on critical boundary regions, so that the subsequent training process receives high-quality, pre-optimized data without delays during training.
3Manufacturing precision
If conventional training data is used, then the training process is fast, but the learned class boundaries have high error compared to true boundaries
Solution Approach 1:
The patent applies local quality by focusing computational resources on generating synthetic data specifically in regions near true class boundaries, rather than uniformly distributing data throughout the entire feature space. This localized approach improves boundary accuracy by concentrating training examples where they are most needed, without requiring a complete overhaul of the entire training dataset.
4Reliability
If training data lacks sufficient boundary representation, then the training process is efficient, but the system becomes vulnerable to adversarial attacks
Solution Approach 1:
The patent applies preliminary anti-action by proactively generating synthetic training data that anticipates potential adversarial attack vectors. By pre-training the model with diverse synthetic examples that stress-test boundary regions, the system builds inherent resistance to adversarial attacks before deployment, rather than adding complex defense mechanisms later.
Data Source
AI summary
The present disclosure provides an apparatus and method for training a machine learning engine configured to determine whether an object in a two dimensional (2D) image is in-scope or out-of-scope relative to the one or more 3D objects that includes receiving a 3D model of each of the one or more 3D objects, for each 3D model receiving a set of specifications and thresholds for the 3D model, augmenting the specifications of the 3D model to generate a plurality of augmented 3D models, and generating auxiliary training data based on the plurality of augmented 3D models, and utilizing the auxiliary training data to train the machine learning engine. The present disclosure also provides an apparatus and method for training a machine learning engine that includes receiving sampling parameters related to a system of interest, sampling a generative model of a generative adversarial network (GAN) based on the sampling parameters to generate auxiliary input data, inputting the auxiliary input data into a discriminator model of the GAN to generate an auxiliary label associated with each of the auxiliary input data element, and utilizing the auxiliary input data and auxiliary labels as auxiliary training data to train the machine learning engine.


