Procedural Models for Synthetic Training Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning systems face challenges with biased, insufficient, or unavailable training data, leading to inefficient and labor-intensive annotation processes, which affect their performance and accuracy, especially in rare scenarios.
Innovation Solution
The method involves creating procedural models for objects and backgrounds, generating a task environment model, and producing annotated synthetic training data by allocating parameters as annotations, allowing for the training of machine learning modules with unbiased and abundant data, enabling efficient processing of operational data sets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If annotated training data is collected from the real world, then the training data reflects actual operational scenarios, but the data suffers from biases, insufficiency, and unavailability especially in rare scenarios
Solution Approach 1:
The patent creates procedural models that generate synthetic training data by copying and simulating real-world operational scenarios. Instead of relying on scarce real annotated data, the system procedurally generates unlimited synthetic data samples that replicate the statistical properties and scenarios of real operations, solving both the quantity and quality issues simultaneously
Solution Approach 2:
The patent performs preliminary annotation during the synthetic data generation process itself. By allocating parameter values as annotations at the time of synthetic data creation, the system eliminates the need for subsequent manual annotation of real data, achieving both abundant data supply and complete annotation without human labor bottlenecks
2Measurement precision
If manual annotation of training data is performed, then accurate annotations are obtained, but the process takes a long time and requires lots of resources
Solution Approach 1:
The patent implements self-service annotation where the synthetic data generation process automatically assigns parameter values as annotations without human intervention. The procedural models inherently know the ground truth values of the parameters they generate, enabling automatic, accurate, and instantaneous annotation that eliminates manual labor entirely
Solution Approach 2:
The patent replaces the mechanical process of manual human annotation with an automated computational system. The procedural generation model substitutes human annotators by algorithmically assigning parameter values as annotations, achieving the same annotation function without the time and resource constraints of manual processes
3Productivity
If training data is collected for a specific purpose, then the data collection process is efficient, but the training data may not be optimal for training the machine learning system for a new task
Solution Approach 1:
The patent creates dynamic procedural models that can adaptively generate training data for different tasks by modifying parameter configurations. The same procedural generation framework can be reconfigured to produce synthetic data for various operational scenarios and machine learning tasks, providing both efficiency and versatility simultaneously
Solution Approach 2:
The patent develops a universal procedural generation system that serves multiple functions: generating synthetic operational data, creating annotated training datasets, and adapting to different machine learning tasks. This multi-functional system replaces multiple specialized data collection processes with a single adaptable generation framework
4Extent of automation
If machine learning systems use methods such as Convolutional Neural Networks, then the systems can process complex data, but the systems have minimal control over the selection of attributes and features
Solution Approach 1:
The patent performs preliminary feature selection and control during the synthetic data generation phase. By defining which parameters to allocate as annotations and how to configure the procedural models before generation, the system establishes controlled feature sets that guide the machine learning training process, enabling feature control before the automated processing begins
Data Source
AI summary
A method of generating an annotated synthetic training data for training a machine learning module for processing an operational data set includes creating a first procedural model for the object, the first procedural model having a first set of parameters relating to the object; creating a second procedural model for the background, the second procedural model having a second set of parameters relating to the background; creating the task environment model pertaining to the machine learning task using the first and the second procedural models; creating a synthetic data set using the task environment model; and allocating at least one parameter of the first set of parameters as an annotation for the simulation data to generate the annotated synthetic training data.


