Procedural Models for Synthetic Training Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning systems face challenges with biased, insufficient, or unavailable training data, leading to inefficient and labor-intensive annotation processes, which affect their performance and accuracy, especially in rare scenarios.

Innovation Solution

The method involves creating procedural models for objects and backgrounds, generating a task environment model, and producing annotated synthetic training data by allocating parameters as annotations, allowing for the training of machine learning modules with unbiased and abundant data, enabling efficient processing of operational data sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If annotated training data is collected from the real world, then the training data reflects actual operational scenarios, but the data suffers from biases, insufficiency, and unavailability especially in rare scenarios

Engineering Contradiction:
Improvetraining data qualityVSAvoidtraining data availability
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent creates procedural models that generate synthetic training data by copying and simulating real-world operational scenarios. Instead of relying on scarce real annotated data, the system procedurally generates unlimited synthetic data samples that replicate the statistical properties and scenarios of real operations, solving both the quantity and quality issues simultaneously

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent performs preliminary annotation during the synthetic data generation process itself. By allocating parameter values as annotations at the time of synthetic data creation, the system eliminates the need for subsequent manual annotation of real data, achieving both abundant data supply and complete annotation without human labor bottlenecks

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If manual annotation of training data is performed, then accurate annotations are obtained, but the process takes a long time and requires lots of resources

Engineering Contradiction:
Improveannotation accuracyVSAvoidannotation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent implements self-service annotation where the synthetic data generation process automatically assigns parameter values as annotations without human intervention. The procedural models inherently know the ground truth values of the parameters they generate, enabling automatic, accurate, and instantaneous annotation that eliminates manual labor entirely

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual human annotation with an automated computational system. The procedural generation model substitutes human annotators by algorithmically assigning parameter values as annotations, achieving the same annotation function without the time and resource constraints of manual processes

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If training data is collected for a specific purpose, then the data collection process is efficient, but the training data may not be optimal for training the machine learning system for a new task

Engineering Contradiction:
Improvedata collection efficiencyVSAvoidtraining data flexibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent creates dynamic procedural models that can adaptively generate training data for different tasks by modifying parameter configurations. The same procedural generation framework can be reconfigured to produce synthetic data for various operational scenarios and machine learning tasks, providing both efficiency and versatility simultaneously

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent develops a universal procedural generation system that serves multiple functions: generating synthetic operational data, creating annotated training datasets, and adapting to different machine learning tasks. This multi-functional system replaces multiple specialized data collection processes with a single adaptable generation framework

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Extent of automation

If machine learning systems use methods such as Convolutional Neural Networks, then the systems can process complex data, but the systems have minimal control over the selection of attributes and features

Engineering Contradiction:
Improveprocessing capabilityVSAvoidfeature control
Core Design Contradiction:
Extent of automationVSEase of operation

Solution Approach 1:

The patent performs preliminary feature selection and control during the synthetic data generation phase. By defining which parameters to allocate as annotations and how to configure the procedural models before generation, the system establishes controlled feature sets that guide the machine learning training process, enabling feature control before the automated processing begins

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20210319363A1Method and system for generating annotated training data
Publication Date: 2021.10.14 YIELD SYST OY
  • US20210319363A1 patent drawing
  • US20210319363A1 patent drawing
  • US20210319363A1 patent drawing

AI summary

A method of generating an annotated synthetic training data for training a machine learning module for processing an operational data set includes creating a first procedural model for the object, the first procedural model having a first set of parameters relating to the object; creating a second procedural model for the background, the second procedural model having a second set of parameters relating to the background; creating the task environment model pertaining to the machine learning task using the first and the second procedural models; creating a synthetic data set using the task environment model; and allocating at least one parameter of the first set of parameters as an annotation for the simulation data to generate the annotated synthetic training data.