Synthetic Data Generation Model Augmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Machine learning models often struggle with accuracy due to insufficient and non-diverse sample data sets, making it challenging to obtain large and varied datasets for training.

Innovation Solution

The use of synthetic data generation models to augment training data by generating synthetic data points that resemble task-specific data, leveraging general data sources to conserve task-specific data points and enhance model training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If synthetic data generation models are used to augment training data, then the diversity and quantity of training data increase, but the complexity of the system increases

Engineering Contradiction:
Improvequantity of training dataVSAvoidsystem complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent uses synthetic data generation models to create copies of task-specific data points from general data sources. These synthetic copies resemble the original task-specific data while providing additional quantity and diversity, effectively multiplying the available training data without requiring collection of entirely new data points.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic data generation models as intermediary components between general data sources and the machine learning model. These intermediaries transform general data into task-specific synthetic data points, facilitating the augmentation process while managing system complexity through a dedicated transformation layer.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If synthetic data generation models are used to augment training data, then the diversity of training data increases, but the computational resources required increase

Engineering Contradiction:
Improvediversity of training dataVSAvoidcomputational resources
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent employs parameter changes in the synthetic data generation models to control the diversity and characteristics of generated data. By adjusting model parameters, the system can generate diverse synthetic data points that match task-specific requirements while optimizing computational resource consumption based on the desired level of diversity.

Inventive Principle:
Principle #35Parameter changes

3Reliability

If task-specific data points are used for training, then the model accuracy improves, but the data collection requirements increase

Engineering Contradiction:
Improvemodel accuracyVSAvoiddata collection
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent creates synthetic copies of task-specific data points from general data sources. These copies maintain the essential characteristics needed for model accuracy while eliminating the need to collect and annotate large amounts of new task-specific data, thereby preserving accuracy requirements while simplifying data collection.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system uses general data sources to automatically generate task-specific synthetic data points through the synthetic data generation models. This self-service approach eliminates the manual data collection and annotation process for task-specific data, as the system autonomously transforms general data into appropriate training examples.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250181673A1Guided Augmentation Of Data Sets For Machine Learning Models
Publication Date: 2025.06.05 ORACLE INT CORP
  • US20250181673A1 patent drawing
  • US20250181673A1 patent drawing
  • US20250181673A1 patent drawing

AI summary

Techniques are disclosed for augmenting data sets used for training machine learning models and for generating predictions by trained machine learning models. These techniques may increase a number and diversity of examples within an initial training dataset of sentences by extracting a subset of words from the existing training dataset of sentences. The techniques may conserve scarce sample data in few-shot situations by training a data generation model using general data obtained from a general data source.