Synthetic Data Generation System with Controlled Variability

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional methods for generating training datasets for machine learning models are expensive, time-consuming, and often lack variability, especially in sensitive areas like healthcare due to privacy concerns, and struggle to include rare or outlier scenarios.

Innovation Solution

A system and method for generating synthetic datasets with user-controlled variability, including a graphical user interface that allows users to select object types, variability parameters, and camera characteristics, enabling the creation of customized datasets that mimic real-world scenarios.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If traditional manual data collection and labeling methods are used, then data quality may be maintained, but the process becomes expensive and time-consuming

Engineering Contradiction:
Improvedata generation speedVSAvoidtime for data collection and labeling
Core Design Contradiction:
ProductivityVSLoss of time

Solution Approach 1:

The patent uses synthetic data generation to create copies of real-world data scenarios through simulation. Instead of collecting and manually labeling actual images, the system generates synthetic images that replicate real-world conditions, object appearances, and scenarios, eliminating the need for time-consuming manual data collection and labeling while maintaining data quality for training machine learning models

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of manual data collection and labeling with an automated computational system. The system uses algorithms to automatically generate synthetic training data, automatically annotate objects within the synthetic images, and simulate various real-world conditions, substituting human labor with automated computational processes that are both faster and more efficient

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Adaptability or versatility

If real-world data is collected for training, then data authenticity is maintained, but privacy concerns and regulations limit data availability

Engineering Contradiction:
Improvedata availability for trainingVSAvoidprivacy concerns
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of real-world data that preserve the statistical properties and visual characteristics of authentic data without containing any actual private or sensitive information. By generating artificial images that simulate real-world scenarios, the system maintains data authenticity for training purposes while eliminating privacy risks associated with using real patient or personal data

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic data as an intermediary between the need for authentic training data and privacy protection requirements. Instead of directly using real-world data that raises privacy concerns, the system uses synthesized representations that mediate between these conflicting requirements, providing the benefits of authentic data while eliminating privacy harms

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If diverse training data is generated, then model robustness improves, but data generation complexity increases

Engineering Contradiction:
Improvemodel training robustnessVSAvoiddata generation system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic control over data generation parameters, allowing the system to adaptively adjust the complexity and diversity of synthetic data based on training needs. The system can dynamically modify object types, environmental conditions, lighting scenarios, and camera parameters to generate appropriately diverse training datasets without requiring a permanently complex generation system

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent segments the data generation process into modular components that can be independently configured and controlled. By dividing the complex generation task into separate controllable parameters (object selection, environment settings, camera characteristics, annotation types), the system achieves high data diversity through simple parameter combinations rather than complex integrated systems

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11755189B2Systems and methods for synthetic data generation
Publication Date: 2023.09.12 DATAGEN TECH LTD
  • US11755189B2 patent drawing
  • US11755189B2 patent drawing
  • US11755189B2 patent drawing

AI summary

A non-transitory, computer-readable medium includes instructions that causes at least one processing device to display a graphical user interface (GUI) configured to facilitate generating a synthetic dataset including a plurality of images. The GUI includes a dataset size selector to receive user input to indicate a number of images to generate and include in the synthetic dataset; a target object type selector to receive user input indicative of at least one selected target object type to feature in the synthetic dataset; one or more image parameter variability controls to receive user input indicative of at least one variation to include in the synthetic dataset relative to target object representations generated based on the at least one selected target object type; and a dataset generation control to initiate generating the synthetic dataset. The synthetic dataset is generated according to the size input, target object type input, and variability input.