Synthetic Data Generation Using Environmental Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing synthetic data generation methods struggle to accurately reflect the features of actual data, leading to suboptimal performance in artificial intelligence models due to the use of virtual environments and statistical models, which can result in cost, time, and security challenges, particularly in data-scarce industrial fields like disaster scenarios.

Innovation Solution

The method involves using environmental models to simulate actual environments, extracting relevant environmental features, and configuring synthetic data generation simulators to reflect these features, ensuring that the generated synthetic data matches the distribution and characteristics of actual data, thereby enhancing the quality and relevance of the synthetic data for AI training.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If virtual environment or statistical model is used for synthetic data generation, then data diversity and cost reduction are improved, but data quality and feature accuracy deteriorate

Engineering Contradiction:
Improvecost reductionVSAvoiddata quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent introduces domain knowledge as an intermediary component between the virtual environment and the synthetic data generation process. This domain knowledge acts as a mediator that guides the generation of synthetic data to better reflect actual data features, thereby improving data quality without sacrificing the cost benefits of virtual environment-based generation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent combines multiple components (virtual environment, statistical model, and domain knowledge) to create a composite synthetic data generation system. This composite approach leverages the strengths of each component: the virtual environment provides cost-effective data generation, the statistical model ensures data distribution characteristics, and domain knowledge enhances feature accuracy, resulting in high-quality synthetic data.

Inventive Principle:
Principle #40Composite materials

2Adaptability or versatility

If virtual environment or statistical model is used for synthetic data generation, then data diversity and cost reduction are improved, but feature reflection accuracy deteriorates

Engineering Contradiction:
Improvedata diversityVSAvoidfeature reflection accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

Domain knowledge serves as an intermediary that bridges the gap between diverse synthetic data generation and accurate feature reflection. It guides the virtual environment to generate data that not only exhibits diversity but also accurately reflects the features of actual data.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent utilizes domain knowledge to adjust and optimize parameters in the synthetic data generation process. By changing key parameters based on domain expertise, the system maintains data diversity while significantly improving feature reflection accuracy.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If actual data is collected in data-scarce industrial fields, then data quality is improved, but cost, time, and security challenges increase

Engineering Contradiction:
Improvedata qualityVSAvoiddata collection complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent creates copies of actual data through synthetic data generation. Instead of collecting expensive and complex actual data from data-scarce environments, the system generates synthetic copies that preserve the essential features and characteristics of actual data, thereby avoiding the complexities of real-world data collection while maintaining data quality.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

Domain knowledge acts as an intermediary that enables high-quality synthetic data generation without requiring actual data collection. It allows the system to bypass the complexities of data collection in data-scarce fields by directly generating synthetic data with appropriate features based on domain expertise.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20240095427A1Apparatus and method of synthetic data generation using environmental models
Publication Date: 2024.03.21 KOREA UNIV OF TECH & EDUCATION IND UNIV COOPERATION FOUND
  • US20240095427A1 patent drawing
  • US20240095427A1 patent drawing
  • US20240095427A1 patent drawing

AI summary

Disclosed are a method and an apparatus of synthetic data generation for training an artificial intelligence model. The method includes: reading a target environmental model simulating a target environment to generate synthetic data among a plurality of environmental models simulating a plurality of actual environments, respectively; extracting an environmental feature which influences generation of data in the target environment; configuring a synthetic data generation function of a synthetic data generation simulator for the target environmental model so that the synthetic data reflects the environmental feature; and generating the synthetic data by using the synthetic data generation simulator for the target environmental model, and has an effect of being capable of generating high-quality learning synthetic data which may be used for training an artificial intelligence model.