Synthetic Data Generation for Rare Scenario Training

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Computer-based reasoning systems face challenges in acquiring extensive and specific training data, particularly for rare scenarios, which are costly, difficult, or dangerous to create, and may not be available due to privacy concerns.

Innovation Solution

The techniques involve generating synthetic data using existing training data and optionally target surprisal, ensuring that the synthetic data meets specific conditions. This includes determining distributions for undetermined features, conditioning values on previous values or derivatives, and applying feature bounds to generate values that differ from existing data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If extensive real training data is collected for rare scenarios, then the training data coverage is improved, but the cost and difficulty of data acquisition increases significantly

Engineering Contradiction:
Improvetraining data coverageVSAvoiddata acquisition cost
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The patent generates synthetic training data by copying and transforming existing real data through generative models. Instead of collecting expensive rare scenario data directly, the system creates artificial copies that simulate rare events by learning from abundant common scenario data, thereby reducing data acquisition costs while maintaining training data coverage.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms existing training data by changing parameters such as adding noise, adjusting distributions, and modifying feature values to generate synthetic rare scenario data. This allows the system to create diverse training examples from limited real data by varying data parameters rather than collecting new data.

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If real training data is used to maintain data quality, then the training accuracy is improved, but privacy concerns and anonymity requirements are violated

Engineering Contradiction:
Improvetraining data qualityVSAvoidprivacy risk
Core Design Contradiction:
Manufacturing precisionVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of real training data that preserve statistical properties and learning value while removing personally identifiable information. The generative model learns from real data and produces artificial data copies that maintain data quality for training purposes without containing actual private information, thus resolving the privacy conflict.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces a generative model as an intermediary between real private data and training processes. This intermediary learns from real data and generates synthetic data, acting as a buffer that protects privacy while maintaining data utility. The intermediary transforms sensitive real data into safe synthetic representations.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If more training data is collected to improve model performance, then the prediction accuracy is improved, but the time and resources required for data collection increase

Engineering Contradiction:
Improveprediction accuracyVSAvoiddata collection time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary data generation by creating synthetic training data in advance using generative models trained on existing data. Instead of collecting data on-demand during deployment, the system pre-generates extensive synthetic datasets that can be used immediately for training and testing, eliminating time-consuming data collection during critical phases.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent efficiently generates diverse training data by changing parameters of existing data through the generative model. This allows rapid creation of large volumes of synthetic training examples with varied characteristics, significantly reducing the time required compared to traditional data collection methods while improving model reliability.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250173590A1Identifier Contribution Allocation in Synthetic Data Generation in Computer-Based Reasoning Systems
Publication Date: 2025.05.29 HOWSO INC
  • US20250173590A1 patent drawing
  • US20250173590A1 patent drawing
  • US20250173590A1 patent drawing

AI summary

Techniques for synthetic data generation in computer-based reasoning systems are discussed and include receiving a request for generation of synthetic data based on a set of training data cases. One or more focal training data cases are determined. For undetermined features (either all of them or those that are not subject to conditions), a value for the feature is determined based on the focal cases. In some embodiments, the generated synthetic data may be checked for similarity against the training data, and if similarity conditions are met, it may be modified (e.g., resampled), removed, and/or replaced.