Synthetic Data Generation for Computer-Based Reasoning Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Computer-based reasoning systems face challenges in acquiring extensive and specific training data, particularly for rare scenarios, which can be costly, difficult, or dangerous to obtain, leading to inefficiencies in training and data anonymization.

Innovation Solution

The techniques involve generating synthetic data using existing training data and target surprisal, where undetermined features are conditioned on previous values or derivatives, and feature bounds are used to ensure data validity and anonymity, allowing for the creation of diverse and anonymized training data sets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If extensive real training data is collected for rare scenarios, then the training quality improves, but the cost and difficulty increase significantly

Engineering Contradiction:
Improvetraining qualityVSAvoiddata acquisition cost
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent generates synthetic training data by copying and transforming existing real data through conditional probability models. Instead of collecting expensive rare scenario data directly, the system creates artificial copies that preserve the statistical properties and relationships of the original data, enabling training on rare events without incurring the high costs of actual data collection

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system changes the parameters of existing data by applying conditional probability distributions to generate variations. By modifying data parameters through controlled random sampling from conditional distributions, the system creates diverse synthetic scenarios that maintain the underlying patterns while providing the variety needed for robust training

Inventive Principle:
Principle #35Parameter changes

2Reliability

If real data is used for training, then the training effectiveness improves, but data anonymity and privacy protection become problematic

Engineering Contradiction:
Improvetraining effectivenessVSAvoiddata privacy risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic copies of real data that preserve the statistical properties and learning signals needed for effective training, while eliminating the privacy risks associated with using actual personal or sensitive information. The synthetic data maintains the structural relationships and distributions of the original data without containing any real individual's information

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If more training data is collected to cover all scenarios, then the model coverage improves, but the data acquisition time and resources increase

Engineering Contradiction:
Improvemodel coverageVSAvoiddata acquisition time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system efficiently generates diverse training scenarios by changing parameters through conditional probability sampling. Instead of collecting data for each scenario individually over time, the system rapidly generates varied scenarios by sampling from conditional distributions, dramatically reducing the time required to achieve comprehensive model coverage

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent performs preliminary analysis of the data distribution and conditional relationships before generating synthetic data. By pre-computing the conditional probability structures from existing data, the system prepares the framework for rapidly generating diverse scenarios without needing to collect each scenario's data in advance

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12008446B2Conditioned synthetic data generation in computer-based reasoning systems
Publication Date: 2024.06.11 HOWSO INC
  • US12008446B2 patent drawing
  • US12008446B2 patent drawing
  • US12008446B2 patent drawing

AI summary

Techniques for synthetic data generation in computer-based reasoning systems are discussed and include receiving a request for generation of synthetic data based on a set of training data cases. One or more focal training data cases are determined. For undetermined features (either all of them or those that are not subject to conditions), a value for the feature is determined based on the focal cases. In some embodiments, the generation of synthetic data may be conditioned on values of features, preserved features, such as unique identifiers, previous-in-time features, and using the other techniques discussed herein.