Synthetic Data Generation for Computer-Based Reasoning Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computer-based reasoning systems face challenges in acquiring extensive and specific training data, particularly for rare scenarios, which can be costly, difficult, or dangerous to obtain, leading to inefficiencies in training and data anonymization.
Innovation Solution
The techniques involve generating synthetic data using existing training data and target surprisal, where undetermined features are conditioned on previous values or derivatives, and feature bounds are used to ensure data validity and anonymity, allowing for the creation of diverse and anonymized training data sets.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If extensive real training data is collected for rare scenarios, then the training quality improves, but the cost and difficulty increase significantly
Solution Approach 1:
The patent generates synthetic training data by copying and transforming existing real data through conditional probability models. Instead of collecting expensive rare scenario data directly, the system creates artificial copies that preserve the statistical properties and relationships of the original data, enabling training on rare events without incurring the high costs of actual data collection
Solution Approach 2:
The system changes the parameters of existing data by applying conditional probability distributions to generate variations. By modifying data parameters through controlled random sampling from conditional distributions, the system creates diverse synthetic scenarios that maintain the underlying patterns while providing the variety needed for robust training
2Reliability
If real data is used for training, then the training effectiveness improves, but data anonymity and privacy protection become problematic
Solution Approach 1:
The patent creates synthetic copies of real data that preserve the statistical properties and learning signals needed for effective training, while eliminating the privacy risks associated with using actual personal or sensitive information. The synthetic data maintains the structural relationships and distributions of the original data without containing any real individual's information
3Adaptability or versatility
If more training data is collected to cover all scenarios, then the model coverage improves, but the data acquisition time and resources increase
Solution Approach 1:
The system efficiently generates diverse training scenarios by changing parameters through conditional probability sampling. Instead of collecting data for each scenario individually over time, the system rapidly generates varied scenarios by sampling from conditional distributions, dramatically reducing the time required to achieve comprehensive model coverage
Solution Approach 2:
The patent performs preliminary analysis of the data distribution and conditional relationships before generating synthetic data. By pre-computing the conditional probability structures from existing data, the system prepares the framework for rapidly generating diverse scenarios without needing to collect each scenario's data in advance
Data Source
AI summary
Techniques for synthetic data generation in computer-based reasoning systems are discussed and include receiving a request for generation of synthetic data based on a set of training data cases. One or more focal training data cases are determined. For undetermined features (either all of them or those that are not subject to conditions), a value for the feature is determined based on the focal cases. In some embodiments, the generation of synthetic data may be conditioned on values of features, preserved features, such as unique identifiers, previous-in-time features, and using the other techniques discussed herein.


