Generative Framework for Synthetic Data Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data generation methods for AI and ML models are inadequate, as they require large amounts of real data, which can be privacy concerns and are not always available.
Innovation Solution
The development of an apparatus and method for synthetic data generation using a generative framework that includes a processor and memory configured to input data into a generative framework, which utilizes various Gen AI architectures and hierarchical modeling algorithms to produce synthetic data that mimics the original data while ensuring anonymization and maintaining inter-table referential integrity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real data is used to train AI and ML models, then model training quality is improved, but privacy concerns arise and data availability is limited
Solution Approach 1:
The patent applies the copying principle by generating synthetic data that replicates the statistical properties and patterns of real data without copying actual sensitive information. The generative framework creates artificial data samples that mimic the distribution, relationships, and features of the original data while ensuring complete anonymization, thus resolving the contradiction between needing high-quality training data and protecting privacy.
2Reliability
If real data is used to train AI and ML models, then model training quality is improved, but data availability is reduced
Solution Approach 1:
The patent applies the self-service principle by enabling the system to generate its own training data from existing data samples. The generative framework processes input data and automatically produces synthetic data samples that can be used for training, eliminating the dependency on external real data sources and expanding data availability without requiring additional real data collection.
Solution Approach 2:
The patent creates synthetic data copies that replicate the statistical properties and patterns of real data. By generating multiple synthetic samples from a limited set of real data, the system effectively multiplies data availability while maintaining training quality, as the synthetic copies preserve the essential data distribution and relationships needed for model training.
3Object-affected harmful factors
If synthetic data is generated using generative frameworks, then privacy is protected and data availability is improved, but data generation complexity increases
Solution Approach 1:
The patent applies segmentation by dividing the data generation process into distinct functional components: data input processing, generative framework execution, and synthetic data output. The generative framework itself is segmented into different generation modes (first category and second category of synthetic data generation), allowing complex privacy-preserving data synthesis to be broken down into manageable, modular operations that reduce overall system complexity.
Data Source
AI summary
In an embodiment, an apparatus for synthetic data generation is presented. The apparatus includes a processor and a memory communicatively connected to the processor. The memory contains instructions configured to the processor to receive data. The processor is configured to input the data into a generative framework. The generative framework includes a first category of synthetic data generation and a second category of synthetic data generation. The generative framework is configured to input data an output synthetic data through at least a category of synthetic data generation. The processor is configured to generate, based on the generative framework, synthetic data from the received data.


