Synthetic Data Generator for Data Lake Diversity
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Data lakes often lack a diverse selection of data types and use cases, which limits the precision, accuracy, and efficiency of tasks such as machine learning model training, as source systems may not have the resources to generate sufficient synthetic data to cover various datatypes and scenarios.
Innovation Solution
A synthetic data generator system that identifies identifier and relationship fields in seed data samples to create synthetic data samples with synthetically generated values, propagating changes hierarchically and sending them to target systems for tasks like reporting, visualization, and machine learning, thereby supplementing the data volume and diversity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If source systems generate synthetic data to cover various datatypes and scenarios, then data diversity and volume are improved, but the resources required to generate such data increase
Solution Approach 1:
The patent applies the copying principle by creating synthetic data samples that replicate the structure and relationships of real data without requiring the source system to generate actual diverse data. The synthetic data generator copies the schema and relationships from seed data samples to produce new data that mimics the original data's characteristics, thereby increasing data volume and diversity without the resource burden of generating authentic diverse data from source systems.
Solution Approach 2:
The patent introduces a synthetic data generator as an intermediary component between source systems and target systems. This intermediary generates synthetic data based on seed data samples from source systems and delivers it to target systems, allowing the source system to maintain low resource requirements while still providing diverse data to support machine learning tasks.
2Adaptability or versatility
If synthetic data is generated based on seed data samples, then data diversity is improved, but the complexity of the data generation process increases
Solution Approach 1:
The patent applies segmentation by dividing the data generation process into distinct operational steps: identifying identifier fields and relationship fields in seed data samples, generating synthetic values for identifier fields, preserving relationship field values, and assembling synthetic data samples. This segmentation simplifies the overall process by breaking down the complex task of generating diverse synthetic data into manageable, sequential operations that can be automated.
Solution Approach 2:
The patent utilizes parameter changes by modifying only the identifier field values while preserving relationship field values when generating synthetic data samples. This selective parameter change approach maintains the structural integrity and relationships of the original data while introducing diversity through new identifier values, thereby achieving data diversity without complex generation logic.
3Measurement precision
If synthetic data samples are sent to target systems for machine learning tasks, then model training accuracy is improved, but data transmission and processing requirements increase
Solution Approach 1:
The patent applies copying by creating synthetic data samples that replicate the essential structure and relationships of real data used for training models. These synthetic copies maintain the necessary patterns and relationships for accurate model training while allowing for flexible generation of large volumes of data without the constraints of transmitting and processing actual diverse real-world data.
Data Source
AI summary
A method may include identifying an identifier field included in a first datatype of a seed data sample associated with a source system. The identifier field may store a first value that enables a differentiation between different instances of the first datatype. A relationship field, which stores a second value that define a relationship between the first datatype and a second data type, may be identified. A synthetic data sample may be generated by populating the identifier field of the synthetic data sample with a synthetically generated value and the relationship field of the synthetic data sample with the second value. The synthetic data sample may be sent to a target system to enable a performance of a task at the target system. The synthetic data sample may supplement a volume and/or a diversity of the data that occurs organically at the source system.


