Synthetic Data Generator for Data Lake Diversity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Data lakes often lack a diverse selection of data types and use cases, which limits the precision, accuracy, and efficiency of tasks such as machine learning model training, as source systems may not have the resources to generate sufficient synthetic data to cover various datatypes and scenarios.

Innovation Solution

A synthetic data generator system that identifies identifier and relationship fields in seed data samples to create synthetic data samples with synthetically generated values, propagating changes hierarchically and sending them to target systems for tasks like reporting, visualization, and machine learning, thereby supplementing the data volume and diversity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If source systems generate synthetic data to cover various datatypes and scenarios, then data diversity and volume are improved, but the resources required to generate such data increase

Engineering Contradiction:
Improvedata volume and diversityVSAvoidsource system resources
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent applies the copying principle by creating synthetic data samples that replicate the structure and relationships of real data without requiring the source system to generate actual diverse data. The synthetic data generator copies the schema and relationships from seed data samples to produce new data that mimics the original data's characteristics, thereby increasing data volume and diversity without the resource burden of generating authentic diverse data from source systems.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces a synthetic data generator as an intermediary component between source systems and target systems. This intermediary generates synthetic data based on seed data samples from source systems and delivers it to target systems, allowing the source system to maintain low resource requirements while still providing diverse data to support machine learning tasks.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If synthetic data is generated based on seed data samples, then data diversity is improved, but the complexity of the data generation process increases

Engineering Contradiction:
Improvedata diversityVSAvoiddata generation process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent applies segmentation by dividing the data generation process into distinct operational steps: identifying identifier fields and relationship fields in seed data samples, generating synthetic values for identifier fields, preserving relationship field values, and assembling synthetic data samples. This segmentation simplifies the overall process by breaking down the complex task of generating diverse synthetic data into manageable, sequential operations that can be automated.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent utilizes parameter changes by modifying only the identifier field values while preserving relationship field values when generating synthetic data samples. This selective parameter change approach maintains the structural integrity and relationships of the original data while introducing diversity through new identifier values, thereby achieving data diversity without complex generation logic.

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If synthetic data samples are sent to target systems for machine learning tasks, then model training accuracy is improved, but data transmission and processing requirements increase

Engineering Contradiction:
Improvemodel training accuracyVSAvoiddata transmission volume
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies copying by creating synthetic data samples that replicate the essential structure and relationships of real data used for training models. These synthetic copies maintain the necessary patterns and relationships for accurate model training while allowing for flexible generation of large volumes of data without the constraints of transmitting and processing actual diverse real-world data.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS12073246B2Data curation with synthetic data generation
Publication Date: 2024.08.27 SAP SE
  • US12073246B2 patent drawing
  • US12073246B2 patent drawing
  • US12073246B2 patent drawing

AI summary

A method may include identifying an identifier field included in a first datatype of a seed data sample associated with a source system. The identifier field may store a first value that enables a differentiation between different instances of the first datatype. A relationship field, which stores a second value that define a relationship between the first datatype and a second data type, may be identified. A synthetic data sample may be generated by populating the identifier field of the synthetic data sample with a synthetically generated value and the relationship field of the synthetic data sample with the second value. The synthetic data sample may be sent to a target system to enable a performance of a task at the target system. The synthetic data sample may supplement a volume and/or a diversity of the data that occurs organically at the source system.