Data Catalog System for Synthetic Dataset Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data cataloging systems have limitations in providing comprehensive and efficient access to data assets, especially in managing distributed and dynamic datasets, and face challenges with data privacy concerns, leading to restricted usage and inaccuracy in machine-learning training due to missing or restricted data.

Innovation Solution

A data catalog system that automatically generates synthetic datasets based on original datasets, using metadata information to fill gaps, replace restricted data, and augment datasets, employing machine-learning techniques for improved accuracy and accessibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If data catalog systems provide comprehensive access to data assets, then data accessibility and analytics speed are improved, but data privacy risks and governance challenges worsen

Engineering Contradiction:
Improvedata accessibilityVSAvoiddata privacy risks
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The system generates synthetic copies of real data that preserve statistical properties and patterns while eliminating sensitive information. These synthetic datasets can be freely accessed and used for analytics and machine learning without privacy concerns, as they are generated from and reflect the original data distribution without containing actual PII or restricted information.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The data catalog system acts as an intermediary between data sources and users. It provides a controlled layer that enables access to data while managing privacy risks through synthetic data generation, data lineage tracking, and governance policies. The system mediates between the need for comprehensive data access and the need to protect privacy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If synthetic datasets are generated to replace restricted data, then data privacy is improved, but data accuracy and training quality may worsen

Engineering Contradiction:
Improvedata privacyVSAvoidtraining accuracy
Core Design Contradiction:
Object-affected harmful factorsVSMeasurement precision

Solution Approach 1:

The system transforms data parameters by generating synthetic values that match the statistical distribution, patterns, and relationships of the original data. By carefully controlling generation parameters and preserving key data characteristics, the system maintains training accuracy while achieving privacy protection. The synthetic data preserves the essential parameters needed for accurate machine learning models.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If data is centralized for unified cataloging, then data governance and organization are improved, but system complexity and implementation difficulty worsen

Engineering Contradiction:
Improvedata governance structureVSAvoidsystem implementation
Core Design Contradiction:
Device complexityVSEase of manufacture

Solution Approach 1:

The data catalog system is designed as a universal platform that performs multiple functions: data discovery, cataloging, synthetic data generation, lineage tracking, and governance enforcement. By consolidating these functions into a single system, the patent reduces overall implementation complexity compared to separate specialized systems, while providing comprehensive data management capabilities.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Quantity of substance

If existing data is used for machine-learning training, then training data availability is improved, but data quality and completeness worsen due to missing or restricted data

Engineering Contradiction:
Improvetraining data availabilityVSAvoiddata quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The system performs preliminary action by generating synthetic training data before actual training begins. The synthetic datasets are prepared in advance, ensuring data quality and completeness issues are resolved prior to model training. This preliminary generation of high-quality synthetic data eliminates the need to work with incomplete or restricted real data during the training process.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11687568B2Data catalog system for generating synthetic datasets
Publication Date: 2023.06.27 ORACLE INT CORP
  • US11687568B2 patent drawing
  • US11687568B2 patent drawing
  • US11687568B2 patent drawing

AI summary

A data catalog system that is configured to automatically generate synthetic datasets based upon original datasets cataloged by the data catalog system, wherein each synthetic dataset comprises synthetic data that is generated using one or more data generation techniques. The data catalog system may access an original dataset and harvest associated metadata information and generate catalog information for the original dataset. The data catalog system may then generate a synthetic dataset based upon the original dataset and its harvested metadata information. The data catalog system may also generate catalog information for the generated synthetic dataset. The catalog information generated for the original dataset may be updated to refer to the newly generated synthetic dataset and its catalog information. The catalog information generated for the synthetic dataset may include references to the original dataset and its catalog information to inform a user of the original dataset about the synthetic dataset.