Data Catalog System for Synthetic Dataset Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data cataloging systems have limitations in providing comprehensive and efficient access to data assets, especially in managing distributed and dynamic datasets, and face challenges with data privacy concerns, leading to restricted usage and inaccuracy in machine-learning training due to missing or restricted data.
Innovation Solution
A data catalog system that automatically generates synthetic datasets based on original datasets, using metadata information to fill gaps, replace restricted data, and augment datasets, employing machine-learning techniques for improved accuracy and accessibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If data catalog systems provide comprehensive access to data assets, then data accessibility and analytics speed are improved, but data privacy risks and governance challenges worsen
Solution Approach 1:
The system generates synthetic copies of real data that preserve statistical properties and patterns while eliminating sensitive information. These synthetic datasets can be freely accessed and used for analytics and machine learning without privacy concerns, as they are generated from and reflect the original data distribution without containing actual PII or restricted information.
Solution Approach 2:
The data catalog system acts as an intermediary between data sources and users. It provides a controlled layer that enables access to data while managing privacy risks through synthetic data generation, data lineage tracking, and governance policies. The system mediates between the need for comprehensive data access and the need to protect privacy.
2Object-affected harmful factors
If synthetic datasets are generated to replace restricted data, then data privacy is improved, but data accuracy and training quality may worsen
Solution Approach 1:
The system transforms data parameters by generating synthetic values that match the statistical distribution, patterns, and relationships of the original data. By carefully controlling generation parameters and preserving key data characteristics, the system maintains training accuracy while achieving privacy protection. The synthetic data preserves the essential parameters needed for accurate machine learning models.
3Device complexity
If data is centralized for unified cataloging, then data governance and organization are improved, but system complexity and implementation difficulty worsen
Solution Approach 1:
The data catalog system is designed as a universal platform that performs multiple functions: data discovery, cataloging, synthetic data generation, lineage tracking, and governance enforcement. By consolidating these functions into a single system, the patent reduces overall implementation complexity compared to separate specialized systems, while providing comprehensive data management capabilities.
4Quantity of substance
If existing data is used for machine-learning training, then training data availability is improved, but data quality and completeness worsen due to missing or restricted data
Solution Approach 1:
The system performs preliminary action by generating synthetic training data before actual training begins. The synthetic datasets are prepared in advance, ensuring data quality and completeness issues are resolved prior to model training. This preliminary generation of high-quality synthetic data eliminates the need to work with incomplete or restricted real data during the training process.
Data Source
AI summary
A data catalog system that is configured to automatically generate synthetic datasets based upon original datasets cataloged by the data catalog system, wherein each synthetic dataset comprises synthetic data that is generated using one or more data generation techniques. The data catalog system may access an original dataset and harvest associated metadata information and generate catalog information for the original dataset. The data catalog system may then generate a synthetic dataset based upon the original dataset and its harvested metadata information. The data catalog system may also generate catalog information for the generated synthetic dataset. The catalog information generated for the original dataset may be updated to refer to the newly generated synthetic dataset and its catalog information. The catalog information generated for the synthetic dataset may include references to the original dataset and its catalog information to inform a user of the original dataset about the synthetic dataset.


