Structured Synthetic Data Generation via Bayesian Network Modeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing synthetic data generation methods struggle to ensure consistency in the joint probability distribution and feature relationships of generated synthetic data with original data, fail to simultaneously process discrete and continuous features, and cannot generate synthetic data under specific conditions for various application scenarios.
Innovation Solution
A structured synthetic data generation system and method that utilizes a data preprocessing unit to transform original data into vector representations and model Bayesian networks for feature relationships, and a training and generation unit to train a generative adversarial network-based model for generating synthetic data records that preserve feature relationships and can include both discrete and continuous features.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional data anonymization techniques are used to protect privacy, then privacy protection is improved, but data availability is greatly reduced
Solution Approach 1:
The patent creates synthetic data copies that replicate the statistical properties and relationships of original data without containing actual sensitive information. The synthetic data generation model learns the joint probability distribution and feature relationships from original data, then generates artificial records that preserve analytical utility while ensuring privacy protection, thus resolving the contradiction between privacy protection and data availability
Solution Approach 2:
The patent transforms the data representation by changing from original sensitive data to synthetic data with modified parameters. The synthetic data maintains the same structural characteristics, feature relationships, and statistical properties as original data, but with fundamentally different underlying values that cannot be reverse-engineered to reveal sensitive information, thereby achieving both privacy protection and data availability
2Productivity
If existing synthetic data generation methods are used, then synthetic data can be generated, but consistency in joint probability distribution and feature relationships with original data cannot be guaranteed
Solution Approach 1:
The patent employs adversarial training where a discriminator network provides feedback to the generator network. The discriminator evaluates whether synthetic data follows the same joint probability distribution and feature relationships as original data, and this feedback continuously refines the generator's output to improve consistency, thereby achieving both high productivity and manufacturing precision
Solution Approach 2:
The patent uses a dynamic training process where the generator and discriminator networks continuously adapt to each other. The generator learns to produce increasingly accurate synthetic data by responding to the discriminator's evaluations, allowing the system to dynamically converge toward perfect consistency in joint probability distribution and feature relationships while maintaining efficient generation capability
3Adaptability or versatility
If existing synthetic data generation methods are used, then some data can be generated, but the ability to process both discrete and continuous features simultaneously is lost
Solution Approach 1:
The patent designs a universal synthetic data generation model that can simultaneously handle both discrete and continuous features through a unified architecture. The model uses continuous relaxation techniques and appropriate probability distributions (categorical for discrete, Gaussian for continuous) within a single generative framework, enabling comprehensive processing of mixed feature types while maintaining high generation efficiency through parallel computation
4Productivity
If existing synthetic data generation methods are used, then general synthetic data can be produced, but synthetic data under specific conditions according to specific application scenarios cannot be generated
Solution Approach 1:
The patent performs preliminary learning of the overall data distribution and feature relationships from the complete original dataset during the training phase. This preliminary action enables the model to quickly generate high-quality synthetic data for general purposes, while the learned representations can be further conditioned or fine-tuned for specific application scenarios without requiring complete retraining, thus achieving both high productivity and adaptability
Data Source
AI summary
Disclosed are a structured synthetic data generation system and method. The structured synthetic data generation system comprises a data preprocessing unit and a training and generation unit. The data preprocessing unit is used for transforming each sample in original data into a vector representation and modeling a Bayesian network for describing a relation between features during the transformation process. The training and generation unit is used for training by means of the vector representation transformed from the original data to obtain a synthetic data generation model and generating a synthetic data record by means of the synthetic data generation model. The system and method provided by the invention can simultaneously generate synthetic data records including continuous features and discrete features; generated synthetic data are identical in data distribution and the relation between features with original data.


