Structured Synthetic Data Generation via Bayesian Network Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing synthetic data generation methods struggle to ensure consistency in the joint probability distribution and feature relationships of generated synthetic data with original data, fail to simultaneously process discrete and continuous features, and cannot generate synthetic data under specific conditions for various application scenarios.

Innovation Solution

A structured synthetic data generation system and method that utilizes a data preprocessing unit to transform original data into vector representations and model Bayesian networks for feature relationships, and a training and generation unit to train a generative adversarial network-based model for generating synthetic data records that preserve feature relationships and can include both discrete and continuous features.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional data anonymization techniques are used to protect privacy, then privacy protection is improved, but data availability is greatly reduced

Engineering Contradiction:
Improveprivacy protectionVSAvoiddata availability
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent creates synthetic data copies that replicate the statistical properties and relationships of original data without containing actual sensitive information. The synthetic data generation model learns the joint probability distribution and feature relationships from original data, then generates artificial records that preserve analytical utility while ensuring privacy protection, thus resolving the contradiction between privacy protection and data availability

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent transforms the data representation by changing from original sensitive data to synthetic data with modified parameters. The synthetic data maintains the same structural characteristics, feature relationships, and statistical properties as original data, but with fundamentally different underlying values that cannot be reverse-engineered to reveal sensitive information, thereby achieving both privacy protection and data availability

Inventive Principle:
Principle #35Parameter changes

2Productivity

If existing synthetic data generation methods are used, then synthetic data can be generated, but consistency in joint probability distribution and feature relationships with original data cannot be guaranteed

Engineering Contradiction:
Improvesynthetic data generation capabilityVSAvoidconsistency in joint probability distribution and feature relationships
Core Design Contradiction:
ProductivityVSManufacturing precision

Solution Approach 1:

The patent employs adversarial training where a discriminator network provides feedback to the generator network. The discriminator evaluates whether synthetic data follows the same joint probability distribution and feature relationships as original data, and this feedback continuously refines the generator's output to improve consistency, thereby achieving both high productivity and manufacturing precision

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent uses a dynamic training process where the generator and discriminator networks continuously adapt to each other. The generator learns to produce increasingly accurate synthetic data by responding to the discriminator's evaluations, allowing the system to dynamically converge toward perfect consistency in joint probability distribution and feature relationships while maintaining efficient generation capability

Inventive Principle:
Principle #15Dynamics

3Adaptability or versatility

If existing synthetic data generation methods are used, then some data can be generated, but the ability to process both discrete and continuous features simultaneously is lost

Engineering Contradiction:
Improveprocessing capability for feature typesVSAvoidcomprehensive data generation efficiency
Core Design Contradiction:
Adaptability or versatilityVSProductivity

Solution Approach 1:

The patent designs a universal synthetic data generation model that can simultaneously handle both discrete and continuous features through a unified architecture. The model uses continuous relaxation techniques and appropriate probability distributions (categorical for discrete, Gaussian for continuous) within a single generative framework, enabling comprehensive processing of mixed feature types while maintaining high generation efficiency through parallel computation

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Productivity

If existing synthetic data generation methods are used, then general synthetic data can be produced, but synthetic data under specific conditions according to specific application scenarios cannot be generated

Engineering Contradiction:
Improvegeneral synthetic data generationVSAvoidscenario-specific data generation capability
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The patent performs preliminary learning of the overall data distribution and feature relationships from the complete original dataset during the training phase. This preliminary action enables the model to quickly generate high-quality synthetic data for general purposes, while the learned representations can be further conditioned or fine-tuned for specific application scenarios without requiring complete retraining, thus achieving both high productivity and adaptability

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250028942A1Structured synthetic data generation system and method
Publication Date: 2025.01.23 HARBIN INST OF TECH (SHENZHEN) (INST OF SCIENCE & TECH INNOVATION HARBIN INST OF TECH)
  • US20250028942A1 patent drawing
  • US20250028942A1 patent drawing
  • US20250028942A1 patent drawing

AI summary

Disclosed are a structured synthetic data generation system and method. The structured synthetic data generation system comprises a data preprocessing unit and a training and generation unit. The data preprocessing unit is used for transforming each sample in original data into a vector representation and modeling a Bayesian network for describing a relation between features during the transformation process. The training and generation unit is used for training by means of the vector representation transformed from the original data to obtain a synthetic data generation model and generating a synthetic data record by means of the synthetic data generation model. The system and method provided by the invention can simultaneously generate synthetic data records including continuous features and discrete features; generated synthetic data are identical in data distribution and the relation between features with original data.