Generative Database Synthetic Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Businesses face challenges in providing timely data modeling and analysis due to regulatory and privacy hurdles when accessing sensitive datasets, requiring manual and time-consuming data de-sensitization processes.

Innovation Solution

A system and method for on-demand generation of anonymized and privacy-compliant synthetic datasets using a generative database, which identifies a query for synthetic data samples statistically representative of a target sensitive dataset, constructs a generative model election request, searches for an appropriate generative model based on efficacy metrics, and generates synthetic data that preserves privacy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual data de-sensitization is performed to ensure privacy compliance, then privacy protection is improved, but processing time increases significantly

Engineering Contradiction:
Improveprivacy protectionVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent creates synthetic copies of sensitive datasets that preserve statistical properties and relationships without containing actual sensitive information. Generative models produce artificial data samples that replicate the structure, distributions, and correlations of the original data, enabling analysis without exposing real PII or PHI.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces manual mechanical de-sensitization processes with automated machine learning systems. Instead of human analysts manually identifying and removing sensitive elements, generative models automatically synthesize privacy-preserving data, dramatically reducing processing time from weeks/months to hours.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If manual data de-sensitization is performed to ensure privacy compliance, then privacy protection is improved, but productivity decreases

Engineering Contradiction:
Improveprivacy complianceVSAvoiddata analysis throughput
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

Synthetic data copies enable multiple analysis teams to work simultaneously on the same dataset without competing for access to sensitive information. Each team receives identical privacy-preserving synthetic data, eliminating bottlenecks and enabling parallel processing that dramatically increases overall productivity.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The generative database system automatically handles privacy compliance without requiring manual intervention. The system self-manages the synthesis process, model selection, and data generation, freeing personnel from time-consuming de-sensitization tasks and enabling them to focus on higher-value analysis work.

Inventive Principle:
Principle #25Self-service

3Reliability

If synthetic data is generated to preserve privacy, then privacy compliance is improved, but data representativeness may be compromised

Engineering Contradiction:
Improveprivacy complianceVSAvoiddata statistical accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent employs feedback mechanisms where generative models are trained and evaluated against the original data distributions, with continuous refinement based on statistical fidelity metrics. The system monitors and adjusts synthetic data quality to ensure it maintains accurate representations of underlying data patterns while preserving privacy.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically adjusts generation parameters such as sample size, diversity controls, and statistical constraints to optimize the balance between privacy protection and data representativeness. By tuning these parameters, the system can produce synthetic data with varying degrees of fidelity depending on specific analytical needs.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11922289B1Machine learning-based systems and methods for on-demand generation of anonymized and privacy-enabled synthetic datasets
Publication Date: 2024.03.05 SUBSALT INC
  • US11922289B1 patent drawing
  • US11922289B1 patent drawing
  • US11922289B1 patent drawing

AI summary

A system and method for generating synthetic datasets includes receiving, via an application programming interface (API) of a remote generative database service, a generative database query for obtaining synthetic data samples statistically representative of a sensitive dataset, searching a generative model data structure comprising a plurality of generative model nexuses based on a generative model election request derived from the generative database query, wherein the searching returns a generative model for fulfilling the generative database query, generating a synthetic dataset using the generative model returned from the searching based on a plurality of generative query parameters extracted from the generative database query, and returning the synthetic dataset as a result to the generative database query.