Synthetic Data Generation via Datapoint Density Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Bioinformatics research faces challenges with High Dimensionality but Low Sample Size (HDLSS) datasets, particularly in class imbalance scenarios, where existing methods like SMOTE and GANs fail to adequately address data imbalance and privacy concerns, leading to biased performance and identifiability risks.
Innovation Solution
The use of Datapoint Density Estimation (DDE) for synthetic data generation, which samples from local kernel density estimates between datapoints and incorporates privacy filtering to minimize identifiability risk, effectively generating synthetic datasets that preserve nonstandard distributions and reduce re-identification risks.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If SMOTE or GANs are used for synthetic data generation, then data imbalance is addressed, but identifiability risk and privacy concerns increase
Solution Approach 1:
The patent segments the synthetic data generation process into two distinct phases: first generating synthetic data using SMOTE or GANs to address class imbalance, then applying a separate privacy filtering step that segments and removes identifiable information. This two-stage approach allows the system to benefit from imbalance correction while subsequently eliminating privacy risks through the filtering mechanism.
2Device complexity
If existing methods are used for HDLSS data, then processing is simpler, but statistical utility and distribution fidelity decrease
Solution Approach 1:
The patent applies preliminary action by first assessing the HDLSS characteristics and data distribution before selecting and applying appropriate synthesis methods. The system performs preliminary density estimation and identifies suitable regions for synthetic data generation, ensuring that the subsequent processing is optimized for the specific data characteristics rather than applying generic methods.
3Quantity of substance
If more synthetic data is generated to improve sample size, then machine learning performance improves, but privacy risk increases
Solution Approach 1:
The patent introduces privacy filtering as an intermediary mechanism between synthetic data generation and final data usage. This intermediary step processes the generated synthetic data to remove or mask identifiable information, allowing the system to generate sufficient sample sizes for machine learning while maintaining privacy protection through the filtering intermediary.
Data Source
AI summary
Techniques for generating synthetic data are described. An exemplary approach includes receiving one or more requests to generate synthetic data based on a first dataset; generating the synthetic dataset is generated according to the request by choosing a set of synthetic datapoints between pairs of datapoints of the first dataset along a line connecting them while sampling a likely value of a local probability distribution; and providing the synthetic dataset as configured by the request.


