Synthetic Data Generation via Datapoint Density Estimation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Bioinformatics research faces challenges with High Dimensionality but Low Sample Size (HDLSS) datasets, particularly in class imbalance scenarios, where existing methods like SMOTE and GANs fail to adequately address data imbalance and privacy concerns, leading to biased performance and identifiability risks.

Innovation Solution

The use of Datapoint Density Estimation (DDE) for synthetic data generation, which samples from local kernel density estimates between datapoints and incorporates privacy filtering to minimize identifiability risk, effectively generating synthetic datasets that preserve nonstandard distributions and reduce re-identification risks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If SMOTE or GANs are used for synthetic data generation, then data imbalance is addressed, but identifiability risk and privacy concerns increase

Engineering Contradiction:
Improvedata balanceVSAvoididentifiability risk
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent segments the synthetic data generation process into two distinct phases: first generating synthetic data using SMOTE or GANs to address class imbalance, then applying a separate privacy filtering step that segments and removes identifiable information. This two-stage approach allows the system to benefit from imbalance correction while subsequently eliminating privacy risks through the filtering mechanism.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If existing methods are used for HDLSS data, then processing is simpler, but statistical utility and distribution fidelity decrease

Engineering Contradiction:
Improveprocessing complexityVSAvoidstatistical utility
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent applies preliminary action by first assessing the HDLSS characteristics and data distribution before selecting and applying appropriate synthesis methods. The system performs preliminary density estimation and identifies suitable regions for synthetic data generation, ensuring that the subsequent processing is optimized for the specific data characteristics rather than applying generic methods.

Inventive Principle:
Principle #10Preliminary action

3Quantity of substance

If more synthetic data is generated to improve sample size, then machine learning performance improves, but privacy risk increases

Engineering Contradiction:
Improvesample sizeVSAvoidprivacy risk
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent introduces privacy filtering as an intermediary mechanism between synthetic data generation and final data usage. This intermediary step processes the generated synthetic data to remove or mask identifiable information, allowing the system to generate sufficient sample sizes for machine learning while maintaining privacy protection through the filtering intermediary.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12153710B2Synthetic data generation
Publication Date: 2024.11.26 AMAZON TECH INC
  • US12153710B2 patent drawing
  • US12153710B2 patent drawing
  • US12153710B2 patent drawing

AI summary

Techniques for generating synthetic data are described. An exemplary approach includes receiving one or more requests to generate synthetic data based on a first dataset; generating the synthetic dataset is generated according to the request by choosing a set of synthetic datapoints between pairs of datapoints of the first dataset along a line connecting them while sampling a likely value of a local probability distribution; and providing the synthetic dataset as configured by the request.