Synthetic Data Generation via Network Topology Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The financial sector faces challenges in accessing granular data due to internal approval procedures and data privacy concerns, leading to latency in operationalizing data for analytics and risk assessment, particularly in fraud detection and understanding risk propagation across complex financial systems.
Innovation Solution
A method and system for generating synthetic data from aggregate datasets using network generation models that define personas, create network topologies, and distribute transactions, mimicking real data properties to reduce latency and security concerns.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If granular confidential datasets are accessed and shared, then data quality and analytical capability are improved, but data privacy security and access complexity worsen
Solution Approach 1:
The patent generates synthetic copies of granular confidential datasets that replicate statistical properties and patterns without containing actual sensitive information. These synthetic datasets serve as substitutes for real data, enabling analytical work while preserving privacy security.
Solution Approach 2:
The patent introduces an intermediary synthetic data generation system that mediates between the need for high-quality analytical data and the requirement for data privacy protection. This intermediary process transforms confidential data into protective synthetic representations.
2Reliability
If internal approval procedures are followed for data access, then data security is maintained, but data operationalization time increases
Solution Approach 1:
The patent performs preliminary synthetic data generation before actual analytical needs arise. By pre-generating synthetic datasets from aggregate statistics, the system eliminates the need for time-consuming approval procedures when data is needed, as the synthetic data can be freely distributed.
Solution Approach 2:
The patent enables organizations to self-generate their own synthetic datasets using publicly available aggregate statistics, eliminating the need to request access to confidential granular data through complex internal approval chains.
3Object-affected harmful factors
If synthetic data is generated from aggregate datasets, then data privacy is protected, but data granularity and detail are reduced
Solution Approach 1:
The patent transforms aggregate dataset parameters (statistics, distributions, relationships) into synthetic granular data points. By changing the parameter representation from aggregated sums to individual synthetic records, the system recovers granularity while maintaining privacy through statistical fidelity rather than actual data copying.
Data Source
AI summary
Disclosed is a method for generating synthetic data from an aggregate dataset related to a plurality of entities. The method comprises defining a set of personas, selecting a network generation model, defining a network topology and generating the synthetic data. The set of personas is defined based on a number of the plurality of entities and/or a type of at least one of the plurality of entities. The network generation model is selected based on the defined set of personas and the aggregate dataset related thereto. The generated network topology comprises nodes and links between nodes using the selected network generation model. The node represents one of the personas in the set of personas. The link between two nodes represent one or more transactions between the personas represented by the two nodes. The generated synthetic data is based on information about distribution of the one or more transactions.


