Synthetic Data Generation via Network Topology Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The financial sector faces challenges in accessing granular data due to internal approval procedures and data privacy concerns, leading to latency in operationalizing data for analytics and risk assessment, particularly in fraud detection and understanding risk propagation across complex financial systems.

Innovation Solution

A method and system for generating synthetic data from aggregate datasets using network generation models that define personas, create network topologies, and distribute transactions, mimicking real data properties to reduce latency and security concerns.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If granular confidential datasets are accessed and shared, then data quality and analytical capability are improved, but data privacy security and access complexity worsen

Engineering Contradiction:
Improvedata qualityVSAvoiddata privacy security
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent generates synthetic copies of granular confidential datasets that replicate statistical properties and patterns without containing actual sensitive information. These synthetic datasets serve as substitutes for real data, enabling analytical work while preserving privacy security.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces an intermediary synthetic data generation system that mediates between the need for high-quality analytical data and the requirement for data privacy protection. This intermediary process transforms confidential data into protective synthetic representations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If internal approval procedures are followed for data access, then data security is maintained, but data operationalization time increases

Engineering Contradiction:
Improvedata securityVSAvoiddata operationalization time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary synthetic data generation before actual analytical needs arise. By pre-generating synthetic datasets from aggregate statistics, the system eliminates the need for time-consuming approval procedures when data is needed, as the synthetic data can be freely distributed.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent enables organizations to self-generate their own synthetic datasets using publicly available aggregate statistics, eliminating the need to request access to confidential granular data through complex internal approval chains.

Inventive Principle:
Principle #25Self-service

3Object-affected harmful factors

If synthetic data is generated from aggregate datasets, then data privacy is protected, but data granularity and detail are reduced

Engineering Contradiction:
Improvedata privacyVSAvoiddata granularity
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The patent transforms aggregate dataset parameters (statistics, distributions, relationships) into synthetic granular data points. By changing the parameter representation from aggregated sums to individual synthetic records, the system recovers granularity while maintaining privacy through statistical fidelity rather than actual data copying.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11381467B2Method and system for generating synthetic data from aggregate dataset
Publication Date: 2022.07.05 FINANCIAL NETWORK ANALYTICS LTD
  • US11381467B2 patent drawing
  • US11381467B2 patent drawing
  • US11381467B2 patent drawing

AI summary

Disclosed is a method for generating synthetic data from an aggregate dataset related to a plurality of entities. The method comprises defining a set of personas, selecting a network generation model, defining a network topology and generating the synthetic data. The set of personas is defined based on a number of the plurality of entities and/or a type of at least one of the plurality of entities. The network generation model is selected based on the defined set of personas and the aggregate dataset related thereto. The generated network topology comprises nodes and links between nodes using the selected network generation model. The node represents one of the personas in the set of personas. The link between two nodes represent one or more transactions between the personas represented by the two nodes. The generated synthetic data is based on information about distribution of the one or more transactions.