Synthetic Transaction Data Generation for Privacy-Preserving Fraud Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Financial crime detection systems require large amounts of sensitive real customer data for training predictive models, but sharing such data is restricted due to privacy concerns, limiting the effectiveness of simulated fraudulent situation detection.

Innovation Solution

A method using unsupervised learning to create standard customer profiles from real customer data, generating synthetic transaction data that mimics real customer behavior without exposing sensitive information, and distributing a detection model across entities for training predictive models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If real customer data is used to train predictive models, then model effectiveness is improved, but privacy protection is compromised

Engineering Contradiction:
Improvemodel effectivenessVSAvoidprivacy risk
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic customer data that copies the statistical patterns and relationships of real customer data without containing actual sensitive information. The synthetic data is generated by learning the underlying distributions and correlations from real data, then reproducing these patterns with artificial records that maintain analytical utility while eliminating privacy risks.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic data as an intermediary between real customer data and the predictive model training process. This intermediary layer allows the model to learn from data that preserves the statistical properties needed for effective training while removing the direct connection to sensitive personal information, thus mediating between model effectiveness and privacy protection.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If simulated customer data is generated to replace real data, then privacy protection is improved, but model training effectiveness deteriorates

Engineering Contradiction:
Improveprivacy riskVSAvoidmodel training effectiveness
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent transforms the parameters and characteristics of real customer data into synthetic representations that preserve the essential statistical properties, distribution patterns, and relationships needed for model training. By carefully controlling the parameter transformations during synthetic data generation, the system maintains training effectiveness while achieving privacy protection.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The synthetic data generation process copies the structural patterns, correlations, and statistical relationships from real customer data, ensuring that the simulated data maintains the same analytical value for model training as the original data, thereby preventing deterioration in model training effectiveness.

Inventive Principle:
Principle #26Copying

3Loss of information

If standard customer profiles are created through unsupervised learning, then data utility is improved, but computational complexity increases

Engineering Contradiction:
Improvedata utilityVSAvoidcomputational complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments the customer data into distinct profiles or clusters through unsupervised learning, organizing the data into manageable groups based on shared characteristics. This segmentation reduces the complexity of processing individual records while preserving the utility of the data by maintaining the distinctive patterns of different customer segments.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The unsupervised learning process transforms the raw customer data by identifying and emphasizing key parameters that define different customer profiles, reducing dimensionality and complexity while retaining the essential information needed for data utility in model training.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11475468B2System and method for unsupervised abstraction of sensitive data for detection model sharing across entities
Publication Date: 2022.10.18 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11475468B2 patent drawing
  • US11475468B2 patent drawing
  • US11475468B2 patent drawing

AI summary

An abstraction system for generating a standard customer profile in a data processing system has a processing device and a memory. The abstraction system may receive customer data from a computing device over a network, the customer data including information for a plurality of customers and perform unsupervised learning on the customer data to produce a plurality of clusters of customers with a plurality of features in common, determine that a cluster represents a standard customer and store a plurality of standard customer profiles based on the determined standard customers. The abstraction system may also provide the standard customer profiles to a cognitive system for generating synthetic transaction data based on the standard customer, generate a detection model for detecting activity based on the synthetic transaction data, and distribute the detection model to each of the computing devices over the network.