Synthetic Customer Data Generation for Financial Crime Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Financial crime detection systems require large amounts of realistic simulated customer transaction data to build effective predictive models, but sensitive real customer data is limited, making it challenging to generate sufficient training data without exposing privacy.

Innovation Solution

A cognitive system employing a multi-layered unsupervised clustering approach combined with interactive reinforcement learning to create standard customer transaction behaviors, which are then used to generate synthetic transaction data that mimics real customer behavior without revealing identifiable information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real customer data is used to train predictive models, then model accuracy is improved, but customer privacy is compromised

Engineering Contradiction:
Improvemodel accuracyVSAvoidprivacy risk
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic customer data that copies the statistical patterns and behavioral characteristics of real customer data without containing any actual personal information. This synthetic data is generated through clustering algorithms that replicate the distribution and relationships in the original data, enabling model training while completely protecting customer privacy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent extracts only the essential statistical patterns and structural relationships from real customer data, separating these useful characteristics from the sensitive personal information. By taking out only the necessary statistical properties needed for model training, the system achieves accurate model training without exposing any actual customer data.

Inventive Principle:
Principle #2Taking out (Extraction)

2Quantity of substance

If more simulated customer data is generated, then training data quantity is improved, but data realism may deteriorate

Engineering Contradiction:
Improvetraining data quantityVSAvoiddata realism
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent performs preliminary clustering and pattern extraction from real customer data before generating synthetic samples. By pre-processing the real data to identify and capture its underlying statistical structures and behavioral patterns, the system ensures that subsequently generated synthetic data maintains high realism while enabling large-scale data generation for comprehensive model training.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If real customer data is shared among banks, then collaborative model training is improved, but data security risks increase

Engineering Contradiction:
Improvecollaborative training capabilityVSAvoiddata security
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces synthetic customer data as an intermediary that enables collaborative model training among multiple banks without requiring direct sharing of sensitive real customer data. This intermediary synthetic data preserves the statistical characteristics needed for effective collaborative training while eliminating security risks associated with data sharing, as no actual personal information is transmitted or stored.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11488185B2System and method for unsupervised abstraction of sensitive data for consortium sharing
Publication Date: 2022.11.01 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11488185B2 patent drawing
  • US11488185B2 patent drawing
  • US11488185B2 patent drawing

AI summary

An abstraction system for generating a standard customer profile in a data processing system has a processing device and a memory. The abstraction system may receive customer data from a computing device over a network and perform unsupervised learning on the customer data to produce a plurality of clusters of customers with a first feature in common. The abstraction system performs unsupervised learning on the plurality of clusters of customers to produce a plurality of sub-clusters of customers with a second feature in common, and repeats the unsupervised learning on the plurality of sub-clusters produced to produce further sub-clusters with a plurality of features in common. The abstraction system determines that a sub-cluster represents a standard customer and stores a plurality of standard customer profiles based on the determined standard customers. The abstraction system provides the standard customer profiles to a cognitive system for generating synthetic transaction data.