Unsupervised Customer Profile Abstraction for Synthetic Data Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Financial crime detection systems require large amounts of realistic simulated transaction data to build effective predictive models, but real customer data is sensitive and limited, making it challenging to generate sufficient training data.

Innovation Solution

A method using unsupervised learning and interactive reinforcement learning to create standard customer profiles, which are then used to generate synthetic transaction data that mimics real customer behavior without exposing sensitive information, by clustering real customer data and applying boundary limitations to produce abstracted customer profiles for use in cognitive systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If real customer data is used to train predictive models, then model accuracy is improved, but customer privacy is compromised

Engineering Contradiction:
Improvemodel accuracyVSAvoidcustomer privacy exposure
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent creates synthetic customer profiles that copy the statistical characteristics and behavioral patterns of real customers without containing actual sensitive data. These synthetic profiles are generated by clustering real data to extract common patterns, then synthesizing new profiles that preserve the underlying distributions while removing identifiable information, thus maintaining model training effectiveness while protecting privacy

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces synthetic customer profiles as an intermediary between real customer data and the predictive model training process. This intermediary layer allows the model to learn from realistic data patterns without direct exposure to sensitive real customer information, acting as a buffer that preserves both model accuracy and privacy

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If more simulated customer data is generated, then training data quantity is improved, but data realism may be compromised

Engineering Contradiction:
Improvetraining data quantityVSAvoiddata realism
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent applies different processing techniques to different aspects of customer data. Clustering is applied to identify common behavioral patterns across customer groups, while synthetic profile generation preserves local characteristics such as transaction amounts, frequencies, and patterns within each cluster. This ensures that generated data maintains realistic local qualities while increasing overall quantity

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent generates synthetic customer profiles by sampling from learned data distributions and applying parameter transformations. By changing parameters such as customer identifiers, account numbers, and specific transaction details while preserving the underlying statistical distributions and behavioral patterns, the system increases data quantity while maintaining realism

Inventive Principle:
Principle #35Parameter changes

3Object-affected harmful factors

If real customer data is abstracted to protect privacy, then privacy protection is improved, but data utility for training may be reduced

Engineering Contradiction:
Improveprivacy protectionVSAvoiddata utility
Core Design Contradiction:
Object-affected harmful factorsVSQuantity of substance

Solution Approach 1:

The patent segments customer data into clusters based on shared characteristics and behavioral patterns. By dividing the data into meaningful segments rather than removing information entirely, the system maintains the utility needed for training while protecting individual privacy through aggregation and pattern-based representation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates synthetic copies of customer profiles that preserve the statistical and behavioral characteristics of real data. These copies maintain data utility for training purposes while eliminating direct links to real customers, thus balancing privacy protection with data usefulness

Inventive Principle:
Principle #26Copying

Data Source

PatentUS11556734B2System and method for unsupervised abstraction of sensitive data for realistic modeling
Publication Date: 2023.01.17 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11556734B2 patent drawing
  • US11556734B2 patent drawing
  • US11556734B2 patent drawing

AI summary

An abstraction system for generating a standard customer profile in a data processing system has a processing device and a memory. The abstraction system may receive customer data from a computing device over a network, perform unsupervised learning on the customer data to produce a plurality of clusters of customers with a plurality of features in common, and determine that a cluster represents a standard customer, and store a plurality of standard customer profiles based on the determined standard customers, wherein the standard customer profiles comprise a plurality of data distributions for the plurality of features in common. The abstraction system also derives additional standard customer profiles by applying a boundary limiter to the customer data. The abstraction system additionally provides the standard customer profiles and the additional standard customer profiles to a cognitive system for generating synthetic transaction data.