Simulated Transactional Data Clustering for Intelligent Agent Training
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for detecting suspicious financial activity rely on real customer data, which is difficult to access and limited in use due to its sensitive nature, and existing methods for addressing privacy concerns, such as data cleansing and anonymization, further limit the data's usefulness for training intelligent agents.
Innovation Solution
A system and method for generating simulated transactional data with sufficient variability to train intelligent agents by forming clusters of similar behavior, increasing variability through unsupervised learning and statistical analysis, and merging clusters to ensure meaningful representation of transactional behavior.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If real customer data is used for training intelligent agents, then the training effectiveness and realism are improved, but data privacy and security risks worsen
Solution Approach 1:
The patent creates simulated copies of real customer transactional data that preserve the statistical properties and behavioral patterns of actual data while removing all personally identifying information. These synthetic data copies enable intelligent agent training with the same realism and effectiveness as real data would provide, without exposing actual customer information to security risks.
Solution Approach 2:
The patent introduces simulated transactional data as an intermediary between the need for realistic training data and the requirement for data privacy protection. This intermediary layer provides the statistical fidelity needed for effective agent training while acting as a buffer that prevents direct access to sensitive real customer information.
2Object-affected harmful factors
If data is anonymized or hashed to protect privacy, then data security is improved, but data usefulness for analytics and training deteriorates
Solution Approach 1:
The patent transforms the parameters of transactional data by generating synthetic data that matches the statistical distribution and behavioral patterns of real data without containing actual customer information. This parameter transformation maintains the analytical value and training utility of the data while eliminating privacy risks associated with using real customer information.
3Measurement precision
If clusters are formed with high similarity for intelligent agent training, then behavioral pattern recognition is improved, but variability and representativeness worsen
Solution Approach 1:
The patent segments transactional data into multiple clusters based on behavioral patterns, then selectively combines clusters that exhibit complementary characteristics. This segmentation approach allows the system to maintain the behavioral precision needed for pattern recognition while incorporating diverse cluster characteristics that enhance overall data variability and representativeness for training robust intelligent agents.
Data Source
AI summary
A method, system, and computer programming product for checking that clusters representative of transactional activity of a group of persons exhibits sufficient variability including: receiving transactional data; forming clusters from the received transactional data representing groups of persons that behave similarly; determining that a cluster representing a group of persons that behave similarly is not sufficiently variable; and increasing, in response to the cluster representing the group of persons behaving similarly not being sufficiently variable, the variability of the cluster. Further including, in an embodiment, creating a superset cluster consisting of both the cluster and the parent of the cluster; creating test data using the superset as a baseline; injecting the test data into the superset cluster; determining if the superset cluster rejects the injected test data as an indication of insufficient variability.


