Systems and methods for unsupervised abstraction of sensitive data shared for federations
By generating standard customer profiles through unsupervised learning and simulating synthetic transaction data using a transaction data simulator, the problem of financial institutions lacking real customer data is solved, thus improving the training effect of financial crime detection models.
Patent Information
- Application Number
- CN202080076408.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-05
- Filing Date
- 2020-11-02
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2040-11-02
AI Technical Summary
Financial institutions lack sufficient real customer data to train financial crime detection models, resulting in poor simulation performance of the predictive models.
Standard customer profiles are generated using unsupervised learning methods, and a large amount of synthetic transaction data is simulated using a transaction data simulator to train a predictive model by simulating customer transaction behavior.
It provides a large amount of seemingly authentic but non-sensitive customer transaction data to train financial crime detection models, thereby improving the models' predictive accuracy and robustness.
Smart Images

Figure CN114730359B_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to a cognitive system for implementing a transaction data simulator, and more specifically to a system and method for unsupervised abstraction of sensitive data for information alliance sharing. Background Technology
[0002] Financial crime detection systems, such as Financial Crime Warning Insights, in conjunction with IBM Cognitive analytics can be used to help banks detect money laundering and terrorist financing. It distinguishes between "normal" financial activities and "suspicious" activities, using this discriminative information to build predictive models for banks. However, training these predictive models requires a large amount of real-world financial customer data.
[0003] Because real customer data is highly sensitive, banks can only provide a limited amount of real customer data. However, to best simulate fraud scenarios and detect different types of financial crimes, more realistic simulated customer data, such as transaction data used for training, can produce better predictive models. IBM and IBM Watson are trademarks of International Business Machines Corporation and are registered in many jurisdictions worldwide. Therefore, there is a need in this field to address the aforementioned issues. Summary of the Invention
[0004] From a first aspect, the present invention provides a computer-implemented method for generating standard customer profiles in a data processing system, the data processing system including a processing device and a memory, the memory including instructions executable by the processing device, the method comprising: receiving customer data from a plurality of computing devices via a network, the customer data including information from a plurality of customers to a plurality of entities; performing unsupervised learning on the customer data by the processing device to generate a plurality of customer clusters having a plurality of common features; determining, by the processing device, that a cluster represents a standard customer, and storing a plurality of standard customer profiles based on the determined standard customers, wherein the standard customer profiles include a plurality of data distributions for the plurality of common features; and providing the plurality of standard customer profiles to each of the plurality of computing devices for generating synthetic transaction data based on the standard customers.
[0005] In another aspect, the present invention provides an abstract system including a processing device and a memory, the memory including instructions executed by the processing device for generating standard customer profiles in a data processing system configured to: receive customer data from a plurality of computing devices via a network, the customer data including information from a plurality of customers to a plurality of entities; perform unsupervised learning on the customer data by the processing device to generate a plurality of customer clusters having a plurality of common features; determine by the processing device that a cluster represents a standard customer, and store a plurality of standard customer profiles based on the determined standard customers, wherein the standard customer profiles include a plurality of data distributions for the plurality of common features; and provide the plurality of standard customer profiles to each of the plurality of computing devices for generating synthetic transaction data based on the standard customers.
[0006] In another aspect, the present invention provides a computer program product for generating standard customer profiles in a data processing system, the computer program product including a computer-readable storage medium that can be read by processing circuitry and stores instructions that are executed by the processing circuitry to perform a method for performing the steps of the present invention.
[0007] In another respect, the present invention provides a computer program that is stored on a computer-readable medium and can be loaded into the internal memory of a digital computer, the computer program including a software code portion that, when the program is run on a computer, performs the steps of the present invention.
[0008] In another aspect, the present invention provides a computer program product including software that, when executed by a processor, performs a method comprising: receiving customer data from a plurality of computing devices via a network, the customer data including information from a plurality of customers to a plurality of entities; performing unsupervised learning on the customer data by the processing device to generate a plurality of customer clusters having a plurality of common features; determining, by the processing device, that the clusters represent standard customers, and storing a plurality of standard customer profiles based on the determined standard customers, wherein the standard customer profiles include a plurality of data distributions for the plurality of common features; and providing the plurality of standard customer profiles to each of the plurality of computing devices for generating synthetic transaction data based on the standard customers.
[0009] According to some embodiments, this disclosure describes a computer-implemented method for generating standard customer profiles in a data processing system. The method includes steps performed by a processing device, including: receiving customer data via a network from a plurality of computing devices, the customer data including information from multiple customers to multiple entities; performing unsupervised learning on the customer data to generate multiple customer clusters having multiple common characteristics; determining that each cluster represents a standard customer; and storing multiple standard customer profiles based on the determined standard customers. The standard customer profiles include multiple data distributions of multiple common characteristics. The method also includes providing the multiple standard customer profiles to each of the plurality of computing devices for generating synthetic transaction data based on the standard customers.
[0010] According to other embodiments, this disclosure describes an abstract system for generating standard customer profiles in a data processing system. The abstract system may include processing devices and memory. The abstract system can receive customer data from multiple computing devices via a network, the customer data including information from multiple customers to multiple entities. The abstract system can also perform unsupervised learning on the customer data to generate multiple customer clusters with multiple common characteristics, determine that each cluster represents a standard customer, and store multiple standard customer profiles based on the determined standard customers, wherein the standard customer profiles include multiple data distributions with multiple common characteristics. The abstract system can also provide the multiple standard customer profiles to each of the multiple computing devices for generating synthetic transaction data based on the standard customers.
[0011] According to another embodiment, this disclosure describes a non-transitory computer-readable medium having instructions stored thereon for generating standard client profiles in a data processing system, the instructions performing the disclosed method consistent with the disclosed embodiments when executed by at least one processing device.
[0012] Additional features and advantages of this disclosure will become apparent from the following detailed description of illustrative embodiments with reference to the accompanying drawings. Attached Figure Description
[0013] The foregoing and other aspects of the invention will be best understood from the following detailed description when read in conjunction with the accompanying drawings. For the purpose of illustrating the invention, presently preferred embodiments are shown in the drawings; however, it should be understood that the invention is not limited to the specific means disclosed. The drawings include the following figures:
[0014] Figure 1 A block diagram illustrating an embodiment of a cognitive system that implements a transaction data simulator in a computer network, according to the disclosed embodiments, is shown.
[0015] Figure 2A block diagram of an example data processing system, in which aspects of the illustrative embodiments may be implemented, is depicted according to the disclosed embodiments;
[0016] Figure 3 A schematic diagram illustrating an embodiment of an abstract system according to the disclosed embodiments is provided.
[0017] Figure 4 An exemplary process from customer data to a standard customer is described according to the disclosed embodiments;
[0018] Figure 5 A flowchart illustrating an illustrative embodiment of a method for abstracting data to generate a standard client according to a disclosed embodiment is provided.
[0019] Figure 6 A schematic diagram of an example standard customer generated by an abstract system according to a disclosed embodiment is depicted;
[0020] Figure 7 A schematic diagram illustrating an embodiment of a transaction data simulator according to the disclosed embodiments is shown.
[0021] Figure 8 A flowchart illustrating an illustrative embodiment of a method for simulating transaction data according to disclosed embodiments is provided; and
[0022] Figure 9 A schematic diagram illustrating multiple synthetic transaction data entries according to a disclosed embodiment is depicted. Detailed Implementation
[0023] In summary, cognitive systems are dedicated computer systems, or groups of computer systems, configured with hardware and / or software logic (combined with hardware logic that executes software thereon) to mimic human cognitive functions. These cognitive systems apply human-like characteristics to communicating and manipulating thoughts, and when combined with the inherent strengths of digital computing, they can solve problems with high accuracy and large-scale resilience. IBM This is an example of a cognitive system that can process human-readable language and, with human-like accuracy, identify inferences between text segments at a much faster speed and on a much larger scale. Typically, such a cognitive system is capable of performing the following functions;
[0024] Navigating the complexities of human language and understanding
[0025] Ingesting and processing large amounts of structured and unstructured data
[0026] Generating and evaluating hypotheses
[0027] Weighted and evaluated responses based solely on relevant evidence.
[0028] Provide situation-specific advice, insights, and guidance
[0029] The machine learning process improves knowledge and learning through each iteration and interaction.
[0030] This enables decisions to be made at the point of influence (contextual guidance).
[0031] Scaling by task size
[0032] Expanding and amplifying human expertise and knowledge
[0033] Identifying human-like attributes and traits that evoke resonance from natural language; inferring specific or unknowable attributes of various languages from natural language.
[0034] Highly correlated recall (memory and recall) based on data points (images, text, speech)
[0035] Based on experience, it imitates human cognition and utilizes situational awareness for prediction and sensing.
[0036] Answering questions based on natural language and specific evidence
[0037] In one aspect, cognitive systems can be augmented with transaction data simulators to simulate sets of customer transaction data from financial institutions (such as banks). Even if the simulated customer transaction data is not from the financial institution's "actual" customer transaction data, it can be used to train predictive models for identifying financial crimes.
[0038] The transaction data simulator combines multi-level unsupervised clustering methods with interactive reinforcement learning (IRL) models to create a large number of intelligent agents that have learned to behave like "standard customers".
[0039] In one embodiment, a multi-level unsupervised clustering method uses information including hundreds of attributes of a “standard customer” over varying time periods to create a large number of standard customer transaction behaviors (extracted from real customer transaction data provided by the bank). Each standard customer transaction behavior can be associated with a group of customers with similar transaction characteristics. An intelligent agent generates a human customer profile and selects one of the standard customer transactions to combine with the generated human customer profile. In this way, the intelligent agent can simulate a “standard customer” and learn to behave like a “standard customer.” The intelligent agent is then given a period of time (e.g., ten years) during which it can observe the environment, such as the past behavior of the represented “standard customer,” and learn to perform “fake” customer transactions similar to the standard customer’s transaction behaviors. Each factor of the standard customer transaction behavior can be statistical data. For example, the transaction amount of a standard customer transaction behavior can be a range of values, such as a transaction amount of 20-3000 yuan. The transaction location of the standard customer transaction behavior can be provided statistically, such as 30% of the transaction locations being shopping malls, 50% being restaurants, and 20% being gas stations. The types of transactions in standard customer transactions can be provided statistically, for example, 20% of transactions are checks, 40% are POS payments, 25% are ATM withdrawals, and 15% are wire transfers. The medium of transaction in standard customer transactions can also be provided statistically, for example, 15% are cash, 45% are credit cards, 25% are checking accounts, and 15% are other forms of payment.
[0040] In one embodiment, a large number of artificial customer profiles are generated based on multiple real customer profile data. The real customer profile data can be provided by one or more banks. Each real customer profile may include the customer's address; the customer's name (the customer can be a legal entity or an individual); contact information such as phone number, email address, etc.; credit information, such as credit score, credit report, etc.; income information (e.g., annual income of a legal entity, or salary of an individual), etc. The real customer profile data is stored under different categories. For example, business customers (i.e., legal entities) can be categorized based on the size of the business customer, products, or services. Artificial customer profiles can be generated by randomly searching all real customer profile data. For example, artificial customer profiles can be generated by combining randomly selected information, including address, first name, last name, phone number, email address, credit score, income, or salary, etc. Therefore, the generated artificial customer profiles extract different pieces of information from the real customer profile data and thus resemble real customer profiles. (Financial transaction artificial customer profiles.)
[0041] In one embodiment, to protect the privacy of real customers, composite information such as address and name can be divided into multiple parts before random selection. For example, the address "2471 George Wallace Street" can be parsed into three parts: [number] "2471", [name] "George Wallace", and [suffix] "street". These parts can be randomly selected individually to form a human customer profile. In another embodiment, the composite information of the human customer profile, such as address and name, is compared with the composite information of the real customer profile. If the similarity is greater than a predetermined threshold, the human customer profile is unacceptable and needs to be updated until the similarity is less than the predetermined threshold.
[0042] Figure 1 A schematic diagram depicts an illustrative embodiment of a cognitive system 100 implementing a transaction data simulator 110 and an abstract system 120 within a computer network 114. The cognitive system 100 is implemented on one or more computing devices 112 (including one or more processing devices and one or more memories, and potentially any other computing device elements generally known in the art, including buses, storage devices, communication interfaces, etc.) connected to the computer network 114. The computer network 114 includes a plurality of computing devices 112 that communicate with each other and with other devices or components via one or more wired and / or wireless data communication links, wherein each communication link includes one or more of wires, routers, switches, transmitters, receivers, etc. Other embodiments of the cognitive system 100 can be used with components, systems, subsystems, and / or devices other than those described herein. In various embodiments, the computer network 114 includes local network connectivity and remote connectivity, enabling the cognitive system 100 to operate in environments of any size, including local and global, such as the Internet. Cognitive system 100 is configured to implement a transaction data simulator 110 capable of simulating standard customer transaction data 106 (i.e., standard customer transaction behavior). Transaction data simulator 110 can generate a large amount of simulated customer transaction data 108 based on the standard customer transaction data 106, making the simulated customer transaction data 108 appear to resemble real customer transaction data. In an embodiment, the standard customer transaction data 106 is obtained through an unsupervised clustering method. The original customer data, comprising a large amount of customer transaction data, is provided by one or more banks, and is clustered or grouped from the original customer data through an unsupervised clustering method into groups representing different characteristics of the bank's customers. Each group includes transaction data from customers with similar characteristics. For example, group A represents a single lawyer practicing patent law in New York, while group B represents a married lawyer practicing commercial law in New York.
[0043] The abstraction system 120 is implemented in hardware and / or software and is configured to perform unsupervised abstraction of standard customer transaction data 106 to generate one or more standard customers that are abstract representations of real customers but do not contain traceable customer information that may expose sensitive information. In an exemplary embodiment, the abstraction system 120 is configured to perform repeated unsupervised learning steps to cluster and sub-cluster the real customer data to generate standard customers representing a small group of customers.
[0044] Figure 2 This is a block diagram of an example data processing system 200 in which various aspects of the illustrative embodiments are implemented. The data processing system 200 is an example of a computer, in which computer-usable code or instructions for implementing the processes of the illustrative embodiments of the invention are located. In one embodiment, Figure 2 This refers to a transaction data simulator 110, which implements at least some aspects of the cognitive system 100 described herein.
[0045] In the described example, the data processing system 200 may employ a hub architecture including a northbridge and memory controller hub (NB / MCH) 201 and a southbridge and input / output (I / O) controller hub (SB / ICH) 202. The processing unit 203, main memory 204, and graphics processor 205 may be connected to the NB / MCH 201. The graphics processor 205 may be connected to the NB / MCH 201 via an Accelerated Graphics Port (AGP).
[0046] In the described example, network adapter 206 is connected to SB / ICH 202. Audio adapter 207, keyboard and mouse adapter 208, modem 209, read-only memory (ROM) 210, hard disk drive (HDD) 211, optical disc drive (CD or DVD) 212, Universal Serial Bus (USB) port and other communication ports 213, and PCI / PCIe device 214 can be connected to SB / ICH 202 via bus system 216. PCI / PCIe device 214 may include Ethernet adapters, add-in cards, and PC cards for notebook computers. ROM 210 may be, for example, a Flash Basic Input / Output System (BIOS). HDD 211 and optical disc drive 212 may use Integrated Drive Electronics (IDE) or Serial Advanced Technology Attachment (SATA) interfaces. Super I / O (SIO) device 215 may be connected to SB / ICH 202.
[0047] An operating system can run on the processing unit 203. The operating system can coordinate and provide control over various components within the data processing system 200. As a client, the operating system can be a commercially available operating system, such as an object-oriented programming system like Java.TM The programming system can run alongside the operating system and provide calls to the operating system from object-oriented programs or applications executed on the data processing system 200. As a server, the data processing system 200 can run a high-level interactive operating system. eServer TM System or Operating system. Registered trademark. This is used under license from the Linux Foundation, the proprietary licensee of Linus Torvalds, the trademark owner worldwide. eServer is a trademark of International Business Machines Corporation, registered in many jurisdictions worldwide. Data processing system 200 may be a symmetric multiprocessor (SMP) system, which may include multiple processors in processing unit 203. Alternatively, it may be a single-processor system.
[0048] Instructions for operating systems, object-oriented programming systems, and applications or programs reside on storage devices such as HDD 211 and are loaded into main memory 204 for execution by processing unit 203. The processes of embodiments of the website navigation system can be executed by processing unit 203 using computer-usable program code, which may reside in memory such as main memory 204, ROM 210, or in one or more peripheral devices.
[0049] Bus system 216 may include one or more buses. Bus system 216 may be implemented using any type of communication architecture or structure that can provide data transmission between different components or devices attached to that architecture or structure. Communication units such as modem 209 or network adapter 206 may include one or more devices that can be used to send and receive data.
[0050] Those skilled in the art will understand that Figure 2 The hardware described herein may vary depending on the implementation. For example, data processing system 200 includes several components that will not be directly included in some embodiments of abstract system 120. However, it should be understood that transaction data simulator 110 may include one or more of the components and configurations of data processing system 200 for performing the processing methods and steps according to the disclosed embodiments.
[0051] In addition to, or in place of, the described hardware, other internal hardware or peripheral devices may be used, such as flash memory, equivalent non-volatile memory, or optical disc drives. Furthermore, the data processing system 200 may take the form of any of a variety of different data processing systems, including but not limited to client computing devices, server computing devices, tablet computers, laptop computers, telephones or other communication devices, personal digital assistants, etc. Essentially, the data processing system 200 can be any known or subsequently developed data processing system without architectural limitations.
[0052] Figure 3 This is a schematic diagram of an illustrative embodiment of the abstract system 120. In some embodiments, the abstract system 120 may include multiple modules stored in main memory 204. The multiple modules may be implemented in hardware and / or software. The abstract system 120 may include a data collection module 310, an unsupervised learning module 320, a standard client module 330, and a boundary module 340. In some embodiments, the abstract system 120 may also include and / or be connected to one or more data repositories 250.
[0053] Data collection module 310 can be configured to receive customer data from computing device 112. Customer data can be actual customer data. For example, customer data 106 may originate from a financial institution and include information such as identification information, transaction information, etc. Customer data 106 may include various features stored separately as individual information categories. For example, customer data 106 may include spending data, payment data, time period data, location data, etc. In some embodiments, data collection module 310 can be configured to collect data from multiple computing devices 112, such as from multiple financial institutions. In some embodiments, data collection module 310 can be configured to perform a filtering process to create data sets for analysis. For example, data collection module 310 can use manual or automatic customer classification to create pools of similar customers (e.g., individuals, companies, retail, services, etc.).
[0054] The unsupervised learning module 320 can be configured to perform unsupervised learning on a dataset. Unsupervised learning can be, for example, a clustering algorithm configured to group one or more subsets of data based on patterns, trends, and / or other similarities found in the data. The unsupervised learning module 320 can be configured to perform the clustering process without requiring manual input to the groups (hence the term "unsupervised" learning). As a result, clusters can be free from the bias of how much a user might believe the data should be grouped.
[0055] The standard customer module 330 can be configured to extract clusters or groups from the output of the unsupervised learning module to generate and store standard customer profiles based on input data from the data collection module. The standard customer module 330 can also be configured to perform general sanity checks on the clusters (e.g., sample size, statistical significance, etc.) to determine when a cluster or sub-cluster can be considered a standard customer.
[0056] Boundary module 340 can be configured to further segment the collected customer data based on one or more boundaries. For example, boundary module 340 can be a statistical and / or time-slicing module configured to further filter data based on one or more parameters, enabling analysis of individual customers and / or standard customers from different viewpoints. For example, boundary module 340 can create subcategories of data based on two or more features (e.g., transaction information and time information). For example, customer data collected by data collection module 310 can provide transaction information for a customer over a year. Boundary module 340 can set time-segment boundaries for the data on a yearly basis to identify additional features that can be considered data points. For example, boundary module 340 can create categories such as "holiday spending," "vacation spending," "lunchtime spending," "savings period," etc. Therefore, boundary module 340 can be used to further segment and classify customer data. In some embodiments, boundary module 340 can apply these principles to standard customers. For example, boundary module 340 can derive additional standard customer behaviors from established customer behaviors by exploring data according to certain time periods or based on other statistical boundaries.
[0057] Figure 4This is a flowchart illustrating the process of using unsupervised learning on customer data 106 to generate one or more standard customer profiles through data abstraction. As a result of data abstraction, customer data 106 is abstracted / aggregated to a level that can be stored locally without privacy concerns. In some embodiments, data collection module 310 may receive customer data 106 from one or more computing devices 112. Data collection module 310 may perform initial data filtering 405. For example, data collection module 310 may perform RFM (Recency, Frequency, Currency Value) analysis on subgroups of data from customer data 106. Unsupervised learning module 320 may perform clustering process 410 to create one or more data clusters 415. One or more data clusters 415 may be customer groupings based on the unsupervised learning algorithm applied as the clustering process 410. Clusters 415 may be based on the similarity of one or more features in the customer data. For example, "Cluster 1" of cluster 415 may be all customers in a specific geographic area, while "Cluster 2" of cluster 415 may be all customers of a specific age, who spend a certain amount annually, or whose annual savings are less than a certain amount, etc. Unsupervised learning 410 can generate any number of clusters 415, and clients can be in more than one cluster.
[0058] Unsupervised learning module 320 can perform additional clustering process 420 to create one or more sub-clusters 425. This unsupervised learning module 320 can generate sub-clusters 425 by further grouping customers based on additional similarity in the data. For example, for customers in an initial location-based cluster 415, sub-clusters could be based on age, occupation, spending, transaction details, etc. The unsupervised learning 420 used to generate sub-clusters 425 can be repeated any number of times until the standard customer module 330 identifies clusters that are considered sub-clusters of standard customers 430. For example, the standard customer module 330 can select clusters that meet specific criteria, such as the number of customers in the group and / or similar characteristics. The customer module 330 can store these customers as profiles of standard customers 430 to be used as “abstract” customers that can be used to reproduce real customer data. For example, standard customers 430 can be provided to cognitive system 100 for use with transaction data simulator 110.
[0059] Figure 5This is an exemplary process 500 for transforming customer data into abstract standard customers to generate synthetic transaction data that is real but cannot be traced back to actual data. In step 510, the data collection module 310 receives and filters customer data. In step 520, the unsupervised learning module 320 applies an algorithm to the data to generate customer clusters based on the similarity of customers in at least one feature. In step 530, the unsupervised learning module performs unsupervised learning on the clusters to generate sub-clusters of customers and customer features. The clustering process can be repeated as needed to generate smaller and more specific customer clusters. In at least some embodiments, each unsupervised learning step adds data features to the customer groupings.
[0060] In step 540, the standard customer module 330 determines standard customers based on data clusters and sub-clusters through unsupervised learning. The standard customer module 330 can use a rule database to determine when a cluster is considered a standard customer. For example, the standard customer module 330 can compare the number of data features and the number of customers in a group with a threshold to determine whether the group has sufficient and / or sufficiently narrow data to be considered a standard customer.
[0061] In step 550, boundary module 340 may further derive additional standard customers. For example, in some embodiments, boundary module 340 may add customers to a standard customer profile based on a portion of customer data that fits the customer profile. For example, boundary module 340 may perform boundary operations on the customer data to identify customers that fit the standard customer profile when certain boundaries are applied. For example, boundary module 340 may select a cluster or a standard customer profile and perform additional analysis to see how customer behavior evolves when the time element is taken into account. In other examples, boundary module 340 may apply statistical boundaries to derive additional standard customers.
[0062] In step 560, the abstract system 120 may provide standard customers to the cognitive system 100. The cognitive system can use the standard customers as input to create new synthetic transaction data 108 that conforms to standard customer behavior but is not traceable to the original real customer data. As a result, real customer data 106 is used to create artificial customer data 108, which can be trusted as real but does not expose actual sensitive customer data.
[0063] Figure 6These are representations of standard customers 610 and 620, which can be generated based on customer data 106 through one or more disclosed processes. In an exemplary embodiment, standard customers 610 and 620 include multiple features that describe customers existing in the groups that make up the standard customers 610 and 620. For example, feature 1 may include customer age, feature 2 may include customer income, feature 3 may include customer spending, and so on. At least some of the features constituting standard customers 610 and 620 can be represented as a distribution of data. For example, the distribution may be a distribution of data having data points for each customer in the standard customer profile. Thus, the distribution is a representation of the actual customer data, but it is an abstract, general statistical representation that does not expose the actual data.
[0064] Figure 7A schematic diagram illustrating an illustrative embodiment of a transaction data simulator 110 is shown. The transaction data simulator 110 utilizes reinforcement learning techniques to simulate financial transaction data. The transaction data simulator 110 includes an intelligent agent 702 and an environment 704. The intelligent agent 702 randomly selects standard transaction behaviors 720 (i.e., targets 720) representing a group of "customers" with similar transaction characteristics and associates these standard transaction behaviors with randomly selected human customer profiles 718. The intelligent agent 702 takes an action 712 in each iteration. In this embodiment, the action 712 taken in each iteration includes conducting multiple transactions throughout the day. Each transaction has information including the transaction type (e.g., Automated Clearing House (ACH) transfer, check payment, wire transfer, ATM withdrawal, point-of-sale (POS) payment, etc.); the transaction amount; the transaction time; the transaction location; the transaction medium (e.g., cash, credit card, debit card, checking account, etc.); and a second party associated with the transaction (e.g., the person receiving the wire transfer payment), etc. Environment 704 takes action 712 as input and returns a reward 714 (or feedback) and a state 716 as output. The reward 714 is feedback that measures the success or failure of action 712. In this embodiment, environment 704 compares action 712 to a target 720 (e.g., standard trading behavior). If action 712 deviates from target 720 by more than a predefined threshold, the intelligent agent 702 is penalized; if action 712 deviates from target 720 within the predefined threshold (i.e., action 712 is similar to target 720), the intelligent agent 702 is rewarded. Action 712 is effectively evaluated so that the intelligent agent 702 can improve its next action 712 based on the reward 714. In this embodiment, environment 704 is the set of all old actions taken by the intelligent agent 702; that is, environment 704 is the set of all old simulated transactions. The intelligent agent 702 observes the environment 704 and obtains information about past transactions, such as the number of transactions made in a day, week, month, or year; the amount of each transaction, account balance, and transaction type. The strategy engine 706 can adjust the strategy based on the observations, enabling the intelligent agent 702 to take better actions in the next iteration 712.
[0065] The intelligent agent 702 also includes a policy engine 706, which is configured to adjust the policy based on state 716 and reward 714. The policy is the strategy used by the intelligent agent 702 to determine the next action 712 based on state 716 and reward 714. The policy is adjusted to obtain a higher reward 714 for the next action 712 taken by the intelligent agent 702. This policy includes a set of different policy probabilities or decision probabilities, which can be used to determine whether to execute a transaction on a specific day, the number of transactions per day, the transaction amount, the transaction type, the trading party, etc. In reinforcement learning models, the outcome of events is random, and a random number generator (RNG) is a system that generates random numbers from a real source of randomness. In one example, the maximum number of transactions per day is 100, and the maximum transaction amount is 15 million yuan. In the first iteration, the intelligent agent 702 performs a random transaction of 15 million yuan to Zimbabwe. Action 712 deviates from target 720 (e.g., a transaction in Maine made by a married lawyer practicing commercial law), and therefore action 712 is penalized (i.e., reward 714 is negative). The strategy engine 706 is trained to adjust the strategy, enabling different transactions to be made that are closer to target 720. Through further iterations, a "smarter" strategy engine 706 can simulate transactions similar to target 720. Figure 8 As shown, multiple transactions from client "James Culley" were simulated, and the simulated transaction data is similar to target 720.
[0066] like Figure 2 As shown, in one embodiment, a feedback loop (i.e., one iteration) corresponds to one “day” of actions (i.e., one “day” of simulated trading). Over a period of time, such as ten years, the intelligent agent 702 learns how to take actions 712 to obtain the highest possible reward 714. The number of iterations corresponds to the duration. For example, ten years corresponds to 10 × 365 = 3650 iterations. The reinforcement learning model judges actions 712 based on the results produced by actions 712. It is objective 720, and its purpose is to learn a sequence of actions 712 that will guide the intelligent agent 702 to achieve its objective 720 or maximize its objective function.
[0067] In one embodiment, the transaction data simulator 110 also includes an updater 710. A new action 712 is executed in each iteration. After each iteration, the updater 710 updates the environment 704 using the action 712 taken by the smart agent 702. The action 712 taken in each iteration is added to the environment 704 by the updater 710. In another embodiment, the transaction data simulator 110 also includes a pruner 708 configured to prune the environment 704. In this embodiment, the pruner 708 may remove one or more unwanted actions. For example, actions 712 taken in the first ten iterations may be removed because these ten iterations deviate significantly from the target 720 and have a similarity below a predefined threshold. In another embodiment, a complete reinitialization of the transaction data simulator 110 may be performed to remove all accumulated actions in the environment 704, allowing the smart agent 702 to restart.
[0068] Figure 8 A flowchart illustrating an illustrative embodiment of a method 800 for simulating transaction data is shown. In step 802, standard customer transaction behavior data is provided as target 720. Standard customer transaction behavior represents a group of customers with similar transaction characteristics. Standard customer transaction behavior is obtained through an unsupervised clustering method.
[0069] In step 804, action 712 is taken to execute multiple transactions in an iteration representing, for example, a day (e.g., 100 transactions per day). Each transaction has information including transaction type, transaction amount, transaction time, transaction location, transaction medium, and a second party associated with the transaction.
[0070] In step 806, environment 704 compares target 720 with action 712 taken in the iteration, and rewards or punishes action 712 based on similarity or deviation from target 720. The threshold or rule used to determine whether action 712 is similar to target 720 is predefined and can be adjusted based on user preferences and the degree of similarity to target 720.
[0071] In step 808, the environment 704 is updated to include action 712 in the current iteration. Environment 704 includes the set of all previous actions.
[0072] In step 810, the strategy engine 706 adjusts the strategy used to determine the next action 712 based on the reward 714 (i.e., reward or penalty). This strategy is formulated based on various factors, such as the probability of a transaction occurring, the number of transactions per day, the transaction amount, the transaction type, the trading party, the transaction frequency for each transaction type, the upper and lower limits for each transaction, the trading medium, etc. The strategy can adjust the weights of these factors based on the reward 714 in each iteration.
[0073] In step 812, in the new iteration, the smart agent 702 takes a new action 712. Steps 804 to 812 are repeated until action 712 is sufficiently similar to target 720 (step 814). For example, the transaction amount specified in target 720 is 20-3000 yuan. If the transaction amount of each transaction in action 712 falls within the range of 20-3000 yuan, then action 712 is sufficiently similar to target 720.
[0074] Since standard customer transaction data 106 can include anomalous data, such as fraudulent transactions, simulated customer transaction data 108 can also include anomalous data, as simulated customer transaction data 108 is similar to standard customer transaction data 106. In the reinforcement learning model, intelligent agent 702 explores environment 704 randomly or stochastically, learns a policy from its experience, and updates the policy during its exploration to improve the behavior (i.e., transactions) of intelligent agent 702. In embodiments, in contrast to random actions, behavioral patterns may emerge during RNG-based exploration (e.g., spending "splurges" until savings are depleted, or experiencing "buyer's regret" after a large purchase, etc.). Anomalous behavioral patterns may indicate fraudulent transactions. For example, simulated customer James Culley typically makes transactions under $1,000. Suddenly, a transaction of $5,000 occurs, and this suspicious transaction may be fraudulent (e.g., James Culley's credit card was stolen, or James Culley's checking account was hacked).
[0075] There are behavioral patterns that naturally emerge during exploration. For example, such as... Figure 9As shown, the simulated customer James Culley received $12,387.71 in his checking account on January 1, 2014. James Culley spent $474.98 on January 9, 2014, $4,400 on January 31, and $3,856.55 on March 2, 2014, using the debit card associated with the checking account. The following month, James Culley received $12,387.71 in his checking account on February 1, 2014. James Culley spent $4,500 on February 2, 2014, $1,713.91 on February 3, and transferred $8,100 out of his checking account on June 27, 2014. In this example, the simulated customer James Culley exhibits a tendency to save and spend money, and occasionally makes large purchases. Behavioral patterns make the simulated customer James Culley behave more realistically (i.e., he looks more like a real customer than a bot). The strategy engine 706 generates multiple parameters, such as "behavioral consistency" (the degree of consistency in behavior over a period of time), "consistency volatility" (the frequency of behavioral changes), and "behavioral anomalies" (deviations from regular trading behavior), and these parameters are used to indicate the different personalities of each simulated customer.
[0076] The transaction data simulator 110 uses abstracted or aggregated real customer data to simulate customer data representing real customers. The transaction data simulator 110 can provide a large amount of simulated customer data (i.e., simulated transaction data combined with human customer profiles), which can be used to train predictive models for detecting anomalous customer behavior. Furthermore, the simulated customer data is generated based on abstract data from real raw customer data rather than the real raw customer data itself, therefore it is impossible to derive any actual transaction actions of real customers.
[0077] In addition, the transaction data simulator 110 allows for the generation of behavioral patterns for each simulated customer during iterations.
[0078] The systems and processes shown in the accompanying drawings are not exclusive. Other systems, processes, and menus can be derived from the principles of the embodiments described herein to achieve the same purpose. It should be understood that the embodiments and variations shown and described herein are for illustrative purposes only. Modifications to the present design can be made by those skilled in the art without departing from the scope of the embodiments. As described herein, various systems, subsystems, agents, managers, and processes can be implemented using hardware components, software components, and / or combinations thereof.
[0079] This invention can be a system, method, and / or computer program product. A computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of the invention.
[0080] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes the following: portable computer disks, head disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices such as punch cards or raised structures in recesses on which instructions are recorded, and any suitable combination of the foregoing. As used herein, computer-readable storage media should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires.
[0081] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device, or downloaded via a network (e.g., the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network) to an external computer or external storage device. The network may include copper cables, optical fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.
[0082] Computer-readable program instructions for performing the operations of this invention may be assembly instructions, instruction set architecture (ISA) instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Java. TMSmalltalk, C++, and other common procedural programming languages, such as the "C" programming language or similar programming languages. Computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer via any type of network, including LANs or WANs, or can be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, to perform aspects of the invention, electronic circuits including, for example, programmable logic circuits, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs) can execute computer-readable program instructions to personalize the electronic circuits by utilizing state information of the computer-readable program instructions.
[0083] Various aspects of the present invention are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0084] These computer-readable program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0085] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0086] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of instructions comprising one or more executable instructions for implementing a specified logical function. In some alternative embodiments, the functions mentioned in the blocks may occur in a non-linear order as shown in the figures. For example, two blocks shown consecutively may actually be executed substantially simultaneously, or these blocks may sometimes be executed in reverse order, depending on the functions involved. It will also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0087] This specification and claims may use the terms "a," "at least one," and "one or more" with respect to specific features and elements of illustrative embodiments. It should be understood that these terms and phrases are intended to indicate the presence of at least one specific feature or element in a particular illustrative embodiment, but more than one may also be present. That is, these terms / phrases are not intended to limit the specification or claims to the presence of a single feature / element or to require the presence of multiple such features / elements. Rather, these terms / phrases only require at least a single feature / element, while the possibility of multiple such features / elements is within the scope of the specification and claims.
[0088] Furthermore, it should be understood that the following description uses various examples of various elements of the illustrative embodiments to further illustrate example implementations of the illustrative embodiments and to aid in understanding the mechanisms of the illustrative embodiments. These examples are intended to be non-limiting and are not an exhaustive list of all possibilities for implementing the mechanisms of the illustrative embodiments. In view of this specification, it will be apparent to those skilled in the art that many other alternative implementations of these various elements may be utilized, in addition to the examples provided herein, or as alternatives, without departing from the spirit and scope of the invention.
[0089] Although the invention has been described with reference to exemplary embodiments, the invention is not limited thereto. Those skilled in the art will understand that many changes and modifications can be made to the preferred embodiments of the invention, and such changes and modifications can be made without departing from the true spirit of the invention. Therefore, the appended claims are intended to be construed as covering all such equivalent changes falling within the true spirit and scope of the invention.
Claims
1. A computer-implemented method, wherein, The data processing system includes a processing device and a memory, the memory including instructions executed by the processing device, the method comprising: Customer data is received from multiple computing devices via a network, and the customer data includes information for multiple customers to multiple entities; The processing device performs unsupervised learning on the customer data to generate multiple customer clusters with multiple common characteristics; The processing device determines clusters representing standard clients and stores multiple standard client profiles based on the determined standard clients, wherein the standard client profiles include multiple data distributions for the multiple common characteristics; and The plurality of standard customer profiles are provided to each of the plurality of computing devices through the following steps to generate synthetic transaction data including actions based on the standard customers: Offering multiple standard customer transactions as a target, Perform multiple iterations until the similarity between the action and the target is higher than a predefined threshold, wherein in each iteration: Take action to execute multiple simulated trades. The action is compared with the target. Provide feedback related to the action based on the similarity to the target, and Based on feedback, adjust the strategy to determine the next action; and The synthetic transaction data is used to train a predictive model to identify unusual customer behavior.
2. The method according to claim 1, wherein, The information for these multiple customers includes identification information and transaction information.
3. The method according to any one of the preceding claims further includes: The customer data is filtered before performing unsupervised learning.
4. The method according to claim 3, wherein, The filtering includes RFM analysis to group customers.
5. The method according to claim 1, wherein, The unsupervised learning process includes clustering customers based on common features and repeating the unsupervised learning to form subclusters of customers based on the multiple common features.
6. The method according to claim 1, wherein, Determining the standard for cluster representation clients includes applying one or more rules.
7. The method according to claim 6, wherein, The one or more rules include size determination, which indicates the minimum or maximum number of customers in a sub-cluster that is determined to be a standard customer.
8. A data processing system comprising a processing device and a memory, the memory including instructions executable by the processing device for generating a standard client profile in the data processing system, the data processing system being configured to: Customer data is received from multiple computing devices via a network, and the customer data includes information for multiple customers to multiple entities; The processing device performs unsupervised learning on the customer data to generate multiple customer clusters with multiple common characteristics; The processing device determines a cluster to represent a standard customer and stores multiple standard customer profiles based on the determined standard customers, wherein the standard customer profiles include multiple data distributions for the multiple common characteristics; as well as The plurality of standard customer profiles are provided to each of the plurality of computing devices through the following steps to generate synthetic transaction data including actions based on the standard customers: Offering multiple standard customer transactions as a target, Perform multiple iterations until the similarity between the action and the target is higher than a predefined threshold, wherein in each iteration: Take action to execute multiple simulated trades. The action is compared with the target. Provide feedback related to the action based on the similarity to the target, and Adjust strategies based on feedback to determine the next action; as well as The synthetic transaction data is used to train a predictive model to identify unusual customer behavior.
9. The system according to claim 8, wherein, The information of the multiple customers includes identification information and transaction information.
10. The system according to any one of claims 8 or 9, further comprising: The customer data is filtered before performing unsupervised learning.
11. The system according to claim 10, wherein, The filtering includes RFM analysis to group customers.
12. The system according to claim 8, wherein, Performing unsupervised learning involves clustering customers based on common features, and repeating unsupervised learning to form subclusters of customers based on the multiple common features.
13. The system according to claim 8, wherein, Determining the standard for cluster representation clients includes applying one or more rules.
14. The system according to claim 13, wherein, The one or more rules include size determination, which indicates the minimum or maximum number of customers in a sub-cluster that is determined to be a standard customer.
15. A computer program product, the computer program product comprising: Program instructions, which can be read by the processing circuit and executed by the processing circuit to perform the method according to any one of claims 1 to 7.
16. A computer-readable storage medium comprising computer instructions that are loadable into the internal memory of a computer and, when executed on the computer, perform the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Automated transaction processing system and approach
CN101484924A
Financial loan big data risk assessment method and system based on privacy-removed data
CN108734021A