Intelligent Agent for Simulating Customer Data
By using reinforcement learning models to generate simulated customer data in the financial crime detection system, the problem of limited real customer data is solved, and the effectiveness and accuracy of financial crime detection is improved.
Patent Information
- Application Number
- CN202080076407.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-11-05
- Filing Date
- 2020-11-02
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-11-02
AI Technical Summary
Financial crime detection systems require a large amount of real financial customer data to train predictive models, but due to the sensitivity of real customer data, banks can only provide a limited amount of data, resulting in poor effectiveness in simulating fraud situations and detecting financial crimes.
By using reinforcement learning models of intelligent agents, policy engines, and environments in the data processing system, artificial customer profiles and simulated customer transaction data are generated to make it similar to real customer data, and are used to train financial crime detection models.
It realizes the generation of realistic simulated customer data without exposing real customer data, which improves the training effect and prediction accuracy of the financial crime detection model.
Smart Images

Figure CN114616546B_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to cognitive systems implementing a transaction data simulator, and more particularly to a transaction data simulator configured to simulate transaction data provided by a financial institution, such as a bank. Background Art
[0002] Financial crime detection systems, such as Financial crime warning insights together with IBM can utilize cognitive analysis to help banks detect money laundering and terrorist financing. Cognitive analysis differentiates "normal" financial activities from "suspicious" activities and uses the differentiation information to build a predictive model for the bank. A large amount of real financial customer data is required to train the predictive model.
[0003] Since real customer data is highly sensitive, banks can only provide a limited amount of real customer data. However, to best simulate fraud scenarios and detect different types of financial crimes, more simulated customer data that appears realistic, such as transaction data for training, can produce a better predictive model. IBM and IBM Watson are trademarks of International Business Machines Corporation and are registered in many jurisdictions throughout the world. Accordingly, there is a need in the art to address the above issues. Summary of the Invention
[0004] From a first aspect, the present invention provides a computer-implemented method in a data processing system including a processor and a memory including instructions, the instructions being executed by the processor to cause the processor to implement a method for simulating customer data using a reinforcement learning model including an intelligent agent, a policy engine, and an environment, the method comprising: generating, by the processor, an artificial customer profile by combining information randomly selected from a set of real customer profile data; providing, by the processor, standard customer transaction data representing a set of real customers having transaction characteristics similar to a target; performing, by the intelligent agent, an action including a plurality of simulated transactions; comparing, by the environment, the action with the target; providing, by the environment, feedback associated with the action based on a similarity with respect to the target; adjusting, by the policy engine, a policy based on the feedback; repeating the steps of performing the action to the step of adjusting the policy until the similarity is higher than a first predetermined threshold; and combining, by the processor, the artificial customer profile with the last action to form simulated customer data.
[0005] From a first aspect, the present invention provides a computer-implemented method in a data processing system, the data processing system comprising a processor and a memory including instructions, the instructions being executed by the processor to cause the processor to implement a method for simulating customer data using a reinforcement learning model including an intelligent agent, a policy engine, and an environment, the method comprising: generating, by the processor, an artificial customer profile by combining information randomly selected from a set of real customer profile data; providing, by the processor, standard customer transaction data, the standard customer transaction data representing a set of real customers having transaction characteristics similar to a target; performing, by the intelligent agent, actions including a plurality of simulated transactions; comparing, by the environment, the actions with the target; providing, by the environment, feedback associated with the actions based on a similarity with respect to the target; adjusting, by the policy engine, a policy based on the feedback; repeating the steps of performing the actions to the step of adjusting the policy until the similarity is higher than a first predetermined threshold; and combining, by the processor, the artificial customer profile with the last action to form simulated customer data.
[0006] From a first aspect, the present invention provides a computer-implemented method in a data processing system, the data processing system comprising a processor and a memory including instructions, the instructions being executed by the processor to cause the processor to implement a method for simulating customer data using a reinforcement learning model including an intelligent agent, a policy engine, and an environment, the method comprising: generating, by the processor, an artificial customer profile by combining information randomly selected from a set of real customer profile data; providing, by the processor, standard customer transaction data, the standard customer transaction data representing a set of real customers having transaction characteristics similar to a target; performing, by the intelligent agent, actions including a plurality of simulated transactions; comparing, by the environment, the actions with the target; providing, by the environment, feedback associated with the actions based on a similarity with respect to the target; adjusting, by the policy engine, a policy based on the feedback; repeating the steps of performing the actions to the step of adjusting the policy until the similarity is higher than a first predetermined threshold; and combining, by the processor, the artificial customer profile with the last action to form simulated customer data.
[0007] From another aspect, the present invention provides a computer program product for simulating customer data using a reinforcement learning model including an intelligent agent, a policy engine, and an environment. The computer program product includes a computer-readable storage medium having program instructions embodied therewith, and the program instructions are executable by a processor to cause the processor to: generate an artificial customer profile by combining information randomly selected from a set of real customer profile data; provide standard customer transaction data, the standard customer transaction data representing a group of customers having transaction characteristics similar to a target; execute actions including a plurality of simulated transactions by the intelligent agent; compare the actions with the target by the environment; provide feedback associated with the actions by the environment based on the similarity to the target; adjust a policy by the policy engine based on the feedback; repeat the steps of executing the actions to the step of adjusting the policy until the similarity is higher than a first predetermined threshold; and combine the artificial customer profile with the last action to form simulated customer data.
[0008] From another aspect, the present invention provides a system for simulating customer data using a reinforcement learning model including an intelligent agent, a policy engine, and an environment. The system includes: a processor configured to: generate an artificial customer profile by combining information randomly selected from a set of real customer profile data; provide standard customer transaction data, the standard customer transaction data representing a group of customers having transaction characteristics similar to a target; execute actions including a plurality of simulated transactions by the intelligent agent; compare the actions with the target by the environment; provide feedback associated with the actions by the environment based on the similarity to the target; adjust a policy by the policy engine based on the feedback; repeat the steps of executing the actions to the step of adjusting the policy until the similarity is higher than a first predetermined threshold; and combine the artificial customer profile with the last action to form simulated customer data.
[0009] From another aspect, the present invention provides a computer program product for simulating transaction data. The computer program product includes a computer-readable storage medium readable by a processing circuit and storing instructions for executing a method for performing the steps of the present invention by the processing circuit.
[0010] From another aspect, the present invention provides a computer program stored on a computer-readable medium and loadable into an internal memory of a digital computer, including software code portions for performing the steps of the present invention when the program runs on the computer.
[0011] An embodiment provides a computer-implemented method in a data processing system including a processor and a memory including instructions that, when executed by the processor, cause the processor to implement a method for simulating customer data using a reinforcement learning model including an intelligent agent, a policy engine, and an environment. The method includes: generating, by the processor, an artificial customer profile by combining information randomly selected from a set of real customer profile data; providing, by the processor, standard customer transaction data representing a set of real customers having transaction characteristics similar to a target; performing, by the intelligent agent, an action including a plurality of simulated transactions; comparing, by the environment, the action with the target; providing, by the environment, feedback associated with the action based on a similarity to the target; adjusting, by the policy engine, a policy based on the feedback; repeatedly performing the action of the step of adjusting the policy until the similarity is higher than a first predetermined threshold; and combining, by the processor, the artificial customer profile with the last action to form simulated customer data.
[0012] The embodiment further provides a computer-implemented method, wherein the real customer profile data includes one or more of customer address, customer name, contact information, credit information, and income information.
[0013] The embodiment further provides a computer-implemented method, wherein each simulated transaction includes a transaction type, a transaction amount, a transaction time, a transaction location, a transaction medium, and a second party associated with the simulated transaction.
[0014] The embodiment further provides a computer-implemented method, wherein the environment includes a set of all previous actions performed by the intelligent agent.
[0015] The embodiment further provides a computer-implemented method, further including: removing, by the processor, a plurality of previous actions having a similarity lower than a second predefined threshold.
[0016] The embodiment further provides a computer-implemented method, further including: obtaining, by the processor, the standard customer transaction data from raw customer transaction data by an unsupervised clustering method.
[0017] The embodiment further provides a computer-implemented method, wherein the feedback is a reward or a penalty.
[0018] In another illustrative embodiment, a computer program product is provided that includes a computer-usable or readable medium having a computer-readable program. When the computer-readable program is executed on a processor, the computer-readable program causes the processor to perform various operations and combinations of operations outlined above in the illustrative embodiments of the method.
[0019] In yet another illustrative embodiment, a system is provided. The system can include a training data acquisition processor configured to perform various operations and combinations of operations as outlined above with respect to the illustrative embodiments of the method.
[0020] Additional features and advantages of the present disclosure will become apparent from the detailed description of the illustrative embodiments that follows, with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The foregoing and other aspects of the present invention are best understood from the following detailed description when read in conjunction with the accompanying drawings. In order to illustrate the present invention, there are shown in the drawings presently preferred embodiments, it being understood, however, that the invention is not limited to the specific means disclosed. The drawings include the following figures:
[0022] Figure 1 A schematic diagram depicting an illustrative embodiment of a cognitive system 100 implementing a transaction data simulator in a computer network;
[0023] Figure 2 A schematic diagram depicting an illustrative embodiment of a transaction data simulator 110;
[0024] Figure 3 A schematic diagram depicting a plurality of simulated transactions from simulated customers in accordance with embodiments herein;
[0025] Figure 4 A flowchart depicting an illustrative embodiment of a method 400 for simulating customer data; and
[0026] Figure 5 A block diagram of an example data processing system 500 in which aspects of the illustrative embodiments can be implemented. DETAILED DESCRIPTION
[0027] As an overview, a cognitive system is a special-purpose computer system, or a group of computer systems, configured with hardware and / or software logic (in combination with the hardware logic on which the software executes) to emulate human cognitive functions. These cognitive systems apply human-like characteristics to convey and manipulate ideas, which, when combined with the inherent strength of digital computing, can solve problems with high accuracy and large-scale resilience. IBM is an example of such a cognitive system that can process human-readable language and identify inferences between text segments with human-like precision at much faster speeds and much larger scales than humans. Generally, such cognitive systems are capable of performing the following functions:
[0028] Navigate the complexities of human language and understanding
[0029] Ingest and process large amounts of structured and unstructured data
[0030] Generate and evaluate hypotheses
[0031] Weighted sum and evaluate responses based only on relevant evidence
[0032] Provide situation-specific advice, insights, and guidance
[0033] Improve knowledge and learning with each iteration and interaction through a machine learning process
[0034] Enable decision-making at the point of impact (context-guided)
[0035] Scale by task
[0036] Extend and amplify human expertise and cognition
[0037] Identify human-like attributes and traits that resonate from natural language
[0038] Infer various language-specific or agnostic attributes from natural language
[0039] Highly relevant recall (memory and recollection) from data points (images, text, speech)
[0040] Predict and sense using situation awareness that emulates human cognition based on experience
[0041] Answer questions based on natural language and specific evidence
[0042] In one aspect, a cognitive system can be augmented with a transaction data simulator to simulate a collection of customer transaction data from a financial institution (e.g., a bank). Even though the simulated customer transaction data is not “actual” customer transaction data from a financial institution, it can be used to train a predictive model for identifying financial crimes.
[0043] The transaction data simulator combines a multi-layer unsupervised clustering method with an interactive reinforcement learning (IRL) model to create a large number of intelligent agents that have learned to behave like “standard customers”.
[0044] In one embodiment, the multi-layer unsupervised clustering method uses information including hundreds of attributes of "standard customers" over a varying time period to create a large number of standard customer transaction behaviors (extracted from real customer transaction data provided by a bank). Each standard customer transaction behavior can be associated with a group of customers with similar transaction characteristics. The intelligent agent generates an artificial customer profile and selects one of the standard customer transaction behaviors to combine with the generated artificial customer profile. In this way, the intelligent agent can simulate a "standard customer" and learn to behave like a "standard customer". Then, the intelligent agent is provided with a period of time (e.g., ten years) during which the intelligent agent can observe the environment, such as the past behavior of the represented "standard customer", and learn to perform "fake" customer transactions similar to the standard customer transaction behaviors of the represented "standard customer". Each factor of the standard customer transaction behavior can be statistical data. For example, the transaction amount of the standard customer transaction behavior can be a range of values, e.g., the transaction amount of the standard customer transaction behavior is 20 yuan - 3,000 yuan. The transaction location of the standard customer transaction behavior can be provided statistically, e.g., 30% of the transaction locations are shopping malls, 50% of the transaction locations are restaurants, and 20% of the transaction locations are gas stations. The transaction type of the standard customer transaction behavior can be provided statistically, e.g., 20% of the transaction types are check payments, 40% of the transaction types are POS payments, 25% of the transaction types are ATM withdrawals, and 15% of the transaction types are wire transfers. The transaction medium of the standard customer transaction behavior can be provided statistically, e.g., 15% of the transaction media are cash, 45% of the transaction media are credit cards, 25% of the transaction media are checking accounts, and 15% of the transaction media are
[0045] In one embodiment, a large number of artificial customer profiles are generated based on multiple real customer profile data. The real customer profile data can be provided by one or more banks. Each real customer profile may include the customer's address; the customer's name (the customer can be a legal entity or an individual); contact information such as a phone number, an email address, etc.; credit information, such as a credit score, a credit report, etc.; income information (e.g., the annual income of a legal entity, or an individual's salary), etc. The real customer profile data is stored under different categories. For example, business customers (i.e., legal entities) can be divided into different categories based on the size, products or services of the business customers. Artificial customer profiles can be generated by randomly searching through all the real customer profile data. For example, an artificial customer profile can be generated by combining randomly selected information, including an address, a first name, a last name, a phone number, an email address, a credit score, income or salary, etc. Thus, the generated artificial customer profiles extract different pieces of information from the real customer profile data and thus appear like real customer profiles. Financial transaction data is further simulated in association with each artificial customer profile. In one embodiment, the simulated customer transaction data can be combined with the artificial customer profile to form simulated customer data.
[0046] In one embodiment, to protect the privacy of real customers, composite information such as an address, a name, etc. can be divided into multiple parts before random selection. For example, the address "2471 George Wallace Street" can be parsed into 3 parts: [number] "2471", [name] "George Wallace", and [suffix] "Street". These parts can be randomly selected individually to form an artificial customer profile. In another embodiment, the synthetic information of the artificial customer profile, such as an address, a name, etc., is compared with the synthetic information of the real customer profile. If the similarity is greater than a predetermined threshold, the artificial customer profile is unacceptable and needs to be updated until the similarity is less than the predetermined threshold.
[0047] Figure 1FIG. 0 depicts a schematic diagram of an illustrative embodiment of a cognitive system 100 implementing a transaction data simulator 110 in a computer network 102. The cognitive system 100 is implemented on one or more computing devices 104 (including one or more processors and one or more memories, and potentially any other computing device elements commonly known in the art, including buses, storage devices, communication interfaces, etc.) connected to the computer network 102. The computer network 102 includes a plurality of computing devices 104 that communicate with each other and with other devices or components via one or more wired and / or wireless data communication links, where each communication link includes one or more of wires, routers, switches, transmitters, receivers, etc. Other embodiments of the cognitive system 100 may be used with components, systems, subsystems, and / or devices other than those described herein. In various embodiments, the computer network 102 includes local network connections and remote connections such that the cognitive system 100 can operate in environments of any size, including local and global, such as the Internet. The cognitive system 100 is configured to implement a transaction data simulator 110 that can simulate standard customer transaction data 106 (i.e., standard customer transaction behavior). The transaction data simulator 110 can generate a large amount of simulated customer transaction data 108 based on the standard customer transaction data 106 such that the simulated customer transaction data 108 appears like real customer transaction data. Then, the simulated customer transaction data 108 is combined with a randomly selected artificial customer profile 112 to obtain complete simulated customer data 114 for the simulated customer.
[0048] In an embodiment, the standard customer transaction data 106 is obtained by an unsupervised clustering method. Raw customer data including a large amount of customer transaction data is provided by one or more banks, and a large number of groups representing different characteristics of bank customers are clustered or grouped from the raw customer data by an unsupervised clustering method. Each group includes transaction data from customers with similar characteristics. For example, the customers represented by group A are individual lawyers practicing patent law in New York, while the customers represented by group B are married lawyers practicing commercial law in New York.
[0049] Figure 2FIG. 0 is a schematic diagram depicting an illustrative embodiment of a transaction data simulator 110. The transaction data simulator 110 utilizes reinforcement learning techniques to simulate financial transaction data. The transaction data simulator 110 includes an intelligent agent 202 and an environment 204. The intelligent agent 202 randomly selects a standard transaction behavior 220 (i.e., a target 220) that represents a group of "customers" with similar transaction characteristics, and associates the standard transaction behavior with a randomly selected artificial customer profile 112. The intelligent agent 202 takes an action 212 in each iteration. In this embodiment, the action 212 taken in each iteration includes performing multiple transactions in a day. Each transaction has information including a transaction type (e.g., automated clearing house (ACH) transfer, check payment, wire transfer, automated teller machine (ATM) withdrawal, point of sale (POS) payment, etc.); a transaction amount; a transaction time; a transaction location; a transaction medium (e.g., cash, credit card, debit card, checking account, etc.); a second party associated with the transaction (e.g., the person receiving a wire transfer payment), etc. The environment 204 takes the action 212 as an input and returns a reward 214 (or feedback) and a state 216 from the environment 204 as outputs. The reward 214 is such feedback by which the success or failure of the action 212 is measured. In this embodiment, the environment 204 compares the action 212 with the target 220 (e.g., the standard transaction behavior). If the action 212 deviates from the target 220 by more than a predefined threshold, the intelligent agent 202 is penalized, while if the action 212 deviates from the target 220 within the predefined threshold (i.e., the action 212 is similar to the target 220), the intelligent agent 202 is rewarded. The action 212 is effectively evaluated such that the intelligent agent 202 can improve the next action 212 based on the reward 214. In this embodiment, the environment 204 is a collection of all the old actions taken by the intelligent agent 202, i.e., the environment 204 is a collection of all the old simulated transactions. The intelligent agent 202 observes the environment 204 and obtains information about the old transactions, such as the number of transactions conducted in a day, a week, a month, or a year; each transaction amount, account balance, each transaction type, etc. A policy engine 206 can adjust the policy based on the observations such that the intelligent agent 202 can take a better action 212 in the next iteration.
[0050] The intelligent agent 202 also includes a policy engine 206, which is configured to adjust the policy based on the state 216 and the reward 214. The policy is the countermeasure that the intelligent agent 202 uses to determine the next action 212 based on the state 216 and the reward 214. The policy is adjusted to obtain a higher reward 214 for the next action 212 taken by the intelligent agent 202. The policy includes a set of different policy probabilities or decision probabilities, which can be used to decide whether to execute a transaction on a specific day, the number of transactions per day, the transaction amount, the transaction type, the counterparty, etc. In a reinforcement learning model, the outcome of an event is random, and a random number generator (RNG) is a system that generates random numbers from a true source of randomness. In one example, the maximum number of transactions per day is 100, and the maximum transaction amount is 15 million yuan. In the first iteration, the intelligent agent 202 makes a random transaction with a transaction amount of 1.5 million yuan to Zimbabwe. This action 212 deviates from the target 220 (e.g., a transaction conducted by a married lawyer practicing commercial law in Maine), and thus this action 212 is penalized (i.e., the reward 214 is negative). The policy engine 206 is trained to adjust the policy so that different transactions closer to the target 220 can be made. Through more iterations, transactions similar to the target 220 can be simulated by a "smarter" policy engine 206. As Figure 3 shown, multiple transactions from the customer "James Culley" are simulated, and the simulated transaction data is similar to the target 220.
[0051] As Figure 2 shown, in one embodiment, a feedback loop (i.e., one iteration) corresponds to the actions of one "day" (i.e., simulated transactions for one "day"). Over a period of time, such as ten years, the intelligent agent 202 learns how to take actions 212 to obtain the highest possible reward 214. The number of iterations corresponds to the duration. For example, ten years corresponds to 10 × 365 = 3650 iterations. Reinforcement learning judges the action 212 based on the result generated by the action 212. It is goal-oriented towards the target 220, and its purpose is to learn the sequence of actions 212 that will guide the intelligent agent 202 to achieve its target 220 or maximize its objective function.
[0052] In an embodiment, the transaction data simulator 110 further includes an updater 210. A new action 212 is executed in each iteration. The updater 210 updates the environment 204 with the action 212 taken by the intelligent agent 202 after each iteration. The action 212 taken in each iteration is added by the updater 210 to the environment 204. In an embodiment, the transaction data simulator 110 further includes a trimmer 208, which is configured to trim the environment 204. In an embodiment, the trimmer 208 may remove one or more undesirable actions. For example, remove the actions 212 taken in the first ten iterations because these ten iterations deviate far from the target 220 and the similarity is lower than a predefined threshold. In another embodiment, a complete re-initialization of the transaction data simulator 110 may be performed to remove all the accumulated actions in the environment 204 so that the intelligent agent 202 can be started again.
[0053] Figure 4 A flowchart showing an illustrative embodiment of a method 400 for simulating transaction data is shown. At step 402, standard customer transaction behavior data is provided as the target 220. Standard customer transaction behavior represents a group of customers with similar transaction characteristics. The standard customer transaction behavior is obtained by an unsupervised clustering method.
[0054] At step 404, an action 212 is taken to perform multiple transactions in an iteration representing, for example, one day (e.g., 100 transactions per day). Each transaction has information including transaction type, transaction amount, transaction time, transaction location, transaction medium, the second party associated with the transaction, etc.
[0055] At step 406, the environment 204 compares the target 220 with the action 212 taken in that iteration and rewards or punishes the action 212 based on the similarity or deviation from the target 220. The threshold or rule for determining whether the action 212 is similar to the target 220 is predefined and can be adjusted based on the degree of similarity of the user preference to the target 220.
[0056] At step 408, the environment 204 is updated to include the action 212 in the current iteration. The environment 204 includes a set of all old actions.
[0057] At step 410, the policy engine 206 adjusts the policy for determining the next action 212 based on the reward 214 (i.e., reward or punishment). The policy is formulated based on various factors such as the probability of a transaction occurring, the number of transactions per day, transaction amount, transaction type, transaction parties, the transaction frequency of each transaction type, the upper and lower limits of each transaction, the transaction medium, etc. The policy can adjust the weights of these factors based on the reward 214 in each iteration.
[0058] In step 412, in a new iteration, the intelligent agent 202 takes a new action 212. Steps 404 to 412 are repeated until the action 212 is similar enough to the target 220 (step 414). For example, the transaction amount specified in the target 220 is 20 yuan - 3000 yuan. If the transaction amount of each transaction in the action 212 falls within the range of 20 yuan - 3000 yuan, then the action 212 is similar enough to the target 220. In step 416, the artificial customer profile 112 is combined with the last action 212 that includes multiple transactions similar enough to the target 220, thereby generating the simulated customer data 114.
[0059] Since the standard customer transaction data 106 can include abnormal data, such as fraudulent transactions, the simulated customer transaction data 108 can also include abnormal data because the simulated customer transaction data 108 is similar to the standard customer transaction data 106. In the reinforcement learning model, the intelligent agent 202 explores the environment 204 randomly or stochastically, learns a policy from its experience, and updates the policy while exploring to improve the behavior (i.e., transactions) of the intelligent agent 202. In an embodiment, contrary to a random action, during RNG-based exploration, behavioral patterns can occur (e.g., spending "prodigally" until savings are used up, or experiencing "buyer's remorse" for one large purchase, etc.). Abnormal behavioral patterns may indicate fraudulent transactions. For example, the simulated customer JamesCulley can typically make transactions with a transaction amount less than 1000 yuan. Suddenly, there is a transaction with a transaction amount of 5000 yuan, and this suspicious transaction may be a fraudulent transaction (e.g., James Culley's credit card was stolen, or James Culley's checking account was hacked). There are behavioral patterns that occur naturally during exploration. For example, as Figure 3As shown, the simulated customer James Culley received an amount of 12,387.71 yuan in his checking account on January 1, 2014. James Culley spent 474.98 yuan on January 3, 4,400 yuan on January 3, and 3,856.55 yuan on January 4, 2014, using the debit card associated with the checking account. In the next month, James Culley received an amount of 12,387.71 yuan in his checking account on February 1, 2014. James Culley spent 4,500 yuan on February 2, 1,713.91 yuan on February 3, and transferred 8,100 yuan out of the checking account on June 27, 2014, using the debit card associated with the checking account. In this example, the simulated customer James Culley has a tendency to save and spend money, and occasionally makes large purchases. The behavioral pattern makes the simulated customer James Culley behave more realistically (i.e., look more like a real customer rather than a robot). The policy engine 206 generates multiple parameters, such as "behavioral consistency" (the degree of behavioral consistency over a period of time), "consistency volatility" (the frequency of behavioral changes), "behavioral anomalies" (deviations from regular transaction behavior), etc., and the multiple parameters are used to show the different personalities of each simulated customer.
[0060] The transaction data simulator 110 uses abstracted or aggregated real customer data to simulate customer data representative of real customers. The transaction data simulator 110 can provide a large amount of simulated customer data (i.e., simulated transaction data combined with artificial customer profiles), which can be used to train a prediction model for detecting abnormal customer behavior. In addition, the simulated customer data is generated based on abstracted data of real raw customer data rather than the real raw customer data itself, so it is impossible to derive the actual transaction actions of any real customer. Additionally, the transaction data simulator 110 allows for the generation of behavioral patterns for each simulated customer during iteration.
[0061] Figure 5 is a block diagram of an example data processing system 500 in which aspects of the illustrative embodiments are implemented. The data processing system 500 is an example of a computer, such as a server or a client, in which computer-usable code or instructions for implementing the processes of the illustrative embodiments of the present invention are located. In one embodiment, Figure 5 represents a server computing device, such as a server, that implements the cognitive system 100 described herein.
[0062] In the described example, data processing system 500 may employ a hub architecture that includes a north bridge and memory controller hub (NB / MCH) 501, and a south bridge and input / output (I / O) controller hub (SB / ICH) 502. A processing unit 503, main memory 504, and graphics processor 505 may be connected to NB / MCH 501. The graphics processor 505 may be connected to NB / MCH 501 via, for example, an Accelerated Graphics Port (AGP).
[0063] In the described example, a network adapter 506 is connected to SB / ICH 502. An audio adapter 507, keyboard and mouse adapter 508, modem 509, read only memory (ROM) 510, hard disk drive (HDD) 511, optical disk drive (e.g., CD or DVD) 512, universal serial bus (USB) ports and other communication ports 513, and PCI / PCIe devices 514 may be connected to SB / ICH 502 via a bus system 516. The PCI / PCIe devices 514 may include an Ethernet adapter, add-in cards, and PC cards for notebook computers. The ROM 510 may be, for example, a flash basic input / output system (BIOS). The HDD 511 and optical disk drive 512 may use an Integrated Drive Electronics (IDE) or Serial Advanced Technology Attachment (SATA) interface. A Super I / O (SIO) device 515 may be connected to SB / ICH 502.
[0064] An operating system may run on the processing unit 503. The operating system may coordinate and provide control over the various components within the data processing system 500. As a client, the operating system may be a commercially available operating system. An object-oriented programming system, such as Java TM programming system, may run with the operating system and provide calls from object-oriented programs or applications executing on the data processing system 500 to the operating system. As a server, the data processing system 500 may be a eServer TM system or a system of an operating system. eServer is a trademark of International Business Machines Corporation and is registered in many jurisdictions worldwide. Registered trademarks are used under sublicense from the Linux Foundation, a proprietary licensee of Linus Torvalds, the trademark owner worldwide. The data processing system 500 may be a Symmetric Multi-Processor (SMP) system that may include multiple processors in the processing unit 503. Alternatively, a single processor system may be employed.
[0065] Instructions for an operating system, an object-oriented programming system, and applications or programs are located on a storage device such as HDD 511 and are loaded into main memory 504 for execution by processing unit 503. The processes of the embodiments of cognitive system 100 described herein can be executed by processing unit 503 using computer-usable program code, which can be located in a memory such as main memory 504, ROM 510, or in one or more peripheral devices.
[0066] Bus system 516 can include one or more buses. Bus system 516 can be implemented using any type of communication structure or architecture that can provide data transfer between different components or devices attached to the structure or architecture. A communication unit such as modem 509 or network adapter 506 can include one or more devices that can be used to send and receive data.
[0067] Those of ordinary skill in the art will understand that Figure 5 the hardware depicted in the figures may vary according to implementation. In addition to or in place of the hardware described, other internal hardware or peripheral devices may be used, such as flash memory, equivalent non-volatile memory, or optical disk drives.
[0068] Furthermore, data processing system 500 can take the form of any of a variety of different data processing systems, including but not limited to client computing devices, server computing devices, tablet computers, laptop computers, telephones or other communication devices, personal digital assistants, etc. In essence, data processing system 500 can be any known or later-developed data processing system without architectural limitations.
[0069] The systems and processes in the figures are not exclusive. Other systems, processes, and menus can be derived in accordance with the principles of the embodiments described herein to achieve the same purpose. It should be understood that the embodiments and variations shown and described herein are for illustrative purposes only. Those skilled in the art can implement modifications to the current design without departing from the scope of the embodiments. As described herein, various systems, subsystems, agents, managers, and processes can be implemented using hardware components, software components, and / or combinations thereof.
[0070] The present invention can be a system, method, and / or computer program product. The computer program product can include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to execute aspects of the present invention.
[0071] A computer-readable storage medium can be a tangible device that is capable of retaining and storing instructions for use by an instruction execution device. A computer-readable storage medium can be, by way of example and not limitation, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanical encoding device such as a punched card or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium shall not be construed to be a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through an optical fiber cable), or an electrical signal transmitted through a wire.
[0072] The computer-readable program instructions described herein can be downloaded to a respective computing / processing device from a computer-readable storage medium or can be downloaded to an external computer or an external storage device via a network, for example, the Internet, a local area network (LAN), a wide area network (WAN), and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium within the respective computing / processing device.
[0073] The computer-readable program instructions for carrying out operations of the present invention can be assembly instructions, instruction set architecture (ISA) instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Java TM, Smalltalk, C++, etc., as well as conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter case, the remote computer can be connected to the user's computer through any type of network, including a LAN or WAN, or can be connected to an external computer (e.g., using an Internet service provider via the Internet). In some embodiments, to perform aspects of the present invention, an electronic circuit, including, for example, a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), can execute the computer-readable program instructions by utilizing the state information of the computer-readable program instructions to personalize the electronic circuit.
[0074] Aspects of the present invention are described herein with reference to the flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0075] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions executed via the processor of the computer or other programmable data processing apparatus create a means for implementing the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. These computer-readable program instructions can also be stored in a computer-readable storage medium, which can direct a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored comprises an article of manufacture, the article of manufacture including instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams.
[0076] The computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable apparatus, or other devices implement the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams.
[0077] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of instructions, which includes one or more executable instructions for implementing the specified logical function. In some alternative embodiments, the functions recited in the blocks may occur out of the order recited in the figures. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustrations, and combinations of blocks in the block diagrams and / or flowchart illustrations, can be implemented by special-purpose hardware-based systems that perform the specified functions or acts, or combinations of special-purpose hardware and computer instructions
[0078] This specification and the claims may use the terms "a," "at least one," and "one or more" with respect to particular features and elements of the illustrative embodiments. It should be understood that these terms and phrases are intended to indicate that there is at least one particular feature or element in a particular illustrative embodiment, but that there may also be more than one. That is, these terms / phrase are not intended to limit the specification or claims to the presence of a single feature / element or to require the presence of multiple such features / elements. Rather, these terms / phrase only require at least a single feature / element, where the possibility of multiple such features / elements is within the scope of the specification and claims
[0079] Furthermore, it should be understood that the following description uses a variety of different examples of various elements of the illustrative embodiments to further illustrate example implementations of the illustrative embodiments and to assist in understanding the mechanisms of the illustrative embodiments. These examples are intended to be non-limiting and are not an exhaustive listing of the various possibilities for implementing the mechanisms of the illustrative embodiments. Given this specification, it will be apparent to those of ordinary skill in the art that, without departing from the spirit and scope of the present invention, many other alternative implementations of these various elements may be utilized in addition to, or in lieu of, the examples provided herein
[0080] Although the present invention has been described with reference to the exemplary embodiments, the present invention is not limited thereto. Those skilled in the art will understand that many changes and modifications can be made to the preferred embodiments of the present invention, and such changes and modifications can be made without departing from the true spirit of the present invention. Therefore, the appended claims are intended to be construed to cover all such equivalent variations that fall within the true spirit and scope of the present invention
Claims
1. A computer-implemented method in a data processing system, the data processing system including a processor and a memory including instructions, the instructions being executed by the processor to cause the processor to implement a method for simulating customer data using a reinforcement learning model including an intelligent agent, a policy engine, and an environment, the method comprising: generating, by the processor, an artificial customer profile by combining information randomly selected from a set of real customer profile data; providing, by the processor, standard customer transaction data, the standard customer transaction data representing a set of real customers having transaction characteristics similar to a target; executing, by the processor, multiple iterations until the similarity of an action including multiple simulated transactions to the target is higher than a first predetermined threshold, this iteration action being the last action, wherein, in each iteration: executing, by the intelligent agent, the action including multiple simulated transactions; comparing, by the environment, the action with the target; providing, by the environment, feedback associated with the action based on the similarity to the target; and adjusting, by the policy engine, the policy based on the feedback to determine the next action; and combining, by the processor, the artificial customer profile with the last action to form simulated customer data.
2. The method according to claim 1, wherein, the real customer profile data includes one or more of customer address, customer name, contact information, credit information, and income information.
3. The method according to claim 1, wherein, each simulated transaction includes a transaction type, a transaction amount, a transaction time, a transaction location, a transaction medium, and a second party associated with the simulated transaction.
4. The method according to claim 1, wherein, the environment includes a set of all previous actions executed by the intelligent agent.
5. The method according to claim 4, further comprising: removing, by the processor, multiple previous actions having a similarity lower than a second predefined threshold.
6. The method according to claim 1, further comprising: obtaining, by the processor, the standard customer transaction data from the original customer transaction data by an unsupervised clustering method.
7. The method according to claim 1, wherein, the feedback is a reward or a penalty.
8. A system for simulating customer data using a reinforcement learning model including an intelligent agent, a policy engine, and an environment, the system comprising: a processor configured to: generate an artificial customer profile by combining information randomly selected from a set of real customer profile data; provide standard customer transaction data, the standard customer transaction data representing a set of customers having transaction characteristics similar to a target; execute, by the processor, multiple iterations until the similarity of an action including multiple simulated transactions to the target is higher than a first predetermined threshold, this iteration action being the last action, wherein, in each iteration: execute, by the intelligent agent, the action including multiple simulated transactions; compare, by the environment, the action with the target; provide, by the environment, feedback associated with the action based on the similarity to the target; and The policy engine adjusts the policy based on the feedback to determine the next action; and combines the artificial customer profile with the last action to form simulated customer data.
9. The system according to claim 8, wherein the real customer profile data includes one or more of customer address, customer name, contact information, credit information, and income information.
10. The system according to claim 8, wherein the environment includes a set of all previous actions performed by the intelligent agent.
11. The system according to claim 10, wherein before the step of adjusting the policy, the processor is further configured to add the action to the environment.
12. The system according to claim 11, wherein the processor is further configured to remove multiple previous actions having a similarity below a second predefined threshold.
13. The system according to claim 8, wherein the feedback is a reward or a penalty.
14. A computer program product for simulating transaction data, the computer program product comprises: a computer-readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing the method according to any one of claims 1 to 7.
15. A computer-readable storage medium storing computer program instructions that can be loaded into the internal memory of a digital computer, the computer program instructions including a software code portion for performing the method according to any one of claims 1 to 7 when the computer program instructions are run on the computer.
Citation Information
Patent Citations
Method for identifying abnormality of customers in futures market
CN102496127A
Detection method and device for cheating behavior in e-commerce industry
CN106934627A