Methods, systems, and computer program products for improving machine learning model tag accuracy using reinforcement learning
By using reinforcement learning agents (RLA) to process training datasets and train deep learning models, the problem of misprediction in machine learning models is solved, improving accuracy and reducing resource consumption.
Patent Information
- Application Number
- CN202380016943.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-22
- Publication Date
- 2025-07-22
AI Technical Summary
The machine learning model has error affirmation predictions and false negative predictions in the training dataset, resulting in insufficient accuracy and the need to eliminate incorrect predictions requires a large number of network resources.
The initial training data set is processed using the reinforcement learning agent (RLA) machine learning model to generate a second training data set with higher accuracy, and the deep learning model is trained using this data set, the detection rate is evaluated through the test data set, and reward parameters are generated to optimize the model.
Improves label accuracy of machine learning models, reduces the number of incorrect predictions, and reduces the network resources required to eliminate incorrect predictions.
Smart Images

Figure CN120359525A_ABST
Abstract
Description
Technical Field
[0001] The disclosed subject matter generally relates to methods, systems, and computer program products for improving the accuracy of machine learning models, and in some particular embodiments or aspects, to methods, systems, and computer program products for using reinforcement learning to improve the label accuracy of machine learning models. Background Art
[0002] Machine learning can be a field of computer science that uses statistical techniques to provide a computer system with the ability to learn (e.g., gradually improve performance) tasks using data, without explicitly programming the computer system to perform the tasks. In some cases, a machine learning model can be developed for a data set such that the machine learning model can perform tasks related to that data set (e.g., tasks associated with prediction). In one example, a machine learning model can be developed to identify fraudulent transactions (e.g., fraudulent payment transactions). In some cases, the machine learning model can include a fraud detection machine learning model that is used to determine whether a transaction from a set of transactions is a fraudulent transaction based on data associated with the set of transactions.
[0003] In some cases, a machine learning model can be used to classify the outcome of a condition or event. For example, a machine learning model can be used to predict whether a transaction is fraudulent. In such examples, the prediction provided by the machine learning model can include the probability that the condition will occur, where the probability is generated based on training the machine learning model on a large data set. Ultimately, the prediction of the machine learning model can be compared to the ground truth in the actual (e.g., real, true, etc.) scenario where the transaction is fraudulent.
[0004] However, if the training data set includes too few or too many incorrect predictions, such as false positive predictions (e.g., positive results incorrectly predicted by the machine learning model) or false negative predictions (e.g., negative results incorrectly predicted by the machine learning model), then the machine learning model may not be able to accurately predict whether a transaction is fraudulent. In some cases, false positive predictions and / or false negative predictions can be inaccurate predictions based on the specific use of the machine learning model. Additionally, false positive predictions and false negative predictions can lead to a large amount of inaccuracies in the large data set that are difficult to eliminate and may require necessary network resources to eliminate such incorrect predictions. Summary of the Invention
[0005] Accordingly, an object of the presently disclosed subject matter is to provide methods, systems, and computer program products for using reinforcement learning to improve the label accuracy of machine learning models.
[0006] According to a non - limiting embodiment or aspect, a computer - implemented method for improving the label accuracy of a machine - learning model using reinforcement learning is provided, including: receiving, by at least one processor, an initial training data set, where the initial training data set includes a plurality of data instances, where each data instance has a label, and where a first percentage of the plurality of data instances is correctly labeled; providing, by at least one processor, the initial training data set as an input to a reinforcement learning agent (RLA) machine - learning model to generate a second training data set, where the second training data set includes the plurality of data instances, where a second percentage of the plurality of data instances is correctly labeled, and where the second percentage is greater than the first percentage; training, by at least one processor, a deep - learning model using the second training data set to provide a trained deep - learning model; testing, by at least one processor, the trained deep - learning model using a test data set to generate a resulting data set, where the resulting data set has a detection rate, and where the detection rate is an indication of the number of data instances from the test data set that are correctly predicted by the trained deep - learning model; and generating, by at least one processor, a reward parameter based on the detection rate of the resulting data set generated using the trained deep - learning model.
[0007] According to a non - limiting embodiment or aspect, a system for improving the label accuracy of a machine - learning model using reinforcement learning is provided, including at least one processor configured to: receive an initial training data set, where the initial training data set includes a plurality of data instances, where each data instance has a label, and where a first percentage of the plurality of data instances is correctly labeled; provide the initial training data set as an input to a reinforcement learning agent (RLA) machine - learning model to generate a second training data set, where the second training data set includes the plurality of data instances, where a second percentage of the plurality of data instances is correctly labeled, and where the second percentage is greater than the first percentage; train a deep - learning model using the second training data set to provide a trained deep - learning model; test the trained deep - learning model using a test data set to generate a resulting data set, where the resulting data set has a detection rate, and where the detection rate is an indication of the number of data instances from the test data set that are correctly predicted by the trained deep - learning model; and generate a reward parameter based on the detection rate of the resulting data set.
[0008] According to a non - limiting embodiment or aspect, there is provided a computer program product for improving the label accuracy of a machine learning model using reinforcement learning. The computer program product includes at least one non - transient computer - readable medium. The at least one non - transient computer - readable medium includes one or more instructions that, when executed by at least one processor, cause the at least one processor to: receive an initial training data set, where the initial training data set includes a plurality of data instances, where each data instance has a label, and where a first percentage of the plurality of data instances is correctly labeled; provide the initial training data set as an input to a reinforcement learning agent (RLA) machine learning model to generate a second training data set, where the second training data set includes the plurality of data instances, where a second percentage of the plurality of data instances is correctly labeled, and where the second percentage is greater than the first percentage; use the second training data set to train a deep - learning model to provide a trained deep - learning model; use a test data set to test the trained deep - learning model to generate a resulting data set, where the resulting data set has a detection rate, and where the detection rate is an indication of the number of data instances from the test data set that are correctly predicted by the trained deep - learning model; and generate a reward parameter based on the detection rate of the resulting data set.
[0009] Other non - limiting embodiments or aspects are set forth in the following numbered clauses:
[0010] Clause 1: A computer - implemented method, comprising: receiving, by at least one processor, an initial training data set, where the initial training data set includes a plurality of data instances, where each data instance has a label, and where a first percentage of the plurality of data instances is correctly labeled; providing, by at least one processor, the initial training data set as an input to a reinforcement learning agent (RLA) machine learning model to generate a second training data set, where the second training data set includes the plurality of data instances, where a second percentage of the plurality of data instances is correctly labeled, and where the second percentage is greater than the first percentage; training, by at least one processor, a deep - learning model using the second training data set to provide a trained deep - learning model; testing, by at least one processor, the trained deep - learning model using a test data set to generate a resulting data set, where the resulting data set has a detection rate, and where the detection rate is an indication of the number of data instances from the test data set that are correctly predicted by the trained deep - learning model; and generating, by at least one processor, a reward parameter based on the detection rate of the resulting data set generated using the trained deep - learning model.
[0011] Clause 2: The computer-implemented method according to Clause 1 further includes: training the RLA machine learning model using a reinforcement-based learning algorithm based on the reward parameter.
[0012] Clause 3: The computer-implemented method according to Clause 1 or 2, wherein training the RLA machine learning model using the reinforcement-based learning algorithm includes: updating the RLA machine learning model to maximize the reward parameter generated by the trained deep learning model.
[0013] Clause 4: The computer-implemented method according to any one of Clauses 1 to 3, wherein the multiple data instances of the test data set are independent of the multiple data instances of the initial training data set.
[0014] Clause 5: The computer-implemented method according to any one of Clauses 1 to 4, wherein the RLA machine learning model includes a neural network machine learning model.
[0015] Clause 6: The computer-implemented method according to any one of Clauses 1 to 5, wherein the multiple data instances of the initial training data set include multiple data records associated with multiple payment transactions; wherein each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; and wherein the deep learning model is a fraud detection deep learning model configured to provide a prediction as to whether a payment transaction is fraudulent or non-fraudulent.
[0016] Clause 7: The computer-implemented method according to any one of Clauses 1 to 6, wherein the multiple data instances of the test data set include a second multiple data records associated with a second multiple payment transactions; wherein a subset of the data records of the second multiple data records has a label indicating whether each data record in the data record subset is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; and wherein the detection rate is an indication of the number of data records in the data record subset from the test data set that are correctly predicted to be associated with fraudulent payment transactions.
[0017] Clause 8. A system includes: at least one processor configured to: receive an initial training dataset, where the initial training dataset includes a plurality of data instances, where each data instance has a label, and where a first percentage of the plurality of data instances are correctly labeled; provide the initial training dataset as input to a reinforcement learning agent (RLA) machine learning model to generate a second training dataset, where the second training dataset includes the plurality of data instances, where a second percentage of the plurality of data instances are correctly labeled, and where the second percentage is greater than the first percentage; use the second training dataset to train a deep learning model to provide a trained deep learning model; use a test dataset to test the trained deep learning model to generate a resulting dataset, where the resulting dataset has a detection rate, and where the detection rate is an indication of the number of data instances from the test dataset correctly predicted by the trained deep learning model; and generate a reward parameter based on the detection rate of the resulting dataset.
[0018] Clause 9: The system according to Clause 8, where the at least one processor is further configured to: train the RLA machine learning model using a reinforcement-based learning algorithm based on the reward parameter.
[0019] Clause 10: The system according to Clause 8 or 9, where when training the RLA machine learning model using the reinforcement-based learning algorithm, the at least one processor is configured to: update the parameters of the RLA machine learning model to maximize the reward parameter generated by the trained deep learning model.
[0020] Clause 11: The system according to any one of Clauses 8 to 10, where the plurality of data instances of the test dataset are independent of the plurality of data instances of the initial training dataset.
[0021] Clause 12: The system according to any one of Clauses 8 to 11, where the RLA machine learning model includes a neural network machine learning model.
[0022] Clause 13: The system according to any one of Clauses 8 to 12, where the plurality of data instances of the initial training dataset include a plurality of data records associated with a plurality of payment transactions; where each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; and where the deep learning model is a fraud detection deep learning model configured to provide a prediction as to whether a payment transaction is fraudulent or non-fraudulent.
[0023] Clause 14: The system according to any one of Clauses 8 to 13, wherein the plurality of data instances of the test data set include a second plurality of data records associated with a second plurality of payment transactions; wherein a subset of the data records of the second plurality of data records has a label indicating whether each data record in the subset of data records is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; and wherein the detection rate is an indication of the number of data records in the subset of data records from the test data set that are correctly predicted to be associated with fraudulent payment transactions.
[0024] Clause 15: A computer program product comprising a non-transitory computer-readable medium including one or more instructions that, when executed by at least one processor, cause the at least one processor to: receive an initial training data set, wherein the initial training data set includes a plurality of data instances, wherein each data instance has a label, and wherein a first percentage of the plurality of data instances is correctly labeled; provide the initial training data set as input to a reinforcement learning agent (RLA) machine learning model to generate a second training data set, wherein the second training data set includes the plurality of data instances, wherein a second percentage of the plurality of data instances is correctly labeled, and wherein the second percentage is greater than the first percentage; use the second training data set to train a deep learning model to provide a trained deep learning model; use a test data set to test the trained deep learning model to generate a resulting data set, wherein the resulting data set has a detection rate, and wherein the detection rate is an indication of the number of data instances from the test data set that are correctly predicted by the trained deep learning model; and generate a reward parameter based on the detection rate of the resulting data set.
[0025] Clause 16: The computer program product according to Clause 15, wherein the one or more instructions further cause the at least one processor to: train the RLA machine learning model using a reinforcement-based learning algorithm based on the reward parameter.
[0026] Clause 17: The computer program product according to Clause 15 or 16, wherein the one or more instructions that cause the at least one processor to train the RLA machine learning model using the reinforcement-based learning algorithm cause the at least one processor to: update the RLA machine learning model to maximize the reward parameter generated by the trained deep learning model.
[0027] Clause 18: The computer program product according to any one of Clauses 15 to 17, wherein the plurality of data instances of the test data set are independent of the plurality of data instances of the initial training data set.
[0028] Clause 19: A computer program product according to any one of clauses 15 to 18, wherein the RLA machine learning model comprises a neural network machine learning model.
[0029] Clause 20: A computer program product according to any one of clauses 15 to 19, wherein the plurality of data instances of the initial training dataset comprises a plurality of data records associated with a plurality of payment transactions; wherein each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; wherein the deep learning model is a fraud detection deep learning model configured to provide a prediction as to whether a payment transaction is fraudulent or non-fraudulent; wherein the plurality of data instances of the test dataset comprises a second plurality of data records associated with a second plurality of payment transactions; wherein a subset of the data records of the second plurality of data records has a label indicating whether each data record in the data record subset is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; and wherein the detection rate is an indication of the number of data records in the data record subset from the test dataset that are correctly predicted to be associated with fraudulent payment transactions.
[0030] After considering the following description and the appended claims, these and other features and characteristics of the presently disclosed subject matter, as well as the methods of operation and functions of the related structural elements and combinations of the various parts and the manufacturing economy, will become more apparent, all of which form part of this specification, wherein like reference numerals refer to corresponding parts in the various figures. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended as a definition of the limits of the disclosed subject matter. Unless the context clearly dictates otherwise, the singular forms "a" and "the" as used in this specification and the claims include plural referents. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The additional advantages and details of the disclosed subject matter are explained in more detail below with reference to the exemplary embodiments or aspects shown in the accompanying drawings, in which:
[0032] Figure 1 is a diagram of a non-limiting embodiment or aspect of an environment in which the methods, systems, and / or computer program products described herein can be implemented in accordance with the principles of the presently disclosed subject matter;
[0033] Figure 2 is Figure 1 a diagram of a non-limiting embodiment or aspect of components of one or more devices;
[0034] Figure 3is a flowchart of a non - limiting embodiment or aspect of a process for using reinforcement learning to improve the label accuracy of a machine learning model; and
[0035] Figures 4A - 4F is a schematic diagram of a non - limiting embodiment or aspect of an implementation of a process for using reinforcement learning to improve the label accuracy of a machine learning model according to some non - limiting embodiments or aspects. Detailed Description
[0036] For the purposes of the description below, the terms "end", "upper", "lower", "right", "left", "vertical", "horizontal", "top", "bottom", "lateral", "longitudinal" and their derivatives shall refer to the orientation of the disclosed subject matter as it appears in the figures. However, it should be understood that the disclosed subject matter may assume various alternative variations and step sequences, unless explicitly specified to the contrary. It should also be understood that the specific devices and processes shown in the figures and described in the following specification are merely exemplary embodiments or aspects of the disclosed subject matter. Accordingly, unless otherwise indicated, the specific dimensions and other physical characteristics associated with the embodiments or aspects disclosed herein should not be considered limiting.
[0037] As used herein, aspects, components, elements, structures, acts, steps, functions, instructions, etc. should not be construed as critical or essential unless explicitly so described. Also, as used herein, the article "a" is intended to include one or more items and may be interchanged with "one or more" and "at least one". Further, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, combinations of related and unrelated items, etc.) and may be interchanged with "one or more" or "at least one". Where only one item is intended, the term "one" or similar language is used. Moreover, as used herein, the term "having" etc. is intended to be an open - ended term. Additionally, unless otherwise explicitly stated, the phrase "based on" is intended to mean "at least partially based on".
[0038] As used herein, the terms "communicate" and "transfer" can refer to the receipt, acceptance, sending, migration, provision, etc. of information (e.g., data, signals, messages, instructions, commands, etc.). A unit (e.g., a device, a system, a component of a device or system, a combination thereof, etc.) communicating with another unit means that the one unit is capable of receiving information from and / or sending information to the other unit, either directly or indirectly. This can refer to a direct or indirect connection, which can be wired and / or wireless in nature (e.g., a direct communication connection, an indirect communication connection, etc.). Additionally, although the information being sent can be modified, processed, relayed, and / or routed between a first unit and a second unit, the two units can still communicate with each other. For example, even if the first unit receives information passively and does not actively send information to the second unit, the first unit can still communicate with the second unit. As another example, if at least one intermediate unit (e.g., a third unit located between the first unit and the second unit) processes the information received from the first unit and transfers the processed information to the second unit, the first unit can communicate with the second unit. In some non-limiting embodiments or aspects, a message can refer to a network packet (e.g., a data packet, etc.) that includes data. It should be understood that there can be many other arrangements.
[0039] As used herein, the terms "issuing institution", "payment device issuer", "issuer", or "issuing bank" can refer to one or more entities that provide an account to a customer for conducting transactions (e.g., payment transactions) (e.g., initiating credit and / or debit payments). For example, an issuing institution can provide an account identifier that uniquely identifies one or more accounts associated with the customer, such as a primary account number (PAN). The account identifier can be implemented on a payment device such as a physical financial instrument like a payment card, and / or can be electronic and used for electronic payments. The terms "issuing institution" and "issuing institution system" can also refer to one or more computer systems operated by or on behalf of the issuing institution, such as a server computer that executes one or more software applications. For example, an issuing institution system can include one or more authorization servers for authorizing transactions.
[0040] As used herein, the term "account identifier" may include one or more types of identifiers associated with a user account (e.g., PAN, card number, payment card number, payment token, etc.). In some non-limiting embodiments or aspects, an issuer institution may provide an account identifier (e.g., PAN, payment token, etc.) to a user, which uniquely identifies one or more accounts associated with the user. The account identifier may be embodied on a physical financial instrument (e.g., a portable financial instrument, payment card, credit card, debit card, etc.), and / or may be electronic information transmitted to the user for the user to use for electronic payments. In some non-limiting embodiments or aspects, the account identifier may be an original account identifier, where the original account identifier is provided to the user when the account associated with the account identifier is created. In some non-limiting embodiments or aspects, the account identifier may be an account identifier provided to the user after the original account identifier is provided to the user (e.g., a supplementary account identifier). For example, if the original account identifier is forgotten, stolen, etc., the supplementary account identifier may be provided to the user. In some non-limiting embodiments or aspects, the account identifier may be directly or indirectly associated with the issuer institution such that the account identifier may be a payment token mapped to a PAN or other type of identifier. The account identifier may be any combination of alphanumeric, characters, and / or symbols, etc. The issuer institution may be associated with a bank identification number (BIN) that uniquely identifies the issuer institution.
[0041] As used herein, the term "merchant" may refer to one or more entities (e.g., an operator of a retail enterprise that provides goods and / or services and / or access to goods and / or services to a user (e.g., a customer, consumer, customer of the merchant, etc.) based on a transaction (e.g., a payment transaction)). As used herein, the term "merchant system" may refer to one or more computer systems operated by or on behalf of a merchant, such as a server computer that executes one or more software applications. As used herein, the term "product" may refer to one or more goods and / or services provided by a merchant.
[0042] As used herein, a "point-of-sale (POS) device" may refer to one or more devices that may be used by a merchant to initiate a transaction (e.g., a payment transaction), participate in a transaction, and / or process a transaction. For example, a POS device may include one or more computers, peripheral devices, card readers, near field communication (NFC) receivers, radio frequency identification (RFID) receivers, and / or other contactless transceivers or receivers, contact-based receivers, payment terminals, computers, servers, input devices, etc.
[0043] As used herein, a "Point of Sale (POS) system" can refer to one or more computers and / or peripheral devices used by a merchant to conduct transactions. For example, a POS system can include one or more POS devices, and / or other similar devices that can be used to conduct payment transactions. A POS system (e.g., a merchant POS system) can also include one or more server computers configured to process online payment transactions via a web page, a mobile application, etc.
[0044] As used herein, the term "transaction service provider" can refer to an entity that receives a transaction authorization request from a merchant or other entity and, in some cases, provides payment guarantee through an agreement between the transaction service provider and the issuer institution. In some non-limiting embodiments or aspects, the transaction service provider can include a credit card company, a debit card company, etc. As used herein, the term "transaction service provider system" can also refer to one or more computer systems operated by or on behalf of the transaction service provider, such as a transaction processing server that executes one or more software applications. The transaction processing server can include one or more processors and, in some non-limiting embodiments or aspects, can be operated by or on behalf of the transaction service provider.
[0045] As used herein, the term "acquirer" can refer to an entity authorized by and approved by a transaction service provider to initiate a transaction (e.g., a payment transaction) using a payment device associated with the transaction service provider. As used herein, the term "acquirer system" can also refer to one or more computer systems, computer devices, etc. operated by or on behalf of the acquirer. A transaction can include a payment transaction (e.g., a purchase, an Original Credit Transaction (OCT), an Account Funding Transaction (AFT), etc.). In some non-limiting embodiments or aspects, the acquirer can be authorized by the transaction service provider to assign a merchant or service provider to use the payment device of the transaction service provider to initiate a transaction. The acquirer can enter into a contract with a payment service provider to enable the payment service provider to provide initiation to the merchant. The acquirer can monitor the compliance of the payment service provider according to the transaction service provider regulations. The acquirer can conduct due diligence on the payment service provider and ensure appropriate due diligence is conducted before signing a contract with the merchant to be initiated. The acquirer may be responsible for all transaction service provider programs operated or initiated by the acquirer. The acquirer can be responsible for the actions of the acquirer payment service provider, the merchants initiated by the acquirer payment service provider, etc. In some non-limiting embodiments or aspects, the acquirer can be a financial institution, such as a bank.
[0046] As used herein, the terms "electronic wallet", "electronic wallet mobile application", and "digital wallet" may refer to one or more electronic devices and / or one or more software applications configured to initiate and / or conduct transactions (e.g., payment transactions, electronic payment transactions, etc.). For example, an electronic wallet may include a user device (e.g., a mobile device) executing an application, as well as server-side software and / or a database for maintaining transaction data and providing the transaction data to the user device. As used herein, the term "electronic wallet provider" may include an entity that provides and / or maintains an electronic wallet and / or an electronic wallet mobile application for a user (e.g., a customer). Examples of electronic wallet providers include, but are not limited to, Google Android Apple and Samsung In some non-limiting examples, a financial institution (e.g., an issuer institution) may be an electronic wallet provider. As used herein, the term "electronic wallet provider system" may refer to one or more computer systems, computer devices, servers, server clusters, etc. operated by or on behalf of an electronic wallet provider.
[0047] As used herein, the term "payment device" may refer to a payment card (e.g., a credit card or a debit card), a gift card, a smart card, a smart medium, a payroll card, a healthcare card, a wristband, a machine-readable medium containing account information, a keychain device or fob, an RFID transponder, a retailer discount or membership card, a cellular phone, an electronic wallet mobile application, a personal digital assistant (PDA), a pager, a security card, a computer, an access card, a wireless terminal, a transponder, etc. In some non-limiting embodiments or aspects, a payment device may include volatile or non-volatile memory for storing information (e.g., an account identifier, an account holder name, etc.).
[0048] As used herein, the term "payment gateway" may refer to an entity and / or a payment processing system operated by or on behalf of such entity, where the entity (e.g., merchant service provider, payment service provider, payment facilitator, payment facilitator under contract with an acquirer, payment aggregator, etc.) provides payment services (e.g., transaction service provider payment services, payment processing services, etc.) to one or more merchants. The payment services may be associated with the use of payment devices managed by the transaction service provider. As used herein, the term "payment gateway system" may refer to one or more computer systems, computer devices, servers, server groups, etc. operated by or on behalf of the payment gateway, and / or the payment gateway itself. As used herein, the term "payment gateway mobile application" may refer to one or more electronic devices and / or one or more software applications configured to provide payment services for transactions (e.g., payment transactions, electronic payment transactions, etc.).
[0049] As used herein, the terms "client" and "client device" may refer to one or more client-side devices or systems (e.g., remotely from the transaction service provider) used to initiate or facilitate a transaction (e.g., a payment transaction). By way of example, "client device" may refer to one or more POS devices used by a merchant, one or more acquirer host computers used by an acquirer, one or more mobile devices used by a user, etc. In some non-limiting embodiments or aspects, the client device may be an electronic device configured to communicate with one or more networks and initiate or facilitate a transaction. For example, the client device may include one or more computers, laptop computers, tablet computers, mobile devices, cellular phones, wearable devices (e.g., watches, glasses, lenses, clothing, etc.), PDAs, etc. Additionally, "client" may also refer to an entity (e.g., merchant, acquirer, etc.) that owns, utilizes, and / or operates the client device to initiate a transaction (e.g., to initiate a transaction with the transaction service provider).
[0050] As used herein, the term "computing device" can refer to one or more electronic devices configured to process data. A computing device can be a mobile device, a desktop computer, and / or any other similar device. Additionally, the term "computer" can refer to any computing device that includes the necessary components for receiving, processing, and outputting data and typically includes a display, a processor, memory, an input device, and a network interface. As used herein, the term "server" can refer to or include one or more processors or computers, storage devices, or similar computer arrangements that are operated or facilitate communication and processing by multiple parties in a network environment such as the Internet, but it should be understood that communication can be facilitated through one or more public or private network environments and various other arrangements are possible. Additionally, multiple computers (e.g., servers) or other computerized devices (e.g., POS devices) that communicate directly or indirectly in a network environment can constitute a "system" such as a merchant's POS system.
[0051] As used herein, the term "processor" can represent any type of processing unit, such as a single processor with one or more cores, one or more cores of one or more processors, multiple processors each having one or more cores, and / or other arrangements and combinations of processing units.
[0052] As used herein, the term "system" can refer to one or more computing devices or a combination of computing devices (e.g., processors, servers, client devices, software applications, components of these devices, etc.). As used herein, references to "device", "server", "processor", etc. can refer to the previously stated device, server, or processor that is stated to perform a previous step or function, a different server or processor, and / or a combination of servers and / or processors. For example, as used in the specification and claims, a first server or first processor stated to perform a first step or first function can refer to the same or different server or the same or different processor stated to perform a second step or second function.
[0053] Non-limiting embodiments or aspects of the disclosed subject matter relate to methods, systems, and computer program products for using reinforcement learning to improve the label accuracy of machine learning models. In some non-limiting embodiments or aspects, a reinforcement learning system may be configured to: receive an initial training data set, wherein the initial training data set includes a plurality of data instances, wherein each data instance has a label, and wherein a first percentage of the plurality of data instances is correctly labeled; provide the initial training data set as input to a reinforcement learning agent (RLA) machine learning model to generate a second training data set, wherein the second training data set includes the plurality of data instances, wherein a second percentage of the plurality of data instances is correctly labeled, and wherein the second percentage is greater than the first percentage; use the second training data set to train a deep learning model to provide a trained deep learning model; use a test data set to test the trained deep learning model to generate a resulting data set, wherein the resulting data set has a detection rate, and wherein the detection rate is an indication of the number of data instances from the test data set correctly predicted by the trained deep learning model; and generate a reward parameter based on the detection rate of the resulting data set.
[0054] In some non-limiting embodiments or aspects, the reinforcement learning system may be configured to train the RLA machine learning model using a reinforcement-based learning algorithm based on the reward parameter. In some non-limiting embodiments or aspects, when training the RLA machine learning model using a reinforcement-based learning algorithm, the reinforcement learning system may update the parameters of the RLA machine learning model to maximize the reward parameter generated by the trained deep learning model. In some non-limiting embodiments or aspects, the plurality of data instances of the test data set are independent of the plurality of data instances of the initial training data set. In some non-limiting embodiments or aspects, the RLA machine learning model includes a neural network machine learning model. In some non-limiting embodiments or aspects, the plurality of data instances of the initial training data set includes a plurality of data records associated with a plurality of payment transactions. In some non-limiting embodiments or aspects, each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non-fraudulent payment transaction. In some non-limiting embodiments or aspects, the deep learning model is a fraud detection deep learning model configured to provide a prediction as to whether a payment transaction is fraudulent or non-fraudulent.
[0055] In some non - limiting embodiments or aspects, multiple data instances of a test data set include a second plurality of data records associated with a second plurality of payment transactions. In some non - limiting embodiments or aspects, a subset of the data records of the second plurality of data records has a label indicating whether each data record in the data record subset is associated with a fraudulent payment transaction or a non - fraudulent payment transaction. In some non - limiting embodiments or aspects, the detection rate is an indication of the number of data records in a data record subset from the test data set that are correctly predicted to be associated with fraudulent payment transactions.
[0056] In this way, non - limiting embodiments or aspects of the present disclosure provide a reinforcement learning system that can provide a data set with a reduced number of incorrect predictions, which can be used to improve the accuracy of predictions provided by a machine learning model, such as a deep - learning fraud detection model. Additionally, the reinforcement learning system can reduce the network resources required to eliminate incorrect predictions in the data set.
[0057] For illustrative purposes, in the following description, although the presently disclosed subject matter is described in terms of methods, systems, and computer program products for using reinforcement learning to improve machine - learning model label accuracy (e.g., during processing of payment transactions), those skilled in the art will recognize that the disclosed subject matter is not limited to the non - limiting embodiments or aspects disclosed herein. For example, the methods, systems, and computer program products described herein can be used with a wide variety of settings, such as using reinforcement learning to improve machine - learning model label accuracy in any suitable setting, such as prediction, regression, classification, fraud prevention, authorization, authentication, identification, feature selection, and the like.
[0058] Now refer to Figure 1 , Figure 1 is a diagram of an example environment 100 in which the apparatuses, systems, and / or methods described herein can be implemented. As Figure 1 shown, the environment 100 includes a reinforcement learning system 102, a data source 102a, a transaction service provider system 104, an issuer system 106, a user device 108, and a communication network 110. The reinforcement learning system 102, the data source 102a, the transaction service provider system 104, the issuer system 106, and / or the user device 108 can be interconnected via a wired connection, a wireless connection, or a combination of wired and wireless connections (e.g., establish connections for communication).
[0059] The reinforcement learning system 102 may include one or more devices configured to communicate with a transaction service provider system 104, an issuer system 106, and / or a user device 108 via a communication network 110. For example, the reinforcement learning system 102 may include a server, a server group, and / or other similar devices. In some non-limiting embodiments or aspects, the reinforcement learning system 102 may be associated with the issuer system 106. For example, the reinforcement learning system 102 may be operated by the issuer system 106. In another example, the reinforcement learning system 102 may be a component of the issuer system 106. In some non-limiting embodiments or aspects, the reinforcement learning system 102 may communicate with a data source 102a, which may be local or remote to the reinforcement learning system 102. In some non-limiting embodiments or aspects, the reinforcement learning system 102 is capable of receiving (e.g., via pull retrieval) information from the data source 102a, storing information in the data source, sending information to the data source, and / or searching for information stored in the data source.
[0060] The transaction service provider system 104 may include one or more devices configured to communicate with the reinforcement learning system 102, the issuer system 106, and / or the user device 108 via the communication network 110. In some non-limiting embodiments or aspects, the transaction service provider system 104 may include a server, a server group, and / or other similar devices. In some non-limiting embodiments or aspects, the transaction service provider system 104 is associated with the issuer. For example, the transaction service provider system 104 may be operated by the issuer.
[0061] The issuer system 106 may include one or more devices configured to communicate with the reinforcement learning system 102, the transaction service provider system 104, and / or the user device 108 via the communication network 110. For example, the issuer system 106 may include a computing device, such as a server, a server group, and / or other similar devices. In some non-limiting embodiments or aspects, the issuer system 106 may be associated with the transaction service provider system.
[0062] The user device 108 may include a computing device configured to communicate with the reinforcement learning system 102, the transaction service provider system 104, and / or the issuer system 106 via the communication network 110. For example, the user device 108 may include a computing device, such as a desktop computer, a portable computer (e.g., a tablet computer, a laptop computer, etc.), a mobile device (e.g., a cellular phone, a smartphone, a personal digital assistant, a wearable device, etc.), and / or other similar devices. In some non-limiting embodiments or aspects, the user device 108 may be associated with a user (e.g., an individual operating the user device 108).
[0063] The communication network 110 may include one or more wired and / or wireless networks. For example, the communication network 110 may include a cellular network (e.g., a Long-Term Evolution (LTE) network, a Third Generation (3G) network, a Fourth Generation (4G) network, a Fifth Generation (5G) network, a Code Division Multiple Access (CDMA) network, etc.), a Public Land Mobile Network (PLMN), a Local Area Network (LAN), a Wide Area Network (WAN), a Metropolitan Area Network (MAN), a telephone network (e.g., a Public Switched Telephone Network (PSTN), etc.), a private network, an ad-hoc network, an intranet, the Internet, a fiber-optic based network, a cloud computing network, etc. and / or a combination of some or all of these or other types of networks.
[0064] Figure 1 The number and arrangement of the systems, devices, and / or networks shown are provided as an example. There may be additional systems, devices, and / or networks, fewer systems, devices, and / or networks, different systems, devices, and / or networks, and / or systems, devices, and / or networks arranged in a different manner than Figure 1 those shown. Additionally, two or more of the systems or devices shown may be implemented within a single system and / or device, or Figure 1 the single system or device shown in Figure 1 may be implemented as multiple distributed systems or devices. Additionally or alternatively, a set of systems (e.g., one or more systems) and / or a set of devices (e.g., one or more devices) of the environment 100 may perform one or more functions described as being performed by another set of systems or another set of devices of the environment 100.
[0065] Now referring to Figure 2 , Figure 2 is a diagram of example components of the device 200. The device 200 may correspond to one or more devices of the reinforcement learning system 102 (e.g., one or more devices of the reinforcement learning system 102), the transaction service provider system 104 (e.g., one or more devices of the transaction service provider system 104), the issuer system 106, and / or the user device 108. In some non-limiting embodiments or aspects, the reinforcement learning system 102, the transaction service provider system 104, the issuer system 106, and / or the user device 108 may include at least one device 200 and / or at least one component of the device 200.
[0066] As Figure 2As shown, the apparatus 200 may include a bus 202, a processor 204, a memory 206, a storage component 208, an input component 210, an output component 212, and a communication interface 214. The bus 202 may include components that permit communication among the components of the apparatus 200. In some non-limiting embodiments or aspects, the processor 204 may be implemented in hardware, software, firmware, and / or any combination thereof. For example, the processor 204 may include a processor (e.g., a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), etc.), a microprocessor, a digital signal processor (DSP), and / or any processing component that can be programmed to perform a certain function (e.g., a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), etc.). The memory 206 may include a random access memory (RAM), a read only memory (ROM), and / or another type of dynamic or static storage device (e.g., a flash memory, a magnetic memory, an optical memory, etc.) that stores information and / or instructions for use by the processor 204.
[0067] The storage component 208 may store information and / or software associated with the operation and use of the apparatus 200. For example, the storage component 208 may include a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, a solid state disk, etc.), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cassette tape, a magnetic tape, and / or another type of computer-readable medium, as well as a corresponding drive.
[0068] The input component 210 may include components that permit the apparatus 200 to receive information, for example, via user input (e.g., a touch screen display, a keyboard, a keypad, a mouse, a button, a switch, a microphone, a camera, etc.). Additionally or alternatively, the input component 210 may include sensors for sensing information (e.g., a global positioning system (GPS) component, an accelerometer, a gyroscope, an actuator, etc.). The output component 212 may include components that provide output information from the apparatus 200 (e.g., a display, a speaker, one or more light emitting diodes (LEDs), etc.).
[0069] The communication interface 214 may include transceiver-type components (e.g., a transceiver, separate receiver and transmitter, etc.) that enable the apparatus 200 to communicate with other devices, for example, via a wired connection, a wireless connection, or a combination of wired and wireless connections. The communication interface 214 may permit the apparatus 200 to receive information from another device and / or provide information to another device. For example, the communication interface 214 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, interface, interface, interface, a cellular network interface, etc.
[0070] Device 200 may perform one or more processes described herein. Device 200 may execute these processes based on software instructions stored on a computer-readable medium such as memory 206 and / or storage component 208 by a processor 204. A computer-readable medium (e.g., a non-transitory computer-readable medium) is defined herein as a non-transitory memory device. A non-transitory memory device includes a memory space located within a single physical storage device or a memory space spread across multiple physical storage devices.
[0071] The software instructions may be read into memory 206 and / or storage component 208 from another computer-readable medium or from another device via communication interface 214. When executed, the software instructions stored in memory 206 and / or storage component 208 may cause processor 204 to perform one or more processes described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with the software instructions to perform one or more processes described herein. Accordingly, embodiments or aspects described herein are not limited to any particular combination of hardware circuitry and software.
[0072] Figure 2 The number and arrangement of the components shown are provided as an example. In some non-limiting embodiments or aspects, device 200 may include additional components, fewer components, different components, or components arranged in a different manner than Figure 2 those shown. Additionally or alternatively, a set of components (e.g., one or more components) of device 200 may perform one or more functions described as being performed by another set of components of device 200.
[0073] Now refer to Figure 3 , Figure 3 is a flowchart of a non-limiting embodiment or aspect of process 300 for using reinforcement learning to improve the label accuracy of a machine learning model. In some non-limiting embodiments or aspects, one or more steps of process 300 may be performed by a reinforcement learning system 102 (e.g., one or more devices of reinforcement learning system 102) (e.g., fully, partially, etc.). In some non-limiting embodiments or aspects, one or more steps of process 300 may be performed by another device or group of devices (e.g., fully, partially, etc.) separate from or including reinforcement learning system 102 (e.g., one or more devices of reinforcement learning system 102), transaction service provider system 104 (e.g., one or more devices of transaction service provider system 104), issuer system 106 (e.g., one or more devices of issuer system 106), and / or user device 108.
[0074] As Figure 3As shown, at step 302, process 300 includes receiving an initial training dataset that includes a certain percentage of correctly labeled data instances. For example, reinforcement learning system 102 may receive an initial training dataset that includes a certain percentage of correctly labeled data instances. In some non-limiting embodiments or aspects, the initial training dataset may include multiple data instances (e.g., data records), where each data instance has a label. For example, the initial training dataset may include multiple data instances, where each data instance has a label indicating whether the data instance is associated with a non-fraudulent data instance (e.g., a data instance associated with a non-fraudulent payment transaction) or a fraudulent data instance (e.g., a data instance associated with a fraudulent payment transaction). In some non-limiting embodiments or aspects, the initial training dataset includes multiple data records associated with multiple payment transactions, where each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non-fraudulent payment transaction.
[0075] In some non-limiting embodiments or aspects, the correctly labeled data instances may include data instances that are verified (e.g., reported, reported by an entity of the payment network, such as reported by transaction service provider system 104, issuer system 106, etc.) as non-fraudulent data instances associated with a transaction and / or fraudulent data instances verified as associated with a transaction. In some non-limiting embodiments or aspects, the initial training dataset may include multiple data instances that are not fully labeled, such as multiple data instances that are not verified as non-fraudulent data instances associated with a transaction and / or multiple data instances that are not verified as fraudulent data instances associated with a transaction. In some non-limiting embodiments or aspects, the initial training dataset may include multiple data instances that are incorrectly labeled, such as multiple data instances that are not labeled as non-fraudulent data instances associated with a transaction but should be labeled as fraudulent data instances associated with a transaction, and / or multiple data instances that are labeled as fraudulent data instances associated with a transaction but should be labeled as non-fraudulent data instances associated with a transaction.
[0076] In some examples, the initial training dataset can include a large number of labeled data instances, such as 100 data instances, 500 data instances, 1,000 data instances, 5,000 data instances, 10,000 data instances, 25,000 data instances, 50,000 data instances, 100,000 data instances, 1,000,000 data instances, and so on. In some non-limiting embodiments or aspects, a percentage (e.g., a first percentage) of the multiple data instances is correctly labeled (e.g., correctly labeled with a positive label for binary classification, correctly labeled with a negative label for binary classification, etc.). In some non-limiting embodiments or aspects, the multiple data instances are labeled based on labels provided as the output of a deep learning model (e.g., a deep learning fraud detection model).
[0077] In some non-limiting embodiments or aspects, the reinforcement learning system 102 can receive the initial training dataset from the data source 102a. Additionally or alternatively, the model management system 102 can receive the initial training dataset from the transaction service provider system 104, the issuer system 106, the user device 108, or other systems or devices. In some non-limiting embodiments or aspects, the initial dataset can include multiple historical data instances. In some non-limiting embodiments or aspects, the initial dataset can include data (e.g., transaction data) associated with historical payment transactions made using one or more payment processing networks (e.g., one or more payment processing networks associated with the transaction service provider system 104).
[0078] In some non-limiting embodiments or aspects, the initial dataset can include multiple data instances associated with multiple features. In some non-limiting embodiments or aspects, the multiple data instances can represent multiple transactions (e.g., electronic payment transactions) made by one or more account holders (e.g., one or more users, such as users associated with the user device 108).
[0079] In some non - limiting embodiments or aspects, a data instance (e.g., a data instance of a data set such as an initial training data set, a second training data set, or a test data set) can include transaction data associated with a payment transaction. In some non - limiting embodiments or aspects, the transaction data can include a plurality of transaction parameters associated with an electronic payment transaction. In some non - limiting embodiments or aspects, a plurality of features can represent the plurality of transaction parameters. In some non - limiting embodiments or aspects, the plurality of transaction parameters can include electronic wallet card data associated with an electronic card (e.g., an electronic credit card, an electronic debit card, an electronic membership card, etc.), decision data associated with a decision (e.g., a decision to approve or reject a transaction authorization request), authorization data associated with an authorization response (e.g., an approved spending limit, an approved transaction value, etc.), PAN, an authorization code (e.g., PIN, etc.), data associated with a transaction amount (e.g., an approved limit, a transaction value, etc.), data associated with a transaction date and time, data associated with a currency exchange rate, data associated with a merchant type (e.g., a merchant category code indicating a type of goods such as groceries, fuel, etc.), data associated with an acquirer country, data associated with an identifier of a country associated with the PAN, data associated with a response code, data associated with a merchant identifier (e.g., a merchant name, a merchant location, etc.), data associated with a currency type corresponding to funds stored associated with the PAN, etc.
[0080] As Figure 3As shown, at step 304, process 300 includes generating a second training dataset that includes a greater percentage of correctly labeled data instances compared to the initial training dataset. For example, the reinforcement learning system 102 can generate a second training dataset that includes a greater percentage of correctly labeled data instances compared to the initial training dataset. In some non-limiting embodiments or aspects, the reinforcement learning system 102 can provide the initial training dataset (e.g., all data instances of the initial training dataset, a subset of all data instances of the initial training dataset, etc.) as input to a reinforcement learning agent (RLA) machine learning model to generate a second training dataset (e.g., an enhanced training dataset). For example, the reinforcement learning system 102 can provide each data instance of the initial training dataset as input to the RLA machine learning model, and the RLA machine learning model can generate an output based on the input. In some non-limiting embodiments or aspects, the output of the RLA machine learning model can include a label for the data instance provided as input. For example, the output of the RLA machine learning model can include a label indicating whether the data instance is predicted to be a non-fraudulent data instance (e.g., a data instance associated with a non-fraudulent payment transaction) or a fraudulent data instance (e.g., a data instance associated with a fraudulent payment transaction). In some non-limiting embodiments or aspects, the RLA machine learning model can include a neural network machine learning model. In some non-limiting embodiments or aspects, the RLA machine learning model can include a neural network machine learning model with a large number of nodes. For example, the RLA machine learning model can include a neural network machine learning model with 10 nodes, 20 nodes, 50 nodes, 100 nodes, 1,000 nodes, etc.
[0081] In some non-limiting embodiments or aspects, the second training dataset includes a plurality of data instances, where each data instance has a label. For example, the second training dataset can include a plurality of data instances, where each data instance has a label indicating whether the data instance is associated with a non-fraudulent data instance (e.g., a data instance associated with a non-fraudulent payment transaction) or a fraudulent data instance (e.g., a data instance associated with a fraudulent payment transaction). In some non-limiting embodiments or aspects, the second training dataset includes a plurality of data records associated with a plurality of payment transactions, where each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non-fraudulent payment transaction.
[0082] In some examples, the second training dataset includes the same number of data instances as the initial dataset, and the labels of the second training dataset may be different from the labels of the initial training dataset (e.g., one or more labels of the second training dataset may be different from one or more labels of the initial training dataset). In such examples, the percentage of the multiple data instances in the second dataset (e.g., the second percentage) is correctly labeled. In some non-limiting embodiments or aspects, the percentage of correctly labeled data instances in the second dataset is greater than the percentage of correctly labeled data instances in the initial dataset.
[0083] As Figure 3 shown, at step 306, process 300 includes training a deep learning model based on the second training dataset. For example, the reinforcement learning system 102 may train a deep learning model based on the second training dataset (e.g., a deep learning model for labeling multiple data instances of the initial training dataset). In some non-limiting embodiments or aspects, the reinforcement learning system 102 may train a deep learning model by providing each data instance of the second training dataset as input and generating an output based on the input. In some non-limiting embodiments or aspects, the output of the deep learning model may include a label indicating whether the data instance is predicted to be a non-fraudulent data instance (e.g., a data instance associated with a non-fraudulent payment transaction) or a fraudulent data instance (e.g., a data instance associated with a fraudulent payment transaction). In some non-limiting embodiments or aspects, the reinforcement learning system 102 may train a deep learning model based on a loss function. In some non-limiting embodiments or aspects, the deep learning model is a fraud detection deep learning model configured to provide a prediction as to whether a payment transaction is fraudulent or non-fraudulent. In some non-limiting embodiments or aspects, the deep learning model may include a neural network machine learning model having a large number of nodes. For example, the deep learning model may include a neural network machine learning model having 10 nodes, 20 nodes, 50 nodes, 100 nodes, 1,000 nodes, etc.
[0084] In some non-limiting embodiments or aspects, the reinforcement learning system 102 may initialize the parameters of the RLA machine learning model. For example, the reinforcement learning system 102 may initialize the parameters of the RLA machine learning model based on a standard normal distribution (e.g., using numbers from a standard normal distribution). In some non-limiting embodiments or aspects, the reinforcement learning system 102 may train a deep learning model based on initializing the parameters of the RLA machine learning model.
[0085] In some non - limiting embodiments or aspects, the reinforcement learning system 102 can train a deep learning model based on a second training dataset each time the reinforcement learning system 102 generates the second training dataset based on an initial training dataset. For example, the reinforcement learning system 102 can update (e.g., train, retrain, test, and / or validate, etc.) the RLA machine learning model based on the reward parameter each time the reward parameter is generated, and the reinforcement learning system 102 can use the updated RLA machine learning model to generate a second training dataset based on the initial training dataset each time the RLA machine learning model is updated.
[0086] As Figure 3 shown, at step 308, process 300 includes using the trained deep learning model to generate a resulting dataset with a detection rate. For example, the reinforcement learning system 102 can use the trained deep learning model to generate a resulting dataset with a detection rate. In some non - limiting embodiments or aspects, the detection rate can be an indication of the number of data instances correctly predicted by the trained deep learning model from a test dataset. In some non - limiting embodiments or aspects, the reinforcement learning system 102 can use the test dataset to test the trained deep learning model. For example, the reinforcement learning system 102 can use a test dataset including multiple data instances (e.g., multiple data records) to test the deep learning model, where each data instance has a label. In some non - limiting embodiments or aspects, the test dataset can include multiple data instances, where each data instance has a label indicating whether the data instance is associated with a non - fraudulent data instance (e.g., a data instance associated with a non - fraudulent payment transaction) or a fraudulent data instance (e.g., a data instance associated with a fraudulent payment transaction). In some non - limiting embodiments or aspects, the test dataset can include multiple data instances, where each data instance has a label indicating that the data instance is associated with a fraudulent data instance.
[0087] In some non - limiting embodiments or aspects, the test dataset includes multiple data records associated with multiple payment transactions, where each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non - fraudulent payment transaction. In some non - limiting embodiments or aspects, the test dataset includes multiple data records associated with multiple payment transactions, where each data record has a label indicating that the data record is associated with a fraudulent payment transaction.
[0088] In one example, the reinforcement learning system 102 can provide each of a plurality of data instances as input to a deep learning model and generate an output of the deep learning model based on the input. In some non-limiting embodiments or aspects, the output can include a label for the corresponding input. In some non-limiting embodiments or aspects, the resulting data set can include the output for each of the plurality of data instances based on a test data set, and the output of the deep learning model can include a label indicating whether the data instance is predicted to be a non-fraudulent data instance (e.g., a data instance associated with a non-fraudulent payment transaction) or a fraudulent data instance (e.g., a data instance associated with a fraudulent payment transaction). In some non-limiting embodiments or aspects, the plurality of data instances of the test data set are independent of the plurality of data instances of the initial training data set. For example, all of the data instances of the plurality of data instances of the test data set can be different from all of the data instances of the plurality of data instances of the initial training data set. In some examples, all of the data instances of the plurality of data instances of the test data set can be associated with a payment transaction that is different from the payment transaction associated with all of the data instances of the plurality of data instances of the initial training data set.
[0089] In some non-limiting embodiments or aspects, the reinforcement learning system 102 can determine whether the deep learning model correctly labels each data instance in the resulting data set. For example, the reinforcement learning system 102 can compare the labels of the plurality of data instances of the test data set with the labels of the plurality of data instances of the resulting data set. In such examples, if the reinforcement learning system 102 determines that the label of the data instance of the test data set matches the label of the corresponding data instance of the resulting data set, then the reinforcement learning system 102 can determine that the trained deep learning model correctly predicts the corresponding data instance of the resulting data set (e.g., correctly predicts the label of the corresponding data instance of the resulting data set). In some non-limiting embodiments or aspects, the reinforcement learning system 102 can determine the detection rate of the resulting data set by comparing the number of data instances of the resulting data set that are correctly predicted by the trained deep learning model with the total number of data instances of the resulting data set. In one example, the reinforcement learning system 102 can calculate the detection rate of the resulting data set as the ratio of the number of data instances of the resulting data set that are correctly labeled by the trained deep learning model to the total number of data instances of the resulting data set.
[0090] As Figure 3As shown, at step 310, process 300 includes generating a reward parameter based on a detection rate. For example, reinforcement learning system 102 can generate a reward parameter based on the detection rate of the obtained data set. In some non-limiting embodiments or aspects, reinforcement learning system 102 can train the RLA machine learning model based on the reward parameter. For example, reinforcement learning system 102 can use a reinforcement-based learning algorithm to train the RLA machine learning model based on the reward parameter. In one example, the reinforcement learning algorithm can be defined as follows:
[0091]
[0092] where r t+1 is the predicted fraudulent reward amount associated with a fraudulent payment transaction to be performed by an agent in the future, and γ is a weight parameter with a value range of 0 < γ < 1. In some non-limiting embodiments or aspects, reinforcement learning system 102 can update the RLA machine learning model (e.g., update the parameters of the RLA machine learning model, such as weight coefficients) to maximize the reward parameter.
[0093] In some non-limiting embodiments or aspects, reinforcement learning system 102 can perform actions in real time using a deep learning model (e.g., a trained deep learning model). For example, reinforcement learning system 102 can perform an action based on the label of an input provided by the deep learning model (e.g., a classification label), where the input includes a data instance associated with a payment transaction being performed in real time. In some non-limiting embodiments or aspects, reinforcement learning system 102 can perform a program associated with protecting the account of a user (e.g., a user associated with user device 108) based on the label of the input. For example, if the label of the input indicates that the program is necessary, reinforcement learning system 102 can perform the program associated with protecting the user's account. In such examples, if the label of the input indicates that the program is not necessary, reinforcement learning system 102 can refrain from performing the program associated with protecting the user's account. In some non-limiting embodiments or aspects, reinforcement learning system 102 can perform a fraud protection program based on the label of the input.
[0094] Now refer to Figures 4A - 4F , Figures 4A - 4FFIG. is a diagram of a non - limiting embodiment or aspect of embodiment 400 related to a process (e.g., process 300) for using reinforcement learning to improve the label accuracy of a machine - learning model. In some non - limiting embodiments or aspects, one or more steps of the process may be performed (e.g., fully, partially, etc.) by a reinforcement learning system 102 (e.g., one or more devices of the reinforcement learning system 102). In some non - limiting embodiments or aspects, one or more steps of the process may be performed (e.g., fully, partially, etc.) by another device or group of devices that is separate from or includes the reinforcement learning system 102 (e.g., one or more devices of the reinforcement learning system 102), the transaction service provider system 104 (e.g., one or more devices of the transaction service provider system 104), the issuer system 106 (e.g., one or more devices of the issuer system 106), and / or the user device 108.
[0095] As Figure 4A shown by reference numeral 405 in the figure, the reinforcement learning system 102 may receive an initial training data set that includes a plurality of data instances. In some non - limiting embodiments or aspects, the initial training data set may include a plurality of data records (e.g., shown as "R00", "R01", "R02", etc.), where each data instance has a label (e.g., shown as "F" for fraudulent and "NF" for non - fraudulent), and a first percentage of the plurality of data records is correctly labeled. As Figure 4B shown by reference numeral 410 in the figure, the reinforcement learning system 102 may provide the initial training data set as an input to an RLA machine - learning model to generate an enhanced training data set. In some non - limiting embodiments or aspects, the reinforcement learning system 102 may generate a second training data set that includes a greater percentage of correctly labeled data records compared to the initial data set.
[0096] As Figure 4C shown by reference numeral 415 in the figure, the reinforcement learning system 102 may use the enhanced training data set to train a deep - learning fraud - detection model. In some non - limiting embodiments or aspects, the reinforcement learning system 102 may train the deep - learning model by providing each data instance of the second training data set as an input and generating an output based on the input. In some non - limiting embodiments or aspects, the output of the deep - learning model may include a label indicating whether the data record is predicted to be a non - fraudulent data instance (e.g., a data instance associated with a non - fraudulent payment transaction) or a fraudulent data instance (e.g., a data instance associated with a fraudulent payment transaction).
[0097] For example, the reinforcement learning system 102 can train a deep learning model based on a loss function. In some non-limiting embodiments or aspects, the deep learning fraud detection model is configured to provide a prediction as to whether a payment transaction is fraudulent or non-fraudulent. As Figure 4D shown by reference numeral 420 in
[0098] As Figure 4E shown by reference numeral 425 in Figure 4F shown by reference numeral 430 in , the reinforcement learning system 102 can generate a reward parameter based on the detection rate of the resulting data set. As
[0099] Although the disclosed subject matter has been described in detail for purposes of illustration based on the currently considered most practical and preferred embodiments or aspects, it is to be understood that such details are for that purpose only and that the disclosed subject matter is not limited to the disclosed embodiments or aspects, but rather, is intended to cover modifications and equivalent arrangements within the spirit and scope of the appended claims. For example, it is to be understood that the presently disclosed subject matter contemplates, to the extent possible, that one or more features of any embodiment or aspect can be combined with one or more features of any other embodiment or aspect.
Claims
1. A computer-implemented method, comprising: Receiving, by at least one processor, an initial training data set, wherein the initial training data set includes a plurality of data instances, wherein each data instance has a label, and wherein a first percentage of the plurality of data instances is correctly labeled; Providing, by at least one processor, the initial training data set as an input to a reinforcement learning agent (RLA) machine learning model to generate a second training data set, wherein the second training data set includes the plurality of data instances, wherein a second percentage of the plurality of data instances is correctly labeled, and wherein the second percentage is greater than the first percentage; Training, by at least one processor, a deep learning model using the second training data set to provide a trained deep learning model; Testing, by at least one processor, the trained deep learning model using a test data set to generate a resulting data set, wherein the resulting data set has a detection rate, and wherein the detection rate is an indication of the number of data instances from the test data set correctly predicted by the trained deep learning model; And Generating, by at least one processor, a reward parameter based on the detection rate of the resulting data set generated using the trained deep learning model.
2. The computer-implemented method according to claim 1, further comprising: Training the RLA machine learning model using a reinforcement-based learning algorithm based on the reward parameter.
3. The computer-implemented method according to claim 2, wherein training the RLA machine learning model using the reinforcement-based learning algorithm comprises: Updating the RLA machine learning model to maximize the reward parameter generated by the trained deep learning model.
4. The computer-implemented method according to claim 1, wherein the plurality of data instances of the test data set are independent of the plurality of data instances of the initial training data set.
5. The computer-implemented method according to claim 1, wherein the RLA machine learning model includes a neural network machine learning model.
6. The computer-implemented method according to claim 1, wherein the plurality of data instances of the initial training data set includes a plurality of data records associated with a plurality of payment transactions; Wherein each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; and Wherein the deep learning model is a fraud detection deep learning model configured to provide a prediction as to whether a payment transaction is fraudulent or non-fraudulent.
7. The computer-implemented method according to claim 6, wherein the plurality of data instances of the test data set includes a second plurality of data records associated with a second plurality of payment transactions; Wherein a subset of the second plurality of data records has a label indicating whether each data record in the data record subset is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; And The detection rate is an indication of the number of data records in the data record subset from the test data set that are correctly predicted to be associated with a fraudulent payment transaction.
8. A system comprising: at least one processor configured to: receive an initial training data set, wherein the initial training data set includes a plurality of data instances, wherein each data instance has a label, and wherein a first percentage of the plurality of data instances is correctly labeled; provide the initial training data set as input to a reinforcement learning agent (RLA) machine learning model to generate a second training data set, wherein the second training data set includes the plurality of data instances, wherein a second percentage of the plurality of data instances is correctly labeled, and wherein the second percentage is greater than the first percentage; use the second training data set to train a deep learning model to provide a trained deep learning model; use a test data set to test the trained deep learning model to generate a resulting data set, wherein the resulting data set has a detection rate, and wherein the detection rate is an indication of the number of data instances from the test data set that are correctly predicted by the trained deep learning model; and generate a reward parameter based on the detection rate of the resulting data set.
9. The system of claim 8, wherein the at least one processor is further configured to: train the RLA machine learning model using a reinforcement-based learning algorithm based on the reward parameter.
10. The system of claim 9, wherein when training the RLA machine learning model using the reinforcement-based learning algorithm, the at least one processor is configured to: update the parameters of the RLA machine learning model to maximize the reward parameter generated by the trained deep learning model.
11. The system of claim 8, wherein the plurality of data instances of the test data set are independent of the plurality of data instances of the initial training data set.
12. The system of claim 8, wherein the RLA machine learning model includes a neural network machine learning model.
13. The system of claim 8, wherein the plurality of data instances of the initial training data set includes a plurality of data records associated with a plurality of payment transactions; wherein each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; and wherein the deep learning model is a fraud detection deep learning model configured to provide a prediction as to whether a payment transaction is fraudulent or non-fraudulent.
14. The system of claim 13, wherein the plurality of data instances of the test data set includes a second plurality of data records associated with a second plurality of payment transactions; wherein a data record subset of the second plurality of data records has a label indicating whether each data record in the data record subset is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; and The detection rate is an indication of the number of data records in the data record subset from the test data set that are correctly predicted to be associated with fraudulent payment transactions.
15. A computer program product comprising a non-transitory computer-readable medium including one or more instructions that, when executed by at least one processor, cause the at least one processor to: Receive an initial training data set, wherein the initial training data set includes a plurality of data instances, wherein each data instance has a label, and wherein a first percentage of the plurality of data instances are correctly labeled; Provide the initial training data set as input to a reinforcement learning agent (RLA) machine learning model to generate a second training data set, wherein the second training data set includes the plurality of data instances, wherein a second percentage of the plurality of data instances are correctly labeled, and wherein the second percentage is greater than the first percentage; Use the second training data set to train a deep learning model to provide a trained deep learning model; Use a test data set to test the trained deep learning model to generate a resulting data set, wherein the resulting data set has a detection rate, and wherein the detection rate is an indication of the number of data instances from the test data set that are correctly predicted by the trained deep learning model; and Generate a reward parameter based on the detection rate of the resulting data set.
16. The computer program product according to claim 15, wherein the one or more instructions further cause the at least one processor to: Train the RLA machine learning model using a reinforcement-based learning algorithm based on the reward parameter.
17. The computer program product according to claim 16, wherein the one or more instructions that cause the at least one processor to train the RLA machine learning model using the reinforcement-based learning algorithm cause the at least one processor to: Update the RLA machine learning model to maximize the reward parameter generated by the trained deep learning model.
18. The computer program product according to claim 15, wherein the plurality of data instances of the test data set are independent of the plurality of data instances of the initial training data set.
19. The computer program product according to claim 15, wherein the RLA machine learning model includes a neural network machine learning model.
20. The computer program product according to claim 15, wherein the plurality of data instances of the initial training data set include a plurality of data records associated with a plurality of payment transactions; wherein each data record has a label indicating whether the data record is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; wherein the deep learning model is a fraud detection deep learning model configured to provide a prediction as to whether a payment transaction is fraudulent or non-fraudulent; wherein the plurality of data instances of the test data set include a second plurality of data records associated with a second plurality of payment transactions; A data record subset of the second plurality of data records has a label indicating whether each data record in the data record subset is associated with a fraudulent payment transaction or a non-fraudulent payment transaction; and wherein the detection rate is an indication of the number of data records in the data record subset from the test data set that are correctly predicted to be associated with fraudulent payment transactions.
Citation Information
Patent Citations
Method and device for predicting sample label based on reinforcement learning model
CN110263979A
Methods, systems, and computer program products for fraud prevention using deep learning and survival models
CN114387074A
Systems, methods, and computer program products for time-based aggregate learning using supervised and unsupervised machine learning models
CN116802648A
Systems and methods for an adaptive sampling of unlabeled data samples for constructing an informative training data corpus that improves a training and predictive accuracy of a machine learning model
US11496501B1
Weakly supervised multi-task learning for concept-based explainability
US20220114345A1