Systems, methods, and computer program products for state correction of artificial intelligence models
Patent Information
- Application Number
- CN202480088758.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-27
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]然而,由于机器学习模型在有偏数据上训练,实时数据分布与训练数据分布不同,和/或在生产环境中对机器学习模型进行目标攻击的情况,部署在系统中的机器学习模型可能处于关于准确性的低性能风险中
Smart Images

Figure CN122826580A_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to artificial intelligence models, and in some non-limiting embodiments or aspects, to systems, methods, and computer program products for state correction of stateful machine learning (ML) models. Background Technology
[0002] Artificial intelligence (AI) can refer to the general ability of computers to simulate human thinking and perform tasks in real-world environments. Machine learning, as a subset of AI, is a field of computer science that uses statistical techniques to enable computer systems to learn tasks (e.g., progressively improve performance) using data without explicitly programming the computer system to perform the task. In some cases, machine learning models can be developed for datasets that can perform tasks related to said dataset (e.g., tasks associated with prediction).
[0003] In some cases, machine learning models, such as predictive machine learning models, can be used to make predictions related to risk or opportunity based on large amounts of data (e.g., large-scale datasets). Predictive machine learning models can be used to analyze the relationship between a unit's performance and one or more known features of the unit, based on a large dataset associated with that unit. The purpose of a predictive machine learning model can be to assess the probability that similar units will exhibit the same or similar performance as the unit. To generate a predictive machine learning model, the large-scale dataset can be segmented so that the predictive machine learning model can be trained on appropriate data.
[0004] However, machine learning models deployed in a system may be at low performance risk regarding accuracy due to factors such as training machine learning models on biased data, real-time data distribution differing from training data distribution, and / or targeted attacks on machine learning models in production environments. Attempting to address such situations may require disrupting the availability of the machine learning model and / or the training environment. Therefore, attempts may include dedicating significant time and resources (e.g., network and computational resources for the training and / or validation processes) to determining whether the machine learning model is functioning as expected. Summary of the Invention
[0005] Therefore, improved systems, methods, and computer program products for state correction of stateful machine learning (ML) models are provided.
[0006] According to a non-limiting embodiment or aspect, a system is provided, the system including at least one processor configured to: receive input for a stateful machine learning (ML) model; generate an output of the stateful ML model based on the input; determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; assign an initial state of the stateful ML model as an active state of the stateful ML model based on the determination that the output of the stateful ML model corresponds to the baseline truth value associated with the input; and assign a corrected state of the stateful ML model as the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input.
[0007] In some non-limiting embodiments or aspects, when the corrected state of the stateful ML model is assigned to the active state of the stateful ML model, the at least one processor is configured to: compute a hash key using a hash function and the initial state of the stateful ML model relative to the input; determine the corrected state of the stateful ML model using a state map, the hash key, and the benchmark truth value associated with the input, based on the hash key computed using the hash function and the state of the stateful ML model relative to the input; and assign the corrected state of the stateful ML model to the active state of the stateful ML model based on the corrected state of the stateful ML model determined using the state map, the hash key, and the benchmark truth value associated with the input.
[0008] In some non-limiting embodiments or aspects, the at least one processor is further configured to generate the state map based on multiple historical outputs of the stateful ML model. In some non-limiting embodiments or aspects, when generating the state map, the at least one processor is configured to: determine multiple values of states of the stateful ML model associated with correct predictions of the stateful ML model; generate clusters based on the multiple values of the states of the stateful ML model; and determine the centroids of the clusters, wherein the centroids of the clusters include values of states associated with benchmark ground truth values for a classification task of the stateful ML model. In some non-limiting embodiments or aspects, the stateful ML model is a recurrent neural network (RNN). In some non-limiting embodiments or aspects, the at least one processor is further configured to receive the benchmark ground truth value associated with the input after generating the output of the stateful ML model based on the input. In some non-limiting embodiments or aspects, the at least one processor is further configured to update the initial state of the stateful ML model relative to the input based on the received input.
[0009] According to a non-limiting embodiment or aspect, a computer-implemented method is provided, the computer-implemented method comprising: receiving input for a stateful machine learning (ML) model; generating an output of the stateful ML model based on the input; determining whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; abandoning the assignment of an initial state of the stateful ML model relative to the input as the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input; and performing an action associated with assigning the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input.
[0010] In some non-limiting embodiments or aspects, performing the action associated with assigning the active state of the stateful ML model includes: calculating a hash key using a hash function and the state of the stateful ML model relative to the input, based on abandoning the assignment of the initial state of the stateful ML model as the active state of the stateful ML model; determining a corrected state of the stateful ML model using a state map, the hash key, and the baseline truth value associated with the input, based on the hash key calculated using the hash function and the state of the stateful ML model relative to the input; and assigning the corrected state of the stateful ML model as the active state of the stateful ML model based on the corrected state determined using the state map, the hash key, and the baseline truth value associated with the input.
[0011] In some non-limiting embodiments or aspects, the computer-implemented method further includes generating the state map based on multiple historical outputs of the stateful ML model. In some non-limiting embodiments or aspects, generating the state map includes: determining multiple values of states of the stateful ML model associated with correct predictions of the stateful ML model; generating clusters based on the multiple values of the states of the stateful ML model; and determining the centroids of the clusters, wherein the centroids of the clusters include values of states associated with benchmark ground truth values for a classification task of the stateful ML model.
[0012] In some non-limiting embodiments or aspects, the stateful ML model is a recurrent neural network (RNN). In some non-limiting embodiments or aspects, the computer-implemented method further includes receiving the baseline truth value associated with the input after generating the output of the stateful ML model based on the input. In some non-limiting embodiments or aspects, the computer-implemented method further includes updating the initial state of the stateful ML model relative to the input based on the received input.
[0013] According to a non-limiting embodiment or aspect, a computer program product is provided, the computer program product comprising at least one non-transient computer-readable medium, the at least one non-transient computer-readable medium comprising program instructions, which, when executed by at least one processor, cause the at least one processor to: receive input for a stateful machine learning (ML) model; generate an output of the stateful ML model based on the input; determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; assign an initial state of the stateful ML model as an active state of the stateful ML model based on the determination that the output of the stateful ML model corresponds to the baseline truth value associated with the input; and assign a corrected state of the stateful ML model as the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input.
[0014] In some non-limiting embodiments or aspects, the program instructions that cause the at least one processor to assign the corrected state of the stateful ML model to the active state of the stateful ML model cause the at least one processor to: compute a hash key using a hash function and the initial state of the stateful ML model relative to the input; determine the corrected state of the stateful ML model using a state map, the hash key, and the benchmark truth value associated with the input based on the hash key computed using the hash function and the stateful ML model relative to the input; and assign the corrected state of the stateful ML model to the active state of the stateful ML model based on the corrected state of the stateful ML model determined using the state map, the hash key, and the benchmark truth value associated with the input.
[0015] In some non-limiting embodiments or aspects, the program instructions further cause the at least one processor to generate the state map based on multiple historical outputs of the stateful ML model. In some non-limiting embodiments or aspects, the program instructions causing the at least one processor to generate the state map further cause the at least one processor to: determine multiple values of states of the stateful ML model associated with correct predictions of the stateful ML model; generate clusters based on the multiple values of the states of the stateful ML model; and determine the centroids of the clusters, wherein the centroids of the clusters include values of states associated with benchmark ground truth values for a classification task of the stateful ML model.
[0016] In some non-limiting embodiments or aspects, the stateful ML model is a recurrent neural network (RNN). In some non-limiting embodiments or aspects, the program instructions further cause the at least one processor to receive the baseline truth value associated with the input after generating the output of the stateful ML model based on the input. In some non-limiting embodiments or aspects, the program instructions further cause the at least one processor to update the initial state of the stateful ML model relative to the input based on the received input.
[0017] Other non-limiting embodiments or aspects are set forth in the following numbered clauses: Clause 1: A system comprising at least one processor configured to: receive input for a stateful machine learning (ML) model; generate an output of the stateful ML model based on the input; determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; assign an initial state of the stateful ML model as an active state of the stateful ML model based on the determination that the output of the stateful ML model corresponds to the baseline truth value associated with the input; and assign a corrected state of the stateful ML model as the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input.
[0018] Clause 2: The system according to Clause 1, wherein when the corrected state of the stateful ML model is assigned to the active state of the stateful ML model, the at least one processor is configured to: compute a hash key using a hash function and the initial state of the stateful ML model relative to the input; determine the corrected state of the stateful ML model using a state map, the hash key, and the benchmark truth value associated with the input, based on the hash key computed using the hash function and the state of the stateful ML model relative to the input; and assign the corrected state of the stateful ML model to the active state of the stateful ML model based on the corrected state of the stateful ML model determined using the state map, the hash key, and the benchmark truth value associated with the input.
[0019] Clause 3: In the system described in Clause 1 or 2, wherein the at least one processor is further configured to generate the state mapping based on multiple historical outputs of the stateful ML model.
[0020] Clause 4: A system according to any one of Clauses 1 to 3, wherein when generating the state mapping, the at least one processor is configured to: determine a plurality of values of states associated with correct predictions of the stateful ML model; generate clusters based on the plurality of values of states of the stateful ML model; and determine the centroids of the clusters, wherein the centroids of the clusters include values of states associated with benchmark ground truth values for a classification task of the stateful ML model.
[0021] Clause 5: The system according to any one of Clauses 1 to 4, wherein the stateful ML model is a recurrent neural network (RNN).
[0022] Clause 6: A system according to any one of Clauses 1 to 5, wherein the at least one processor is further configured to: after generating the output of the stateful ML model based on the input, receive the baseline truth value associated with the input.
[0023] Clause 7: A system according to any one of Clauses 1 to 6, wherein the at least one processor is further configured to: update the initial state of the stateful ML model relative to the input based on the received input.
[0024] Clause 8: A computer-implemented method comprising: receiving input for a stateful machine learning (ML) model; generating an output of the stateful ML model based on the input; determining whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; abandoning the assignment of an initial state of the stateful ML model relative to the input as the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input; and performing an action associated with assigning the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input.
[0025] Clause 9: A computer-implemented method according to Clause 8, wherein performing the actions associated with assigning the active state of the stateful ML model includes: calculating a hash key using a hash function and the state of the stateful ML model relative to the input, based on abandoning the assignment of the initial state of the stateful ML model as the active state of the stateful ML model; determining a corrected state of the stateful ML model using a state map, the hash key, and the baseline truth value associated with the input, based on the hash key calculated using the hash function and the state of the stateful ML model relative to the input; and assigning the corrected state of the stateful ML model as the active state of the stateful ML model based on the corrected state determined using the state map, the hash key, and the baseline truth value associated with the input.
[0026] Clause 10: The computer-implemented method according to Clause 8 or 9 further includes: generating the state mapping based on multiple historical outputs of the stateful ML model.
[0027] Clause 11: A computer-implemented method according to any one of Clauses 8 to 10, wherein generating the state mapping comprises: determining a plurality of values of states of the stateful ML model associated with correct predictions of the stateful ML model; generating clusters based on the plurality of values of states of the stateful ML model; and determining the centroids of the clusters, wherein the centroids of the clusters comprise values of states associated with benchmark ground truth values for a classification task of the stateful ML model.
[0028] Clause 12: A computer-implemented method according to any one of Clauses 8 to 11, wherein the stateful ML model is a recurrent neural network (RNN).
[0029] Clause 13: The computer-implemented method according to any one of Clauses 8 to 12 further includes: after generating the output of the stateful ML model based on the input, receiving the baseline truth value associated with the input.
[0030] Clause 14: The computer-implemented method according to any one of Clauses 8 to 13 further includes: updating the initial state of the stateful ML model relative to the input based on the received input.
[0031] Clause 15: A computer program product comprising at least one non-transient computer-readable medium, the at least one non-transient computer-readable medium comprising program instructions that, when executed by at least one processor, cause the at least one processor to: receive input for a stateful machine learning (ML) model; generate an output of the stateful ML model based on the input; determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; assign an initial state of the stateful ML model as an active state of the stateful ML model based on the determination that the output of the stateful ML model corresponds to the baseline truth value associated with the input; and assign a corrected state of the stateful ML model as the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input.
[0032] Clause 16: A computer program product according to Clause 15, wherein the program instructions that cause the at least one processor to assign the corrected state of the stateful ML model to the active state of the stateful ML model cause the at least one processor to: compute a hash key using a hash function and the initial state of the stateful ML model relative to the input; determine the corrected state of the stateful ML model using a state map, the hash key, and the benchmark truth value associated with the input based on the hash key computed using the hash function and the stateful ML model relative to the input; and assign the corrected state of the stateful ML model to the active state of the stateful ML model based on the corrected state of the stateful ML model determined using the state map, the hash key, and the benchmark truth value associated with the input.
[0033] Clause 17: A computer program product according to Clause 15 or 16, wherein the program instructions further cause the at least one processor to generate the state mapping based on a plurality of historical outputs of the stateful ML model.
[0034] Clause 18: A computer program product according to any one of Clauses 15 to 17, wherein the program instructions that cause the at least one processor to generate the state mapping further cause the at least one processor to: determine a plurality of values of states of the stateful ML model associated with a correct prediction of the stateful ML model; generate clusters based on the plurality of values of states of the stateful ML model; and determine the centroids of the clusters, wherein the centroids of the clusters include values of states associated with a benchmark truth value for a classification task of the stateful ML model.
[0035] Clause 19: A computer program product pursuant to any one of Clauses 15 to 18, wherein the stateful ML model is a recurrent neural network (RNN).
[0036] Clause 20: A computer program product according to any one of Clauses 15 to 19, wherein the program instructions further cause the at least one processor to: receive the baseline truth value associated with the input after generating the output of the stateful ML model based on the input.
[0037] Clause 21: A computer program product according to any one of Clauses 15 to 20, wherein the program instructions further cause the at least one processor to update the initial state of the stateful ML model relative to the input based on receiving the input.
[0038] These and other features and characteristics of this disclosure, as well as the operational methods and manufacturing economies of combinations of related structural elements and parts, will become more apparent when considered with reference to the accompanying drawings, all of which form part of this specification, wherein like reference numerals denote corresponding parts in the figures. However, it should be clearly understood that the drawings are for illustrative and descriptive purposes only and are not intended to be construed as limiting the scope of this disclosure. Attached Figure Description
[0039] Further advantages and details are explained in more detail below with reference to the non-limiting exemplary embodiments shown in the illustrative drawings, in which: Figure 1 This is a schematic diagram of a system for state correction of a stateful machine learning (ML) model according to a non-limiting embodiment or aspect; Figure 2 This is a flowchart of a method for state correction of a stateful ML model according to some non-limiting embodiments or aspects; Figure 3A to 3G This is a schematic diagram of an exemplary embodiment of the present disclosure for state correction of a stateful ML model according to some non-limiting embodiments or aspects; Figure 4 These are diagrams illustrating exemplary environments in which the systems, methods, and / or computer program products described herein may be implemented, based on some non-limiting embodiments or aspects; and Figure 5 Based on some non-limiting embodiments or aspects Figure 1 and / or Figure 4 A schematic diagram of example components of one or more devices. Detailed Implementation
[0040] For the purposes of the following description, the terms “end,” “upper,” “lower,” “right,” “left,” “vertical,” “horizontal,” “top,” “bottom,” “lateral,” “longitudinal,” and their derivatives should be associated with the orientation of the embodiments in the accompanying drawings. However, it should be understood that various alternative variations and sequences of steps may be employed in this disclosure, except where explicitly specified otherwise. It should also be understood that the specific apparatus and processes shown in the drawings and described in the following specification are merely exemplary and non-limiting embodiments or aspects of the disclosed subject matter. Therefore, specific dimensions and other physical characteristics relating to the embodiments or aspects disclosed herein should not be considered limiting.
[0041] This document describes some non-limiting embodiments or aspects in conjunction with thresholds. As used herein, satisfying a threshold can refer to a value greater than a threshold, more than a threshold, higher than a threshold, greater than or equal to a threshold, less than a threshold, less than a threshold, lower than a threshold, less than or equal to a threshold, equal to a threshold, etc.
[0042] The terms "aspect," "component," "element," "structure," "action," "step," "function," and "instruction" used herein should not be construed as critical or essential unless explicitly stated otherwise. Furthermore, as used herein, the articles "a" and "an" are intended to include one or more items and may be used interchangeably with "one or more" and "at least one." Additionally, as used herein, the term "set" is intended to include one or more items (e.g., related items, unrelated items, a combination of related and unrelated items, etc.) and may be used interchangeably with "one or more" or "at least one." The term "a" or similar language is used where only one item is intended. Moreover, as used herein, the terms "has," "have," and "having" are intended to be open-ended terms. Additionally, unless explicitly stated otherwise, the phrase "based on" is intended to mean "at least partially based on." Furthermore, a reference to an action based on a condition may indicate that the action is "in response to" the condition. For example, in some non-limiting embodiments or aspects, the phrases “based on” and “in response to” may refer to conditions for automatically triggering actions (e.g., specific operations of electronic devices such as computing devices, processors, etc.).
[0043] As used herein, the term "communication" can refer to the receiving, accepting, sending, transmitting, providing, etc., of data (e.g., information, signals, messages, instructions, commands, etc.). For one unit (e.g., apparatus, system, component of an apparatus or system, combination thereof, etc.) to communicate with another unit means that the first unit is able to receive information directly or indirectly from and / or send information to the other unit. This can refer to a direct or indirect connection that is inherently wired and / or wireless (e.g., a direct communication connection, an indirect communication connection, etc.). Furthermore, the two units can communicate with each other even though the transmitted information may be modified, processed, relayed, and / or routed between them. For example, the first unit can communicate with the second unit even if it passively receives information and does not actively send information to the second unit. As another example, the first unit can communicate with the second unit if at least one intermediate unit processes information received from the first unit and transmits the processed information to the second unit. In some non-limiting embodiments or aspects, a message can refer to a network packet (e.g., a data packet, etc.) that includes data. It should be understood that many other arrangements are possible.
[0044] As used herein, the term "computing device" can refer to one or more electronic devices configured to process data. In some examples, a computing device may include essential components for receiving, processing, and outputting data, such as a processor, display, memory, input device, network interface, etc. A computing device can be a mobile device. As examples, a mobile device may include a cellular phone (e.g., a smartphone or standard cellular phone), a portable computer, a wearable device (e.g., a watch, glasses, lenses, clothing, etc.), a personal digital assistant (PDA), and / or other similar devices. A computing device may also be a desktop computer or other form of non-mobile computer.
[0045] As used herein, the term "server" may refer to or include one or more computing devices operated by or facilitating communication and processing among multiple parties in a network environment such as the Internet, but it should be understood that communication may be facilitated through one or more public or private network environments, and various other arrangements may be possible. Furthermore, multiple computing devices (e.g., servers, point-of-sale (POS) devices, mobile devices, etc.) communicating directly or indirectly in a network environment may constitute a "system".
[0046] As used herein, the term "system" may refer to one or more computing devices or a combination of computing devices (e.g., processor, server, client device, software application, components of such computing devices, etc.). As used herein, references to "device," "server," "processor," etc., may refer to a previously described device, server, or processor, different devices, servers, or processors, and / or combinations of devices, servers, and / or processors, described as performing a preceding step or function. For example, as used in the specification and claims, a first device, first server, or first processor described as performing a first step or a first function may refer to the same or different devices, servers, or processors described as performing a second step or a second function.
[0047] As used herein, the term "acquiring party" can refer to an entity authorized and approved by a transaction service provider to initiate transactions (e.g., payment transactions) involving payment devices associated with the transaction service provider. As used herein, the term "acquiring party system" can also refer to one or more computer systems, computer devices, etc., operated by or on behalf of the acquiring party. Transactions that an acquiring party can initiate can include payment transactions (e.g., purchases, Original Credit Transactions (OCT), Account Funds Transactions (AFT), etc.). In some non-limiting embodiments or aspects, the transaction service provider may authorize the acquiring party to assign merchants or service providers to initiate transactions involving payment devices associated with the transaction service provider. The acquiring party may contract with payment service providers to enable them to sponsor merchants. The acquiring party may monitor the compliance of payment service providers in accordance with transaction service provider regulations. The acquiring party may conduct due diligence on payment service providers and ensure that appropriate due diligence occurs before contracting with sponsored merchants. The acquiring party may be responsible for all transaction service provider programs operated or sponsored by the acquiring party. The acquiring party may be responsible for the actions of its payment service provider, merchants sponsored by the payment service provider, etc. In some non-limiting embodiments or aspects, the acquiring party may be a financial institution, such as a bank.
[0048] As used herein, the terms “issuer,” “issuer institution,” “issuer bank,” or “payment device issuer” can refer to one or more entities that provide accounts to individuals (e.g., users, customers, etc.) for making payment transactions such as credit payment transactions and / or debit payment transactions. For example, an issuer institution may provide a customer with an account identifier, such as a primary account number (PAN), that uniquely identifies one or more accounts associated with said customer. In some non-limiting embodiments or aspects, an issuer may be associated with a bank identification number (BIN) that uniquely identifies an issuer institution. As used herein, “issuer system” can refer to one or more computer systems operated by or on behalf of an issuer, such as a server executing one or more software applications. For example, an issuer system may include one or more authorization servers for authorizing transactions.
[0049] As used herein, the term "merchant" can refer to one or more entities (e.g., operators of retail businesses) that provide goods and / or services and / or access to goods and / or services to users (e.g., customers, consumers) based on transactions such as payment transactions. As used herein, "merchant system" can refer to one or more computer systems operated by or on behalf of a merchant, such as servers executing one or more software applications. As used herein, the term "product" can refer to one or more goods and / or services offered by a merchant.
[0050] As used herein, the term "transaction service provider" can refer to an entity that receives transaction authorization requests from merchants or other entities and, in some cases, provides payment guarantees through an agreement between the transaction service provider and the issuing entity. For example, a transaction service provider may include payment networks such as Visa®, MasterCard®, American Express®, or any other entity that processes transactions. As used herein, the term "transaction service provider system" can refer to one or more computer systems operated by or on behalf of a transaction service provider, such as a transaction service provider system executing one or more software applications. A transaction service provider system may include one or more processors and, in some non-limiting embodiments or aspects, may be operated by or on behalf of a transaction service provider.
[0051] Non-limiting embodiments or aspects of this disclosure relate to methods, systems, and computer program products for state correction of stateful machine learning (ML) models. In some non-limiting embodiments or aspects, a state management system may be configured to: receive input for a stateful ML model; generate an output of the stateful ML model based on the input; determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; assign an initial state of the stateful ML model as an active state of the stateful ML model based on the determination that the output of the stateful ML model corresponds to the baseline truth value associated with the input; and assign a corrected state of the stateful ML model as the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input.
[0052] In some non-limiting embodiments or aspects, when assigning the corrected state of the stateful ML model to the active state of the stateful ML model, the state management system may use a hash function and the initial state of the stateful ML model relative to the input to calculate a hash key; based on calculating the hash key using the hash function and the state of the stateful ML model relative to the input, the corrected state of the stateful ML model is determined using a state map, the hash key, and the baseline truth value associated with the input; and based on determining the corrected state of the stateful ML model using the state map, the hash key, and the baseline truth value associated with the input, the corrected state of the stateful ML model is assigned to the active state of the stateful ML model. In some non-limiting embodiments or aspects, the state management system may be further configured to generate the state map based on multiple historical outputs of the stateful ML model.
[0053] In some non-limiting embodiments or aspects, when generating the state mapping, the state management system may determine multiple values of states associated with correct predictions of the stateful ML model; generate clusters based on the multiple values of the states of the stateful ML model; and determine the centroids of the clusters, wherein the centroids of the clusters include values of states associated with benchmark ground truth values for a classification task of the stateful ML model. In some non-limiting embodiments or aspects, the stateful ML model is a recurrent neural network (RNN).
[0054] In some non-limiting embodiments or aspects, the state management system may be further configured to receive the baseline truth value associated with the input after generating the output of the stateful ML model based on the input. In some non-limiting embodiments or aspects, the state management system may be further configured to update the initial state of the stateful ML model relative to the input based on the received input.
[0055] In this way, a state management system can accurately determine whether a stateful ML model is providing accurate results with reduced resources (e.g., network and / or computational resources). Additionally, the state management system can determine when a stateful ML model provides inaccurate results and provide compensation for inaccurate results based on corrective states using the stateful ML model. In some examples, the state management system can reduce and / or prevent the need for training and / or validating (e.g., retraining and / or revalidating) the stateful ML model. Furthermore, the state management system can reduce the need for access to stateful ML in environments using the model, and thus reduce the number of vulnerabilities that lead to low performance of stateful ML models deployed in systems that use them for operations.
[0056] For illustrative purposes, while the subject matter of this disclosure is described in the following description with respect to systems, methods, and computer program products for state correction of stateful ML models, those skilled in the art will recognize that the disclosed subject matter is not limited to the non-limiting embodiments or aspects disclosed herein. For example, the systems, methods, and computer program products described herein can be used with a variety of settings, such as determining the performance of a stateful ML model and / or determining whether the state of the stateful ML model should be corrected based on performance in any suitable setting (e.g., an online setting (e.g., a production setting)) and for any suitable purpose (e.g., prediction, regression, classification, fraud prevention, transaction authorization, user authentication, user identification, feature selection, recommendation, etc.).
[0057] Now for reference Figure 1 , Figure 1 This is a schematic diagram of a system 100 for state correction of a stateful ML model, according to some non-limiting embodiments or aspects. For example... Figure 1 As shown, system 100 may include a state management system 102, a machine learning (ML) model repository 104, a user device 106, and a communication network 108. The state management system 102, the ML model repository 104, and / or the user device 106 may be interconnected via wired connections, wireless connections, or a combination of wired and wireless connections (e.g., establishing connections for communication).
[0058] The state management system 102 may include one or more devices configured to communicate with the ML model repository 104 and / or user device 106 via a communication network 108. For example, the state management system 102 may include computing devices, such as servers (e.g., a single server), server groups, and / or other similar devices. In some non-limiting embodiments or aspects, the state management system 102 may include a processor and / or memory as described herein. In some non-limiting embodiments or aspects, the state management system 102 may include one or more software instructions (e.g., one or more software applications) that execute on a server (e.g., a single server), server group, computing device (e.g., a single computing device), computing device group, and / or other similar devices. In some non-limiting embodiments or aspects, the state management system 102 may be configured to perform one or more steps of the methods described herein. In some non-limiting embodiments or aspects, the state management system 102 may be configured to communicate with a data storage device (e.g., the ML model repository 104). In some non-limiting embodiments or aspects, the state management system 102 may communicate with the ML model repository 104 and / or the user device 106, such that the state management system 102 is separate from the ML model repository 104 and / or the user device 106. In some non-limiting embodiments or aspects, the user device 106 and / or the ML model repository 104 may be implemented by the state management system 102 (e.g., may be part of the state management system).
[0059] In some non-limiting embodiments or aspects, the state management system 102 may generate (e.g., train, validate, retrain, etc.), store, and / or implement one or more machine learning models (e.g., operate one or more machine learning models, provide input to one or more machine learning models, and / or provide output from one or more machine learning models, etc.). For example, the state management system 102 may generate one or more machine learning models by fitting (e.g., validating, testing, etc.) one or more machine learning models against data used for training (e.g., training data). In some non-limiting embodiments or aspects, the state management system 102 may generate, store, and / or implement one or more machine learning models for a production environment (e.g., runtime environment, real-time environment, etc.) for providing inference (e.g., secure inference) based on data input in a live (e.g., real-time) situation. Alternatively or additionally, the state management system 102 may generate, store, and / or implement one or more machine learning models for a non-production environment (e.g., offline environment, training environment, etc.) for providing inference based on data input in a non-live situation. In some non-limiting embodiments or aspects, the state management system 102 may communicate with a data storage device (ML model repository 104), which may be local or remote to the state management system 102.
[0060] The ML model repository 104 may include one or more devices capable of communicating with the state management system 102 and / or user device 106 via a communication network 108. For example, the ML model repository 104 may include computing devices such as servers (e.g., a single server), server groups, and / or other similar devices. In some non-limiting embodiments or aspects, the ML model repository 104 may receive, store, and / or provide (e.g., send) one or more machine learning models. In some non-limiting embodiments or aspects, the ML model repository 104 may be associated with one or more computing devices that provide an interface, allowing users (e.g., administrators, users using service accounts, etc.) to interact with the ML model repository 104 via said one or more computing devices. The ML model repository 104 may communicate with the state management system 102 and / or user device 106, such that the ML model repository 104 is separate from the state management system 102 and / or user device 106. Alternatively, in some non-limiting embodiments or aspects, the ML model repository 104 may be implemented by the state management system 102 and / or the user device 106 (e.g., it may be part of the state management system and / or the user device).
[0061] User device 106 may include a computing device configured to communicate with state management system 102 and / or ML model repository 104 via communication network 108. For example, user device 106 may include computing devices such as desktop computers, portable computers (e.g., tablet computers, laptop computers, etc.), mobile devices (e.g., cellular phones, smartphones, personal digital assistants, wearable devices, etc.) and / or other similar devices. In some non-limiting embodiments or aspects, user device 106 may be associated with a user (e.g., an individual operating user device 106).
[0062] The communication network 108 may include one or more wired and / or wireless networks. For example, the communication network 108 may include cellular networks (e.g., Long Term Evolution (LTE) networks, third-generation (3G) networks, fourth-generation (4G) networks, fifth-generation (5G) networks, Code Division Multiple Access (CDMA) networks, etc.), Public Land Mobile Networks (PLMN), Local Area Networks (LAN), Wide Area Networks (WAN), Metropolitan Area Networks (MAN), Telephone Networks (e.g., Public Switched Telephone Network (PSTN), etc.), Private Networks, Ad Hoc Networks, Intranets, the Internet, Fiber-based Networks, Cloud Computing Networks, etc., and / or combinations of some or all of these or other types of networks.
[0063] supply Figure 1 The number and arrangement of the devices and networks shown are for illustrative purposes only. It is possible that... Figure 1 The devices and / or networks shown are those that, compared to additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks arranged differently. Furthermore, Figure 1 The two or more devices shown can be implemented within a single device, or Figure 1 The single device shown may be implemented as multiple distributed devices. Alternatively, a group of devices (e.g., one or more devices) of system 100 may perform one or more functions described as being performed by another group of devices of system 100.
[0064] Now for reference Figure 2 , Figure 2This is a flowchart of a non-limiting embodiment or aspect of a process 200 for state correction of a stateful ML model. In some non-limiting embodiments or aspects, one or more steps of process 200 may be performed (e.g., wholly, partially, etc.) by a state management system 102 (e.g., one or more means of state management system 102). In some non-limiting embodiments or aspects, one or more steps of process 200 may be performed (e.g., wholly, partially, etc.) by another means or group of means separate from or including the state management system 102 (e.g., one or more means of state management system 102), the ML model repository 104 (e.g., one or more means of ML model repository 104), and / or the user device 106.
[0065] like Figure 2 As shown, at step 202, process 200 includes receiving input for a stateful ML model. For example, the state management system 102 may receive input for a stateful ML model from an ML model repository 104, a user device 106, and / or another system or device. In some non-limiting embodiments or aspects, the state management system 102 may receive input for a stateful ML model from a system or device as part of an inference request. For example, the stateful ML model may be provided in an online setup (e.g., a production environment), and the state management system 102 may receive input for a stateful ML model as part of a real-time inference request in the online setup. In some non-limiting embodiments or aspects, the real-time inference request may be based on a request to determine the characteristics of a transaction (e.g., an electronic payment transaction). For example, the real-time inference request may be based on a request to determine whether a transaction is a fraudulent transaction.
[0066] In some non-limiting embodiments or aspects, a stateful ML model may include an ML model that identifies dependencies between inferences. For example, a stateful ML model may maintain a state (e.g., a hidden state) between successive inferences such that a later inference depends on the result of a previous inference, which may be provided as a state. In some non-limiting embodiments or aspects, the state of a stateful ML model may be referred to as an active state, and the active state may include a value of a state to be passed between successive inferences (e.g., passed as feedback to the stateful ML model). In some non-limiting embodiments or aspects, with respect to an initial inference, the state of the stateful model may include an initial state associated with an initial output (e.g., the result of the initial inference) (e.g., equal to the initial output, based on the initial output, based on a value within a certain amount of the initial output, etc.). In some non-limiting embodiments or aspects, a stateful ML model may include a neural network ML model. In some examples, a stateful ML model may include a recurrent neural network (RNN) ML model, such as a long short-term memory (LSTM) network ML model, a gated recurrent unit (GRU) ML model, a hidden Markov model (HMM) ML model, or any combination thereof. In some non-limiting embodiments or aspects, a stateful ML model may include a neural network ML model with a large number of nodes. In this example, one or more transformer machine learning models may include neural network ML models with at least 5 nodes, at least 10 nodes, at least 20 nodes, at least 50 nodes, at least 100 nodes, at least 1,000 nodes, etc. (e.g., neural network ML models with one or more layers, such as one or more input layers, one or more hidden layers, and / or one or more output layers). In some non-limiting embodiments or aspects, the configuration of one or more layers of the neural network ML model may be based on the inputs that the neural network ML model is designed to receive and / or the outputs that the neural network ML model is designed to provide.
[0067] In some non-limiting embodiments or aspects, a stateful ML model may include a classification ML model configured to receive input and provide output based on the input, wherein the output includes a prediction of classification (e.g., a predicted score) for a classification (e.g., a label for classification).
[0068] In some non-limiting embodiments or aspects, the state management system 102 may receive input for a stateful ML model, the input including a dataset, and the dataset may include multiple data points (e.g., data instances, data examples, etc.). In some non-limiting embodiments or aspects, the dataset may be associated with one or more entities (e.g., one or more users, one or more account holders, one or more merchants, one or more issuers, etc.). For example, multiple data points may represent multiple transactions (e.g., electronic payment transactions) involving entities (e.g., transactions performed by entities). In some examples, multiple data points may include a large number of data points, such as at least 50 data points, at least 100 data points, at least 500 data points, at least 1,000 data points, at least 5,000 data points, at least 10,000 data points, at least 25,000 data points, at least 50,000 data points, at least 100,000 data points, at least 1,000,000 data points, etc. In some examples, multiple data points may be associated with multiple features in the form of multiple transaction parameters (e.g., transaction variables). In some non-limiting embodiments or aspects, the stateful ML model may have already been trained and / or validated on multiple data points.
[0069] In some non-limiting embodiments or aspects, each data point may include transaction data associated with a transaction. In some non-limiting embodiments or aspects, the transaction data may include multiple transaction parameters associated with an electronic payment transaction. In some non-limiting embodiments or aspects, multiple features may represent multiple transaction parameters. In some non-limiting embodiments or aspects, multiple transaction parameters may include e-wallet card data associated with an e-card (e.g., e-credit card, e-debit card, e-loyalty card, etc.), decision data associated with a decision (e.g., a decision to approve or reject a transaction authorization request), authorization data associated with an authorization response (e.g., approved spending limit, approved transaction value, etc.), master account (PAN), authorization code (e.g., personal identification number (PIN), etc.), data associated with the transaction amount (e.g., approved limit, transaction value, etc.), data associated with the transaction date and time, data associated with the currency exchange rate, data associated with the merchant type (e.g., a merchant category code indicating the type of goods such as groceries, fuel, etc.), data associated with the acquiring institution's country, data associated with an identifier of the country associated with the PAN, data associated with the response code, data associated with the merchant identifier (e.g., merchant name, merchant location, etc.), and data associated with the currency type corresponding to the funds stored in association with the PAN, etc.
[0070] like Figure 2As shown, at step 204, process 200 includes generating the output of a stateful ML model based on the input. For example, the state management system 102 may generate the output of a stateful ML model based on the input. In some non-limiting embodiments or aspects, the state management system 102 may provide input to the stateful ML model, and the stateful ML model may provide output based on the input. The output of the stateful ML model may correspond to a baseline truth value associated with the input. For example, the value of the output of the stateful ML model may be the same as the baseline truth value of the output generated based on the input.
[0071] In some non-limiting embodiments or aspects, the state management system 102 may provide multiple data points as input to a stateful ML model to provide multiple outputs. For example, the state management system 102 may provide one or more data points (e.g., all data points, some data points, a subset of all data points, etc.) as input to a machine learning model to provide one or more outputs based on the input. Each of the multiple outputs may have a classification. For example, each of the multiple initial outputs may have a label (e.g., a label score) associated with a prediction of the output's classification (e.g., a predicted classification).
[0072] like Figure 2 As shown, at step 206, process 200 includes determining whether the output of the stateful ML model corresponds to a baseline truth value associated with the input. For example, the state management system 102 may determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input based on the output generated by the stateful ML model. In some non-limiting embodiments or aspects, the state management system 102 may determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input based on multiple inferences generated using the stateful ML model. For example, the state management system 102 may determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input after the stateful ML model has been used to generate at least 5 inferences, at least 10 inferences, at least 15 inferences, at least 20 inferences, at least 50 inferences, at least 100 inferences, etc.
[0073] In some non-limiting embodiments or aspects, the state management system 102 may receive a baseline truth value associated with the input after generating the output of the stateful ML model based on the input. For example, the state management system 102 may receive the baseline truth value associated with the input in real time after generating the output. In some non-limiting embodiments or aspects, the state management system 102 may receive the baseline truth value within a predetermined time interval after generating the output. For example, the state management system 102 may receive the baseline truth value within 1 second, 5 seconds, 10 seconds, etc., after generating the output.
[0074] In some non-limiting embodiments or aspects, the state management system 102 can determine whether the output of the stateful ML model corresponds to a reference truth value associated with the input by comparing the output of the stateful ML model with a reference truth value. If the state management system 102 determines that the output of the stateful ML model corresponds to the reference truth value, then the state management system 102 can perform an action associated with the initial state of the stateful ML model. For example, if the state management system 102 determines that the output of the stateful ML model corresponds to the reference truth value, then the state management system 102 can assign the initial state of the stateful ML model as the active state of the stateful ML model. If the state management system 102 determines that the output of the stateful ML model does not correspond to the reference truth value, then the state management system 102 can abandon the execution of the action associated with the initial state of the stateful ML model. For example, if the state management system 102 determines that the output of the stateful ML model does not correspond to the reference truth value, then the state management system 102 can abandon the assignment of the initial state of the stateful ML model as the active state of the stateful ML model. In some non-limiting embodiments or aspects, if the state management system 102 determines that the output of the stateful ML model does not correspond to the baseline truth, the state management system 102 may assign the corrected state of the stateful ML model as the active state of the stateful ML model.
[0075] like Figure 2 As shown, at step 208 (“No”), process 200 includes assigning a corrected state to the active state of the stateful ML model. For example, the state management system 102 may assign a corrected state to the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to a baseline truth value. In some non-limiting embodiments or aspects, the state management system 102 may abandon assigning the initial state of the stateful ML model to the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to a baseline truth value associated with the input. In some non-limiting embodiments or aspects, the state management system 102 may perform actions associated with assigning the active state of the stateful ML model based on the determination that the output of the stateful ML model does not correspond to a baseline truth value associated with the input.
[0076] In some non-limiting embodiments or aspects, the state management system 102 can determine the corrected state of the stateful ML model. For example, the state management system 102 can compute a hash key (e.g., hash key value) using a hash function and the state of the stateful ML model relative to an input (e.g., initial input). The state management system 102 can compute the hash key as the output of the hash function based on the state (e.g., initial state) provided as input to the hash function. The hash function may include a locality-sensitive hashing algorithm. In some non-limiting embodiments or aspects, the state management system 102 can compute the hash key based on abandoning (e.g., determining to abandon) the initial state of the stateful ML model as the active state of the stateful ML model.
[0077] In some non-limiting embodiments or aspects, the state management system 102 can determine the corrected state of the stateful ML model based on a state map. For example, the state management system 102 can generate a mapping key, which can be a combination of a hash key (e.g., a hash key generated using the initial state of the stateful ML model) and a baseline truth value for the inputs of the stateful ML model (e.g., a baseline truth value used to provide the inputs for the initial state). Furthermore, the state management system 102 can use the state map and the mapping key for locating (e.g., looking up) the corrected state within the state map to determine the corrected state of the stateful ML model. The state management system 102 can match the mapping key with a location in the state map that corresponds to the mapping key and specifies the corrected state (e.g., cluster, centroid of a cluster, distance to values within a cluster, etc.).
[0078] In some non-limiting embodiments or aspects, the mapping key may include a hash key (e.g., as a first part of the mapping key) and a baseline truth value associated with the input (e.g., as a second part of the mapping key). In some non-limiting embodiments or aspects, the state mapping may include multiple corrected states stored in association (e.g., within a predetermined distance, a predetermined neighborhood, etc.) with multiple hash keys and multiple baseline truth values (e.g., multiple baseline truth values known from the training dataset). In some non-limiting embodiments or aspects, the state management system 102 may determine the corrected states of the stateful ML model based on hash keys calculated using a hash function and the initial state of the stateful ML model relative to the input.
[0079] In some non-limiting embodiments or aspects, the state management system 102 may assign the corrected state of the stateful ML model to the active state of the stateful ML model based on determining the corrected state of the stateful ML model (e.g., using a state map, hash key, and baseline truth associated with the input).
[0080] In some non-limiting embodiments or aspects, the state management system 102 can generate state maps. For example, the state management system 102 can generate state maps based on multiple historical outputs of a machine learning model.
[0081] In some non-limiting embodiments or aspects, the state management system 102 can determine multiple values of the states of the machine learning model associated with correct predictions (e.g., correct predictions for a classification task of the machine learning model) based on multiple historical outputs of the machine learning model. A correct prediction may correspond to a baseline truth value for the output of the stateful ML model. In some non-limiting embodiments or aspects, the state management system 102 can generate clusters based on the multiple values of the states of the machine learning model and determine the centroids of the clusters. In some non-limiting embodiments or aspects, the centroids of the clusters may include the values of the states associated with the baseline truth value for a classification task of the machine learning model.
[0082] like Figure 2 As shown, at step 210 (“Yes”), process 200 includes assigning the initial state of the stateful ML model to the active state of the stateful ML model. For example, the state management system 102 may assign the initial state of the stateful ML model to the active state of the stateful ML model based on determining that the output of the stateful ML model corresponds to a baseline truth value associated with the input.
[0083] In some non-limiting embodiments or aspects, the state management system 102 may repeat steps 204-208 or 204-210 based on receiving input other than the initial input.
[0084] In some non-limiting embodiments or aspects, the state management system 102 may use a stateful ML model (e.g., a trained machine learning model in an online environment) to perform actions such as fraud prevention procedures, reputation procedures, and / or recommendation procedures. For example, the state management system 102 may perform actions based on determining that an action needs to be performed. In some non-limiting embodiments or aspects, the state management system 102 may perform fraud prevention procedures associated with protecting the account of a user (e.g., a user associated with user device 106) based on the output of a machine learning model (e.g., including outputs of predictions associated with a user's account). For example, if the output of a stateful ML model indicates that fraud prevention procedures are necessary, the state management system 102 may perform fraud prevention procedures associated with protecting the user's account. In such examples, if the output of a stateful ML model indicates that fraud prevention procedures are not necessary, the state management system 102 may abandon the execution of fraud prevention procedures associated with protecting the user's account. In some non-limiting embodiments or aspects, the state management system 102 may perform fraud prevention procedures based on classifications of inputs provided by a stateful ML model.
[0085] Now for reference Figure 3A to 3G A schematic diagram of an embodiment 300 for a state correction process (e.g., process 200) for a stateful ML model is shown. In some non-limiting embodiments or aspects, one or more steps of the process may be performed (e.g., wholly, partially, etc.) by a state management system 102 (e.g., one or more means of the state management system 102). In some non-limiting embodiments or aspects, one or more steps of the process may be performed (e.g., wholly, partially, etc.) by another means or group of means separate from or including the state management system 102 (e.g., one or more means of the state management system 102), the ML model repository 104, and / or the user device 106.
[0086] like Figure 3A As indicated by reference numeral 305 in the accompanying drawings, the state management system 102 can receive first input for a stateful ML model from a remote system 302. In some examples, the remote system 302 may be the same as or similar to the ML model repository 104, user device 106, and / or other systems (e.g., transaction service provider system 402, issuer system 404, etc.). In some non-limiting embodiments or aspects, the state management system 102 may receive input for the stateful ML model from the remote system 302 as part of an inference request. For example, the stateful ML model may be provided in an online setup, and the state management system 102 may receive input for the stateful ML model as part of a real-time inference request in the online setup.
[0087] like Figure 3B As shown by reference numeral 310 in the attached figure, the state management system 102 can generate a first output of a stateful ML model based on a first input. For example... Figure 3BAs shown, a stateful ML model may include a recurrent neural network (RNN) ML model. In some non-limiting embodiments or aspects, the stateful ML model may include a classification ML model configured to receive input and provide output based on the input, wherein the output includes a prediction (e.g., a predicted score) of a classification (e.g., a label for classification). The classification may be a binary classification, where a first classification is represented as "1" and a second classification is represented as "0". In some non-limiting embodiments or aspects, the state management system 102 may generate a first output of the stateful ML model as an initial inference based on a first input. In some non-limiting embodiments or aspects, the state of the stateful ML model with respect to the initial inference may include an initial state associated with the first output (e.g., the result of the initial inference) (e.g., equal to the first output, based on the first output, based on a value within a certain amount of the first output, etc.). In some non-limiting embodiments or aspects, the active state of the stateful ML model may be an empty state (e.g., no value, having a zero value, etc.) before or during the initial inference. In some non-limiting embodiments or aspects, the state management system 102 may update the initial state of the stateful ML model relative to the input (e.g., a first input) based on the received input.
[0088] like Figure 3C As shown by reference numeral 315 in the accompanying drawings, the state management system 102 can determine whether a first output of a stateful ML model corresponds to a ground truth (GT) value associated with a first input (e.g., a GT value used for the first output). For example, the state management system 102 can compare the output of the stateful ML model with the GT value used for the first output to determine whether the output of the stateful ML model corresponds to a GT value.
[0089] like Figure 3D As shown by reference numeral 320 in the accompanying drawings, the state management system 102 can assign an initial state as the active state of a stateful ML model based on the first output corresponding to the GT value used for the first output. In this way, the state management system 102 can provide the initial state (e.g., as the active state) as feedback for the next inference.
[0090] like Figure 3EAs shown by reference numeral 325 in the accompanying drawings, the state management system 102 can determine a correction state based on the determination that a first output does not correspond to a ground truth (GT) value used for the first output. In some non-limiting embodiments or aspects, the state management system 102 can determine the correction state of the stateful ML model based on a state mapping. In some non-limiting embodiments or aspects, the state mapping can include multiple correction states stored in association (e.g., within a predetermined distance, a predetermined neighborhood, etc.) with multiple hash keys and multiple benchmark truths (e.g., multiple benchmark truths known from the training dataset). For example, the state management system 102 can compute hash keys (e.g., hash key values) using a hash function and the state of the stateful ML model relative to an input (e.g., an initial input). The state management system 102 can compute hash keys as the output of the hash function based on the state (e.g., the initial state) provided as input to the hash function.
[0091] In some non-limiting embodiments or aspects, the state management system 102 can generate a mapping key, which may be a combination of a hash key (e.g., a hash key generated using the initial state of the stateful ML model) and a baseline truth value for the inputs of the stateful ML model (e.g., a baseline truth value for providing the inputs of the initial state). The state management system 102 can use the state map and the mapping key for locating the correction state within the state map to determine the correction state of the stateful ML model. The state management system 102 can associate the mapping key with the centroid of the cluster in the state map that corresponds to the mapping key and specifies the correction state (e.g., shown as...). h c1 or h c0 This depends on the GT value used for the first output to match.
[0092] like Figure 3F As shown by reference numeral 330 in the attached figure, the state management system 102 can assign a corrected state as the active state of a machine learning model based on the determined corrected state. In this way, the state management system 102 can provide the corrected state (e.g., as the active state) as feedback for the next inference.
[0093] like Figure 3G As shown by reference numeral 335 in the accompanying drawings, the state management system 102 can repeat the previous steps for additional inputs and outputs to determine whether to assign a correction state. For example, when a stateful ML model receives a previously active state and a second input for the next inference, the state management system 102 can determine whether the second output of the stateful ML model corresponds to a reference truth (GT) value associated with the second input. In this example, a second state or a correction state as described above can be assigned to the state management system 102.
[0094] Now for reference Figure 4 , Figure 4 A diagram is a non-limiting embodiment or aspect of an environment 400 in which the systems, products, and / or methods described herein can be implemented. Figure 4 As shown, environment 400 may include a transaction service provider system 402, an issuer system 404, a user device 406, a merchant system 408, an acquirer system 410, and a communication network 412. In some non-limiting embodiments or aspects, Figure 1 Each of the state management system 102, the ML model repository 104, and / or the user device 106 may be implemented by the transaction service provider system 402 (e.g., a portion thereof). In some non-limiting embodiments or aspects, Figure 1 At least one of the state management system 102, ML model repository 104, and / or user device 106 may be implemented by another system, another device, another group of systems, or another group of devices (e.g., a portion thereof) that is separate from or includes the transaction service provider system 402 (e.g., issuer system 404, user device 406, merchant system 408, acquirer system 410, etc.).
[0095] Transaction service provider system 402 may include one or more devices capable of receiving information from and / or transmitting information to the issuer system 404, user device 406, merchant system 408, and / or acquirer system 410 via a communication network 412. For example, transaction service provider system 402 may include computing devices such as servers (e.g., transaction processing servers), server clusters, and / or other similar devices. In some non-limiting embodiments or aspects, transaction service provider system 402 may be associated with the transaction service provider described herein. In some non-limiting embodiments or aspects, transaction service provider system 402 may communicate with a data storage device, which may be local or remote to transaction service provider system 402. In some non-limiting embodiments or aspects, transaction service provider system 402 may be able to receive information from the data storage device, store information in the data storage device, transmit information to the data storage device, or search for information stored in the data storage device.
[0096] The issuer system 404 may include one or more devices capable of receiving and / or transmitting information to the transaction service provider system 402, the user device 406, the merchant system 408, and / or the acquiring system 410 via a communication network 412. For example, the issuer system 404 may include computing devices such as servers, server clusters, and / or other similar devices. In some non-limiting embodiments or aspects, the issuer system 404 may be associated with the issuer institution described herein. For example, the issuer system 404 may be associated with an issuer institution that issues credit accounts, debit accounts, credit cards, debit cards, etc., to users associated with the user device 406.
[0097] User device 406 may include one or more devices capable of receiving and / or transmitting information to and from transaction service provider system 402, issuer system 404, merchant system 408, and / or acquirer system 410 via communication network 412. Alternatively, each user device 406 may include devices capable of receiving and / or transmitting information to other user devices 406 via communication network 412, another network (e.g., temporary network, local network, private network, virtual private network, etc.), and / or any other suitable communication technology. For example, user device 406 may include client devices, etc. In some non-limiting embodiments or aspects, user device 406 may or may not be able to receive information via short-range wireless communication connections (e.g., near field communication (NFC) connections, radio frequency identification (RFID) communication connections, Bluetooth® communication connections, Zigbee® communication connections, etc.) (e.g., from merchant system 408 or from another user device 406), and / or transmit information via short-range wireless communication connections (e.g., to merchant system 408).
[0098] Merchant system 408 may include one or more devices capable of receiving information from transaction service provider system 402, issuer system 404, user device 406, and / or acquiring system 410 via communication network 412 and / or transmitting information to said transaction service provider system, the issuer system, the user device, and / or the acquiring system. Merchant system 408 may also include devices capable of receiving information from user device 406 via communication network 412, communication connections with user device 406 (e.g., NFC connection, RFID communication connection, Bluetooth® communication connection, Zigbee® communication connection, etc.), and / or transmitting information to user device 406 via communication network 412, communication connections, etc. In some non-limiting embodiments or aspects, merchant system 408 may include computing devices, such as servers, server groups, client devices, client device groups, and / or other similar devices. In some non-limiting embodiments or aspects, merchant system 408 may be associated with the merchant described herein. In some non-limiting embodiments or aspects, merchant system 408 may include one or more client devices. For example, merchant system 408 may include a client device that allows merchants to transmit information to transaction service provider system 402. In some non-limiting embodiments or aspects, merchant system 408 may include one or more devices, such as computers, computer systems, and / or peripheral devices, that can be used by merchants to conduct transactions with users. For example, merchant system 408 may include a POS device and / or a POS system.
[0099] Acquiring system 410 may include one or more devices capable of receiving and / or transmitting information to transaction service provider system 402, issuer system 404, user device 406, and / or merchant system 408 via communication network 412. For example, acquiring system 410 may include computing devices, servers, server clusters, etc. In some non-limiting embodiments or aspects, acquiring system 410 may be associated with the acquiring party described herein.
[0100] The communication network 412 may include one or more wired and / or wireless networks. For example, the communication network 412 may include cellular networks (e.g., Long Term Evolution (LTE) networks, third-generation (3G) networks, fourth-generation (4G) networks, fifth-generation (5G) networks, Code Division Multiple Access (CDMA) networks, etc.), Public Land Mobile Networks (PLMNs), Local Area Networks (LANs), Wide Area Networks (WANs), Metropolitan Area Networks (MANs), telephone networks (e.g., Public Switched Telephone Networks (PSTN)), private networks (e.g., private networks associated with transaction service providers), temporary networks, intranets, the Internet, fiber-optic networks, cloud computing networks, etc., and / or combinations of these or other types of networks.
[0101] Figure 4 The number and arrangement of systems, devices, and / or networks shown are provided as examples. Figure 4 Compared to those shown, there may be additional systems, devices, and / or networks; fewer systems, devices, and / or networks; different systems, devices, and / or networks; and / or systems, devices, and / or networks arranged differently. Furthermore, implementation may be within a single system and / or device. Figure 4 The two or more systems or devices shown, or Figure 4 The single system or device shown may be implemented as multiple distributed systems or devices. Alternatively or additionally, a group of systems (e.g., one or more systems) and / or a group of devices (e.g., one or more devices) of environment 400 may perform one or more functions described as being performed by another group of systems or devices of environment 400.
[0102] Now for reference Figure 5 The diagram illustrates example components of a device 500 according to some non-limiting embodiments or aspects. As an example, device 500 may correspond to... Figure 1 At least one of the state management system 102, ML model repository 104, and / or user device 106, and / or Figure 4 At least one of the following: transaction service provider system 402, issuer system 404, user device 406, merchant system 408, and / or acquirer system 410. In some non-limiting embodiments or aspects, Figure 1 or Figure 4 Such systems or apparatuses may include at least one device 500 and / or at least one component of device 500. Figure 5 The number and arrangement of components shown are provided as examples. In some non-limiting embodiments or aspects, with Figure 5Compared to those shown, device 500 may include additional components, fewer components, different components, or components arranged in a different manner. Alternatively, a set of components of device 500 (e.g., one or more components) may perform one or more functions described as being performed by another set of components of device 500.
[0103] like Figure 5 As shown, device 500 may include bus 502, processor 504, memory 506, storage component 508, input component 510, output component 512, and communication interface 514. Bus 502 may include components that allow communication between components of device 500. In some non-limiting embodiments or aspects, processor 504 may be implemented in hardware, firmware, or a combination of hardware and software. For example, processor 504 may include a processor (e.g., central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), etc.), microprocessor, digital signal processor (DSP), and / or any processing component that can be programmed to perform functions (e.g., field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), etc.). Memory 506 may include random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, optical memory, etc.) that stores information and / or instructions for use by processor 504.
[0104] Continue to refer to Figure 5 Storage component 508 may store information and / or software related to the operation and use of device 500. For example, storage component 508 may include a hard disk (e.g., magnetic disk, optical disk, magneto-optical disk, solid-state disk, etc.) and / or another type of computer-readable medium. Input component 510 may include components that allow device 500 to receive information, for example, through user input (e.g., touch screen display, keyboard, keypad, mouse, buttons, switches, microphone, etc.). Alternatively, input component 510 may include sensors for sensing information (e.g., Global Positioning System (GPS) components, accelerometers, gyroscopes, actuators, etc.). Output component 512 may include components that provide output information from device 500 (e.g., display, speaker, one or more light-emitting diodes (LEDs), etc.). Communication interface 514 may include transceiver-like components (e.g., transceivers, separate receivers and transmitters, etc.) that enable device 500 to communicate with other devices, for example, via wired connection, wireless connection, or a combination of wired and wireless connection. Communication interface 514 allows device 500 to receive information from another device and / or provide information to another device. For example, communication interface 514 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi® interface, a cellular network interface, etc.
[0105] Apparatus 500 can perform one or more processes described herein. Apparatus 500 can perform these processes based on processor 504 executing software instructions stored in a computer-readable medium, such as memory 506 and / or storage component 508. The computer-readable medium may contain any non-transitory memory device. The memory device includes memory space located within a single physical storage device or memory space extended across multiple physical storage devices. Software instructions may be read into memory 506 and / or storage component 508 via communication interface 514 from another computer-readable medium or from another device. When executed, the software instructions stored in memory 506 and / or storage component 508 may cause processor 504 to perform one or more processes described herein. Additionally or alternatively, hard-wired circuitry may be used in place of or in conjunction with the software instructions to perform one or more processes described herein. Therefore, the embodiments described herein are not limited to any particular combination of hardware circuitry and software. As used herein, the term “configured to” may refer to an arrangement of software, apparatus, and / or hardware for performing and / or implementing one or more functions (e.g., actions, processes, steps of processes, etc.). For example, "a processor configured to..." can refer to a processor that executes software instructions (such as program code) that cause the processor to perform one or more functions.
[0106] Although embodiments have been described in detail for illustrative purposes, it should be understood that such details are for the purposes described only, and this disclosure is not limited to the disclosed embodiments or aspects, but rather is intended to cover modifications and equivalent arrangements that fall within the spirit and scope of the appended claims. For example, it should be understood that this disclosure contemplates, as far as possible, that one or more features of any embodiment or aspect may be combined with one or more features of any other embodiment or aspect.
Claims
1. A system comprising: At least one processor is configured to: Receives input for stateful machine learning (ML) models; The output of the stateful ML model is generated based on the input; Determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; Based on the determination that the output of the stateful ML model corresponds to the baseline truth value associated with the input, the initial state of the stateful ML model is assigned as the active state of the stateful ML model; as well as Based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input, the correction state of the stateful ML model is assigned as the active state of the stateful ML model.
2. The system of claim 1, wherein when the corrected state of the stateful ML model is assigned to the active state of the stateful ML model, the at least one processor is configured to: The hash key is computed using a hash function and the initial state of the stateful ML model relative to the input; The corrected state of the stateful ML model is determined by computing the hash key using the hash function and the state of the stateful ML model relative to the input, and by using the state mapping, the hash key, and the baseline truth value associated with the input; and The corrected state of the stateful ML model is determined by using the state mapping, the hash key, and the baseline truth value associated with the input, and the corrected state of the stateful ML model is assigned as the active state of the stateful ML model.
3. The system of claim 2, wherein the at least one processor is further configured to: The state mapping is generated based on multiple historical outputs of the stateful ML model.
4. The system of claim 3, wherein when generating the state mapping, the at least one processor is configured to: Determine multiple values of the state associated with the correct prediction of the stateful ML model; Generate clusters of multiple values based on the states of the stateful ML model; and Determine the centroid of the cluster, wherein the centroid of the cluster includes the value of the state associated with the baseline truth value for the classification task of the stateful ML model.
5. The system according to claim 1, wherein the stateful ML model is a recurrent neural network (RNN).
6. The system of claim 1, wherein the at least one processor is further configured to: After generating the output of the stateful ML model based on the input, the baseline truth value associated with the input is received.
7. The system of claim 1, wherein the at least one processor is further configured to: Based on the received input, update the initial state of the stateful ML model relative to the input.
8. A computer-implemented method comprising: Receives input for stateful machine learning (ML) models; The output of the stateful ML model is generated based on the input; Determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; Based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input, the assignment of the initial state of the stateful ML model relative to the input as the active state of the stateful ML model is abandoned. as well as Based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input, an action associated with the active state of the stateful ML model is performed.
9. The computer-implemented method of claim 8, wherein performing the action associated with assigning the active state of the stateful ML model comprises: Based on abandoning the assignment of the initial state of the stateful ML model to the active state of the stateful ML model, a hash key is calculated using a hash function and the state of the stateful ML model relative to the input; The state of the stateful ML model is determined by using the hash function and the state of the stateful ML model relative to the input to compute the hash key, and using the state mapping, the hash key, and the baseline truth value associated with the input to determine the corrected state of the stateful ML model. as well as The corrected state of the stateful ML model is determined by using the state mapping, the hash key, and the baseline truth value associated with the input, and the corrected state of the stateful ML model is assigned as the active state of the stateful ML model.
10. The computer-implemented method according to claim 9, further comprising: The state mapping is generated based on multiple historical outputs of the stateful ML model.
11. The computer-implemented method of claim 10, wherein generating the state mapping comprises: Determine multiple values of the state associated with the correct prediction of the stateful ML model; Generate clusters based on the multiple values of the state of the stateful ML model; as well as Determine the centroid of the cluster, wherein the centroid of the cluster includes the value of the state associated with the baseline truth value for the classification task of the stateful ML model.
12. The computer-implemented method of claim 8, wherein the stateful ML model is a recurrent neural network (RNN).
13. The computer-implemented method according to claim 8, further comprising: After generating the output of the stateful ML model based on the input, the baseline truth value associated with the input is received.
14. The computer-implemented method according to claim 8, further comprising: Based on the received input, update the initial state of the stateful ML model relative to the input.
15. A computer program product comprising at least one non-transitory computer-readable medium, said at least one non-transitory computer-readable medium comprising program instructions that, when executed by at least one processor, cause the at least one processor to: Receives input for stateful machine learning (ML) models; The output of the stateful ML model is generated based on the input; Determine whether the output of the stateful ML model corresponds to a baseline truth value associated with the input; Based on the determination that the output of the stateful ML model corresponds to the baseline truth value associated with the input, the initial state of the stateful ML model is assigned as the active state of the stateful ML model; as well as Based on the determination that the output of the stateful ML model does not correspond to the baseline truth value associated with the input, the correction state of the stateful ML model is assigned as the active state of the stateful ML model.
16. The computer program product of claim 15, wherein the program instruction that causes the at least one processor to assign the corrected state of the stateful ML model to the active state of the stateful ML model causes the at least one processor to: The hash key is computed using a hash function and the initial state of the stateful ML model relative to the input; The corrected state of the stateful ML model is determined by computing the hash key using the hash function and the state of the stateful ML model relative to the input, and by using the state mapping, the hash key, and the baseline truth value associated with the input; and The corrected state of the stateful ML model is determined by using the state mapping, the hash key, and the baseline truth value associated with the input, and the corrected state of the stateful ML model is assigned as the active state of the stateful ML model.
17. The computer program product of claim 16, wherein the program instructions further cause the at least one processor to: The state mapping is generated based on multiple historical outputs of the stateful ML model.
18. The computer program product of claim 17, wherein the program instructions that cause the at least one processor to generate the state mapping further cause the at least one processor to: Determine multiple values of the state associated with the correct prediction of the stateful ML model; Generate clusters of multiple values based on the states of the stateful ML model; and Determine the centroid of the cluster, wherein the centroid of the cluster includes the value of the state associated with the baseline truth value for the classification task of the stateful ML model.
19. The computer program product of claim 15, wherein the stateful ML model is a recurrent neural network (RNN).
20. The computer program product of claim 15, wherein the program instructions further cause the at least one processor to: After generating the output of the stateful ML model based on the input, the baseline truth value associated with the input is received.
21. The computer program product of claim 15, wherein the program instructions further cause the at least one processor to: Based on the received input, update the initial state of the stateful ML model relative to the input.