In-context learning through interaction
By generating feature vectors and aggregating data, and using pre-trained models for task-specific predictions, this approach addresses the problem of limited model performance improvement in existing technologies, achieving efficient adaptation to different tasks and saving computational resources.
Patent Information
- Application Number
- CN202580001622.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-13
- Filing Date
- 2025-03-12
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies offer limited performance improvements when using more data and larger models for interactive data prediction, and natural language models require extensive fine-tuning and training heads to perform different tasks.
By generating feature vectors and aggregating data, pre-trained machine learning models are used for task-specific predictions, reducing reliance on model weight updates and employing a context-based learning approach to adapt to different tasks.
It improves the performance of the model across various tasks, reduces the need for fine-tuning, and enables the model to be adapted to new tasks with a small number of instances, thus saving computational resources.
Smart Images

Figure CN120958446A_ABST
Abstract
Description
[0001] Cross-referencing related applications
[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 564,835, filed March 13, 2024, which is incorporated herein by reference in its entirety for all purposes. Background Technology
[0003] Interaction data used for interactions between devices (e.g., data access requests, transactions, etc.) can have many different characteristics and can include different data based on the type of interaction. Different machine learning models can be used to make predictions based on these characteristics to make the interaction process more efficient. Different machine learning models can be trained for each different type of interaction.
[0004] To improve the performance of all different machine learning models, it is common practice to utilize more data and make the models larger. This has been successful across various fields, especially in natural language processing with large language models (LLMs).
[0005] However, due to the interactive data used, more data and a larger model may not improve model performance, as predicting arbitrary numbers is unrealistic. Simply using more data and a larger model may not improve predictive performance on interactive data. Furthermore, even with pre-trained models, current use of natural language models requires numerous fine-tuning and training heads to perform different tasks.
[0006] The embodiments of this disclosure address this problem and other problems individually or collectively. Summary of the Invention
[0007] One embodiment relates to a method in which a computer obtains interaction data and aggregated data for a current interaction. The current interaction relates to means of requesting access to a resource. The computer can use the interaction data and aggregated data to generate feature vectors. The computer can then generate a cue containing a set of feature-label pairs. Each feature-label pair may include a task-specific label and additional feature vectors. The additional feature vectors may include additional interaction data and additional aggregated data. The additional interaction data and the additional aggregated data are associated with other interactions having the same classification type as the current interaction. The computer can load a pre-trained machine learning model into memory, the pre-trained machine learning model being trained to determine predictions for a specific task. The computer can input the cue and a query including the feature vectors into the pre-trained machine learning model. The computer can then use the pre-trained machine learning model to determine a task-specific prediction for the query and the cue. The task-specific prediction may specify a response state to the current interaction.
[0008] In some embodiments, a computer may use feature vectors to generate queries.
[0009] In some embodiments, the computer can obtain multiple historical interaction data, multiple historical aggregate data, and multiple historical task-specific labels from multiple historical interactions. The computer can use the multiple historical interaction data, multiple historical aggregate data, and multiple historical task-specific labels to generate multiple historical feature vectors. The computer can then use the multiple historical feature vectors to generate historical prompts and historical queries. The computer can use the historical prompts and historical queries to train a pre-trained machine learning model.
[0010] In some embodiments, the interaction data includes an indication of a certain interaction type among a plurality of interaction types. Generating a feature vector using the interaction data and aggregated data may include the following steps: A computer may generate an interaction data portion of the feature vector. The interaction data portion includes a plurality of interaction type entries. The computer may set the value of the interaction type entry corresponding to the stated interaction type as the value of the interaction data. The computer may set the values of other interaction type entries that do not correspond to the stated interaction type as default values.
[0011] Another embodiment relates to a computer including a processor and a computer-readable medium coupled to the processor. The computer-readable medium contains code that can be executed by the processor to implement any of the methods herein.
[0012] Another embodiment relates to a system comprising an interactive database storing multiple interactive data and multiple aggregated data, and an evaluation computer communicating with the interactive database. The evaluation computer includes a processor, memory, and a computer-readable medium coupled to the processor. The computer-readable medium contains code executable by the processor to implement any of the methods herein.
[0013] Further details regarding embodiments of this disclosure can be found in the detailed description and accompanying drawings. Attached Figure Description
[0014] Figure 1 A block diagram of a system according to an embodiment is shown.
[0015] Figure 2 A block diagram illustrating unified data according to an embodiment is shown.
[0016] Figure 3 A block diagram illustrating the model inputs of a machine learning model according to an embodiment is shown.
[0017] Figure 4 A flowchart of the training and inference method according to an embodiment is shown.
[0018] Figure 5 A block diagram illustrating a machine learning model according to an embodiment is shown.
[0019] Figure 6 A flowchart illustrating a prediction determination method according to an embodiment is shown.
[0020] Figure 7 A block diagram illustrating masking uniform data according to an embodiment is shown.
[0021] Figure 8 A block diagram of the components of an evaluation computer according to an embodiment is shown.
[0022] Figure 9 A block diagram of the components of a computer according to an embodiment is shown.
[0023] the term
[0024] Before discussing the embodiments of this disclosure, some terms may be described in more detail.
[0025] A "machine learning computer" can include devices for creating, training, and / or otherwise manipulating models. A machine learning computer can train machine learning models.
[0026] A “machine learning model” (ML model) can include a software module configured to run on one or more processors to provide categorical or numerical properties of one or more samples. An ML model can include various parameters (e.g., coefficients, weights, thresholds, functional properties of functions, such as activation functions). As an example, an ML model can include at least 10, 100, 1,000, 5,000, 10,000, 50,000, 100,000, or one million parameters. Sample data (e.g., training samples) can be used to generate an ML model to make predictions on test data. Various numbers of training samples can be used, such as at least 10, 100, 1,000, 5,000, 10,000, 50,000, 100,000, or at least 200,000 training samples. One example is unsupervised learning models such as Hidden Markov Models (HMMs), clustering (e.g., hierarchical clustering, k-means, mixture models, model-based clustering, density-based noisy applied spatial clustering (DBSCAN), and the OPTICS algorithm), methods for learning latent variable models such as the Expectation-Maximization (EM) algorithm, the Method of Moments, and blind signal separation techniques (e.g., principal component analysis, independent component analysis, nonnegative matrix factorization, singular value decomposition), and anomaly detection (e.g., local anomaly factors and isolated forests). Another example model type is supervised learning that can be used with embodiments of this disclosure. Example supervised learning models can include various methods and algorithms, including analytical learning, statistical models, artificial neural networks that can have 1-10 layers as instances (e.g., including convolutional and / or transformer layers), recurrent neural networks (e.g., Long Short-Term Memory, LSTM), reinforcement (meta-algorithms), bootstrapping aggregation (bagging) such as random forests, support vector machines (SVM), support vectors (SVR), Bayesian statistics, case-based reasoning, decision tree learning, inductive logic programming, linear regression, logistic regression, Gaussian process regression, genetic programming, grouping methods for data processing, kernel estimators, learning automata, learning classifier systems, minimum message length (decision trees, decision graphs, etc.), multilinear subspace learning, Naive Bayes classifiers, maximum entropy classifiers, conditional random fields, nearest neighbor algorithms, probabilistic approximate correct learning (PAC) learning, chain wave descent rules, knowledge acquisition methods, symbolic machine learning algorithms, sub-symbolic machine learning algorithms, minimum complexity machines (MCM), ordinal classification, data preprocessing, handling imbalanced datasets, statistical relational learning, or Proaftn (a multi-criteria classification algorithm), or any combination of these types. Supervised learning models can be trained in various ways using a variety of cost / loss functions (e.g., least squares and absolute difference from known classifications) that define the error from known labels and various optimization techniques, such as backpropagation, steepest descent, conjugate gradients, and Newton and quasi-Newton techniques.
[0027] A "deep neural network (DNN)" can be a neural network with multiple layers between its input and output. Each layer of a deep neural network may represent a mathematical operation used to transform the input into the output. Specifically, a "recurrent neural network (RNN)" can be a deep neural network in which data can move forward and backward between the layers of the neural network.
[0028] An encoder processes an input sequence to create a vector. An encoder can process an input sequence to generate an embedding vector. An encoder can encode data from a higher dimension to a lower dimension. A decoder processes vectors to create an output sequence. A decoder can process an embedding vector or a vector modified by it to generate an output sequence. A decoder can decode data from a lower dimension to a higher dimension. Both the encoder and decoder can be separate fully connected neural networks. The encoder and decoder can be recurrent neural networks (RNNs) or variants thereof (e.g., Long Short-Term Memory (LSTM), Gated Recurrent Units (GRUs), etc.), convolutional neural networks (CNNs), and transformer models. An encoder-decoder model can include several encoders and several decoders.
[0029] A "model database" can include a database that stores machine learning models. Machine learning models can be stored in a model database in various forms, such as a set of parameters or other values that define the model. Models in a model database can be stored in association with keywords that convey certain aspects of the model. For example, a model used to evaluate news articles could be stored in a model database associated with the keywords "news," "propaganda," and "information." Machine learning computers can access and retrieve models from the model database, modify models in the model database, delete models from the model database, or add new models to the model database.
[0030] A "feature vector" can include a set of measurable attributes (or "features") representing an object or entity. A feature vector can include a collection of data represented numerically as an array or vector structure. A feature vector can also include a collection of data that can be represented as mathematical vectors on which vector operations such as scalar products can be performed. Feature vectors can be determined or generated from input data. Feature vectors can be used as input to machine learning models, enabling the models to produce a certain output or classification. Based on the properties of the input data, feature vectors can be constructed in various ways. For example, for a machine learning classifier that classifies words as correctly spelled or incorrectly spelled, the feature vector corresponding to a word such as "LOVE" can be represented as a vector (12, 15, 22, 5), corresponding to the alphabetical index of each letter in the input data word. For more complex "inputs," such as human entities, exemplary feature vectors can include features such as a person's age, height, weight, and a numerical representation of relative happiness. Feature vectors can be represented and stored electronically in a feature storage area. Furthermore, feature vectors can be normalized (i.e., given unit values). As an example, the feature vector (12, 15, 22, 5) corresponding to "LOVE" can be normalized to approximately (0.40, 0.51, 0.74, 0.17).
[0031] "Interaction" can include mutual action or influence. Interaction can include communication, contact, or exchange between parties, devices, and / or entities. Example interactions include transactions between two parties and data exchange between two devices. In some embodiments, interaction can include a user requesting access to secure data, secure web pages, secure locations, etc. In other embodiments, interaction can include payment transactions, in which two devices may interact to facilitate payment.
[0032] An "access device" can be any suitable device that provides access to a remote system. Access devices can also be used to communicate with a coordinating computer, communication network, or any other suitable system. Access devices can typically be located anywhere suitable, such as at the merchant's location. Access devices can take any suitable form. Some examples of access devices include POS or point-of-sale devices (e.g., POS terminals), cellular phones, personal digital assistants (PDAs), personal computers (PCs), tablet PCs, handheld dedicated readers, set-top boxes, electronic cash registers (ECRs), vending machines, automatic teller machines (ATMs), virtual cash registers (VCRs), kiosks, security systems, access systems, etc.
[0033] Access devices can use any suitable contact or contactless operating mode to send or receive data from or associated with mobile communication devices or payment devices. For example, an access device may have a card reader, which may include electrical contacts, radio frequency (RF) antennas, optical scanners, barcode readers, or magnetic stripe readers to interact with portable devices such as payment cards.
[0034] The term "resource" can include any asset that can be used or consumed. For example, a resource can be an electronic resource (e.g., stored data, received data, computer accounts, network-based accounts, email inboxes), a physical resource (e.g., tangible objects, buildings, safes, or physical locations), or other electronic communications between computers (e.g., communication signals corresponding to accounts used to execute transactions).
[0035] A "resource provider" can be an entity that can provide resources such as goods, services, information, and / or access. Examples of resource providers include merchants, data providers, transportation departments, government entities, site and residential operators, etc.
[0036] An "authorization request message" can be an electronic message requesting authorization for an interaction. In some embodiments, the message is sent to a transaction processing computer and / or the issuer of a payment card to request authorization for the transaction. Authorization request messages according to some embodiments may conform to International Organization for Standardization (ISO) 8583, a standard for systems that exchange information related to electronic transactions associated with payments made by a user using a payment device or payment account. An authorization request message may include an issuer account identifier that may be associated with a payment device or payment account. The authorization request message may also contain additional data elements corresponding to "identification information," including (by way of example only): service code, CVV (card verification value), dCVV (dynamic card verification value), PAN (primary account number or "account number"), payment token, username, expiration date, etc. The authorization request message may also contain "transaction information," such as any information associated with the current transaction, such as transaction value, merchant identifier, merchant location, acquiring bank identifier (BIN), card acceptor ID, information identifying the item being purchased, and any other information that may be used to determine whether to identify and / or authorize the transaction.
[0037] An "authorization response message" can be a message responding to an authorization request. In some cases, an authorization response message can be an electronic message reply to an authorization request message generated by the issuing financial institution or transaction processing computer. As an example only, an authorization response message may include one or more of the following status indicators: Approved – the transaction is approved; Rejected – the transaction is not approved; or Call Center – further information is pending, and the resource provider will call the toll-free authorization number. An authorization response message may also include an authorization code, which can be a code indicating approval of the transaction returned by the credit card issuing bank to the merchant's access device (e.g., POS equipment) in response to the authorization request message in the electronic message (directly or via the transaction processing computer). This code can serve as evidence of authorization.
[0038] "Authorizing entity" can be an entity that authorizes a request. Instances of authorizing entities can be issuers, government agencies, document repositories, access administrators, etc. Authorizing entities can operate authorizing entity computers. "Issuer" can refer to a commercial entity (e.g., a bank) that issues and optionally maintains user accounts. Issuers can also issue payment credentials to consumers that are stored on user devices (such as cellular phones, smart cards, tablets, or laptops, or in some embodiments, portable devices).
[0039] "Interaction data" may include data related to and / or related to the processing of an interaction. Interaction data may include data used for authorizing interactions, clearing interactions, and any other types of interactions. Interaction data may include amounts, timestamps, user identifiers, resource provider identifiers, transmitter computer identifiers, network processing computer identifiers, authorizing entity identifiers, interaction types, interaction results, and / or any other data related to the interaction between the user operating the user device and the resource provider computer operating the resource provider computer to allow the user to obtain resources from the resource provider.
[0040] "Aggregated data" can include data related to entities, computers, and accounts associated with the interaction. Aggregated data can relate to things involved in the interaction as indicated by the interaction data. Aggregated data can include data related to user accounts (e.g., primary user accounts), resource providers, transmission computers, and authorizing entities involved in the interaction. For example, aggregated data related to user accounts can include historical data associated with user accounts involved in the interaction. Aggregated data related to resource providers can include data related to resource providers and / or resource provider computers, such as the number of interactions performed by the resource provider in the past day. Aggregated data related to transmission computers can include data related to the acquiring party and / or transmission computers, such as the number of interactions processed by the transmission computer in the past month. Aggregated data related to authorizing entities can include data related to authorizing entities and / or authorizing entity computers, such as data used to determine whether to authorize an interaction.
[0041] "Hints" can include inputs to a machine learning model. Hints can be submitted to a machine learning model, such as a neural network or a large language model, to modify the processing and determine the response to a query. Hints can be formulated to guide the behavior of a machine learning model. Hints can include questions, instructions, contextual information, few-shot instances, and partial inputs for the machine learning model to complete or continue.
[0042] A query can include a request. A query can include a request for information. A query can seek information retrieval. A query can be submitted to a machine learning model, such as a neural network or a large language model, to receive a response about the query.
[0043] A “task” can include a specific problem or prediction to be evaluated. Tasks can be performed by a machine learning model. A task can be a process with a specific objective. Example tasks could include: 1) predicting whether an interaction is authorized, 2) predicting whether an interaction is liquidated, 3) predicting whether an interaction is fraudulent, 4) determining the predicted ratings (e.g., credit scores) of users involved in an interaction, 5) determining the category of an interaction (e.g., spending category), 6) determining the category of users involved in an interaction (e.g., user archetypes), and 7) identifying anomalous ratings for an interaction.
[0044] "Labels" can include data that identifies associated data. Labels can indicate information about associated feature vectors. Labels can indicate a specific category of associated feature vectors. For example, a label can indicate a "non-fraudulent" category for feature vectors that include data used for interaction. As another example, a label can indicate the category of "transportation" or "entertainment" spending to be posted to a user's account when an associated interaction is posted to that account.
[0045] "Task-specific tags" can include identifying data specific to a particular problem or prediction. Task-specific tags can include tags associated with a specific task. Each task can involve multiple task-specific tags. Task-specific tags can indicate both the labeling of the data and the task itself. For example, a task-specific tag could be "authorized," which can be specific to the task determining whether to authorize interaction.
[0046] A “processor” can include means for processing things. In some embodiments, a processor can include any suitable one or more data computing means. A processor can include one or more microprocessors that work together to perform a desired function. A processor can include a CPU that includes at least one high-speed data processor sufficient to execute program components for performing user and / or system-generated requests. A CPU can be a microprocessor such as AMD’s Athlon, Duron and / or Opteron; IBM and / or Motorola’s PowerPC; IBM and Sony’s Cell processor; Intel’s Celeron, Itanium, Pentium, Xeon and / or XScale; and / or similar processors.
[0047] "Memory" can be any suitable one or more devices capable of storing electronic data. Suitable memory can contain non-transitory computer-readable media whose storage can be executed by a processor to implement desired methods. Instances of memory can include one or more memory chips, disk drives, etc. Such memory can be operated using any suitable electrical, optical, and / or magnetic modes of operation.
[0048] A "server computer" can include a powerful computer or cluster of computers. For example, a server computer can be a mainframe, a small cluster of computers, or a group of servers operating as a single unit. In one instance, a server computer can be a database server coupled to a web server. A server computer can contain one or more computing devices and can use any of a variety of computing architectures, arrangements, and compilations to serve requests from one or more client computers. Detailed Implementation
[0049] Embodiments of this disclosure can provide a context-based learning framework for: 1) improving model performance across various tasks, and 2) utilizing data from different interaction phases to adapt a trained model to different tasks without requiring fine-tuning for each task. To this end, embodiments may provide: 1) a method for unifying interactions from different sources and phases into uniform data; 2) a model capable of processing uniform data (e.g., represented as feature vectors); and 3) a set of context-based learning objectives such that different tasks can be performed by the same machine learning model.
[0050] By leveraging this framework, implementations can pre-train a single, larger model (e.g., larger than a current task-specific model) using a large amount of interactive data. Compared to single-task-based models such as Smarter Posting, this model offers improved performance and requires no fine-tuning, adapting to new tasks with a small number of instances instead of training new tasks with many instances.
[0051] Some previous methods utilize few-shot learning. Few-shot learning allows a model to learn to make predictions from only a few instances (e.g., a limited amount of training data). The idea behind few-shot learning is learning to learn. Instead of training on individual samples from each task and then inferring from samples from the same task, few-shot learning trains a model on a variety of tasks and learns how to learn from a few samples within that task. Its success demonstrates the possibility of training a model for different tasks and generating task-specific predictions from a few instances of that task. The embodiments provide advantages over previous methods that utilize few-shot learning.
[0052] According to embodiments, machine learning models can utilize context-based learning instead of few-shot learning. Few-shot learning models require updating model weights based on a few instances provided to the model. However, in context-based learning utilized in some embodiments, the machine learning model does not need to update model weights based on a few instances; instead, it can make predictions based on a few instances and the current instance, considering both as input. It is advantageous that the machine learning model does not need to be updated in this manner during use, as the computer can save computational resources each time the machine learning model is used.
[0053] Large Language Models (LLMs) are based on transformer models and are pre-trained on large amounts of text data. The typical training objective of a LLM is next-word prediction. For few-shot learning, LLMs can take several instances and questions together and generate correct responses; this is called learning in context.
[0054] An embodiment may provide a computer that can generate a unified data structure for feature vectors of a learning model within a context. The feature vectors can represent features of interactive data from multiple different learning tasks and interaction types. The computer can generate feature vectors from interactive data and aggregated data. Interactive data relates to data within the current interaction, while aggregated data relates to data surrounding the interactive data.
[0055] Computers can utilize feature vectors from the current interaction and feature vectors from other interactions related to the current interaction in some way. Computers can use pre-trained machine learning models to determine task-specific predictions for a given response state to the current interaction (e.g., authorized or unauthorized, fraudulent or non-fraudulent, risk scoring, etc.).
[0056] I. Example Network Architecture
[0057] Examples can utilize the system described herein to train, maintain, and leverage machine learning models capable of generating predictions (e.g., task-specific labels) about the current interaction based on the current interaction and other interactions.
[0058] Figure 1 A system 100 according to an embodiment of the present disclosure is illustrated. The system 100 includes an evaluation computer 102, an interactive database 104, a downstream device 106, a network processing computer 108, a user device 110, an access device 112, a resource provider computer 114, a transmission computer 116, and an authorization entity computer 118.
[0059] Evaluation computer 102 can operationally communicate with interactive database 104 and downstream device 106. Network processing computer 108 can operationally communicate with interactive database 104, transmission computer 116, and authorizing entity computer 118. Transmission computer 116 can operationally communicate with resource provider computer 114, which can operationally communicate with access device 112 and user device 110. Access device 112 can operationally communicate with user device 110.
[0060] To simplify the explanation, Figure 1 A specific number of components are shown. However, it should be understood that embodiments of the invention may include more than one of each component. Additionally, some embodiments of the invention may include more than one of each component. Figure 1 The components shown are fewer or more components.
[0061] Figure 1Messages between devices in the system 100 shown can be transmitted using secure communication protocols, such as, but not limited to, File Transfer Protocol (FTP); Hypertext Transfer Protocol (HTTP); Secure Hypertext Transfer Protocol (HTTPS); SSL; ISO (e.g., ISO 8583), etc. The communication network can include any and / or a combination of the following: direct interconnection; the Internet; a local area network (LAN); a metropolitan area network (MAN); Operational Mission as a Node on the Internet (OMNI); a secure custom connection; a wide area network (WAN); a wireless network (e.g., employing protocols such as, but not limited to, Wireless Application Protocol (WAP), I-mode, etc.); etc. The communication network can use any suitable communication protocol to generate one or more secure communication channels. In some cases, the communication channel can contain a secure communication channel, which can be established in any known manner, such as by using mutual authentication and session keys, and establishing a Secure Sockets Layer (SSL) session.
[0062] Evaluation computer 102 can create, train, utilize, and / or otherwise manipulate machine learning models to evaluate data. Evaluation computer 102 can obtain interactive and aggregated data stored in interactive database 104. Evaluation computer 102 can use the interactive and aggregated data for interaction to generate feature vectors (e.g., using a uniform data format). Evaluation computer 102 can leverage a pre-trained machine learning model to determine task-specific predictions for prompts and queries created from interactive and aggregated data from interactive database 104. In some embodiments, evaluation computer 102 can use the interactive and aggregated data obtained from interactive database 104 to train a machine learning model.
[0063] Interactive database 104 can store interactive and aggregated data obtained from network processing computer 108. Interactive database 104 can include any suitable database. The database can be a general-purpose, fault-tolerant, relational, scalable, and secure database, such as one available from Oracle. TM or Sybase TM Databases acquired through commercial purchases.
[0064] Downstream device 106 may include means that can utilize a machine learning model trained by evaluation computer 102. Downstream device 106 may obtain inference data from evaluation computer 102 based on a query, or may generate inference data using a machine learning model trained by evaluation computer 102.
[0065] User device 110, access device 112, resource provider computer 114, transmission computer 116, network processing computer 108, and authorization entity computer 118 can handle interactions between the user of the user device and the resource provider of the resource provider computer.
[0066] User device 110 may include one or more computers, laptops, tablets, mobile devices, cellular phones, wearable devices (e.g., watches, glasses, lenses, clothing, etc.), personal digital assistants (PDAs), Internet of Things (IoT) devices, etc. User device 110 may initiate interactions (e.g., transactions) with resource provider computers and / or access devices. For example, user device 110 may select one or more items for interaction at a resource provider location (e.g., a grocery store). During checkout, the user may be instructed to tap (e.g., enter near-field communication range) user device 110 towards access device 112.
[0067] Access device 112 may include a device operated by a resource provider. For example, access device 112 may include a mobile device, a POS terminal, a laptop computer, etc. Access device 112 may communicate with another device (e.g., user device 110) to perform an interaction. During the interaction, access device 112 may receive credentials from the user device and may provide interaction data to resource provider computer 114 for authorization of the interaction. In some embodiments, access device 112 may generate an authorization request message containing at least the interaction data. Access device 112 may provide the authorization request message to resource provider computer 114.
[0068] Resource provider computer 114 may include any suitable computing device operated by the resource provider (e.g., a merchant). In some embodiments, resource provider computer 114 may include one or more server computers that may host one or more websites associated with the resource provider (e.g., a merchant). In some embodiments, resource provider computer 114 may be configured to transmit data to network processing computer 108 via transmission computer 116 as part of a payment verification and / or authentication process for transactions between a user (e.g., a consumer) and the resource provider. Resource provider computer 114 may also be configured to generate authorization request messages for transactions between the resource provider and the user and route these authorization request messages to authorization entity computer 118 for transaction processing.
[0069] The transmission computer 116 may include a server computer. The transmission computer 116 may be associated with an acquirer, which is typically an entity with a business relationship with a particular merchant or other entity (e.g., a commercial bank). Some entities may perform both issuer and acquirer functions. Some embodiments may cover such a single entity as an issuer-acquirer.
[0070] Network processing computer 108 may include a server computer. Network processing computer 108 may be located between transport computer 116 and authorization entity computer 118. Network processing computer 108 may include a data processing subsystem, a network, and operations for supporting and delivering authorization services, exception document services, and clearing and settlement services. For example, network processing computer 108 may include (e.g., via an external communication interface) a server and an information database coupled to a network interface. Network processing computer 108 may represent a transaction processing network. An exemplary transaction processing network may include VisaNet. TM For example, VisaNet TM Transaction processing networks like VisaNet can handle credit card transactions, debit card transactions, and other types of commercial transactions. Specifically, VisaNet... TM This includes the VIP system (Visa Integrated Payment System) for processing authorization requests and the Base II system for performing clearing and settlement services. The network processing computer 108 can use any suitable wired or wireless network, including the Internet.
[0071] When the interaction is processed by the network processing computer 108, the network processing computer 108 can store the interaction data in the interaction database 104 for use in the interaction.
[0072] Authorized entity computer 118 may include a computer operated by an authorized entity. Authorized entity computer 118 may be associated with an authorized entity, which may be an entity that authorizes requests. An example of an authorized entity may be an issuer, which may include an entity that maintains user accounts (e.g., a bank). The issuer may also issue and manage accounts associated with user device 110.
[0073] Interactions can include different types of interactions. Interactions can be authorization interactions, settlement interactions, etc. Interactions can be transactions such as payment transactions (e.g., for purchasing goods or services), access transactions (e.g., for accessing a transportation system), or any other suitable transaction.
[0074] As an illustrative example of authorized interaction, a user of user device 110 can use user device 110 to interact with a resource provider (e.g., a merchant). User device 110 can interact with access device 112 at a resource provider location associated with resource provider computer 114. For example, a user can tap user device 110 against an NFC reader in access device 112. Alternatively, user device 110 can electronically indicate payment account information to resource provider computer 114, as in an online transaction. In some cases, user device 110 can transmit an account identifier, such as a payment token, to resource provider computer 114 or access device 112.
[0075] To authorize the interaction, an authorization request message may be generated by the access device 112 or the resource provider computer 114, and then forwarded to the transmission computer 116 (e.g., the acquiring computer). Upon receiving the authorization request message, it may be sent to the network processing computer 108. The network processing computer 108 then forwards the authorization request message to the corresponding authorizing entity computer 118 associated with the authorizing entity linked to the user's payment account.
[0076] After receiving the authorization request message, the authorizing entity computer 118 can determine whether to authorize the interaction. The authorizing entity computer 118 can generate an indication of whether the interaction is authorized. The authorizing entity computer 118 can generate an authorization response message containing the indication and can provide the authorization response message to the network processing computer 108 to indicate whether the current transaction is authorized (or not authorized). The network processing computer 108 can then provide the authorization response message to the transport computer 116. In some embodiments, even if the authorizing entity computer 118 has authorized the interaction, the network processing computer 108 can refuse the interaction, for example, based on the value of a fraud risk score.
[0077] Authorized entity computer 118 may store interaction data related to the interaction in interaction database 104. Interaction data may include amount, date, time, account (e.g., master account (PAN)), user identifier, user device identifier, resource provider identifier, authorized entity computer identifier, transmission computer identifier, network processing computer identifier, list of resources involved in the interaction, indication of whether the interaction is authorized, and / or any other data related to the interaction and / or the processing of the interaction.
[0078] The transmitting computer 116 can then provide the authorization response message to the resource provider computer 114. After receiving the authorization response message, the resource provider computer 114 can provide it to the user of the user device 110. The response message can be displayed to the user by the access device 112 or printed on a physical receipt. Alternatively, if the interaction is online, the resource provider computer 114 can provide the user device 110 with a webpage or other instructions for the authorization response message as a virtual receipt.
[0079] At the end of the day, the network processing computer 108 can perform clearing and settlement interactions. The clearing interaction process involves exchanging financial details between the acquiring party and the authorizing entity to facilitate posting to the user's payment account and verifying the user's settlement position. Clearing interactions are another instance of interactions that can be processed by the system. The network processing computer 108 can store data related to clearing interactions in the interaction database 104.
[0080] II. Unify the data into feature vectors
[0081] Different types of interactions can have different feature sets. For example, interactions in the context of a transaction can be authorization interactions, clearing interactions, etc., and can have different features for each stage of transaction processing. To address the technical challenge of having machine learning models applicable to all interaction data, computers can unify interactions into feature vectors of the same format. Figure 2 A block diagram illustrating unified data according to an embodiment is shown.
[0082] Figure 2 This includes interaction data 202 for interaction, which can be used to generate unified data for machine learning models. Interaction data 202 can be used for authorization interactions, settlement interactions, etc., and can be generated by... Figure 1 The system displays interactive data processed by the system.
[0083] Interaction data 202 may include interaction identifiers, amounts, timestamps, user identifiers, user device identifiers, account numbers, resource provider identifiers, resource provider computer identifiers, acquiring party identifiers, transmitter computer identifiers, network processor identifiers, network processing computer identifiers, authorizing entity identifiers, authorizing entity computer identifiers, interaction types, interaction results, and / or any other data related to the interaction between the user operating the user device and the resource provider computer operating the resource provider computer to allow the user to obtain resources from the resource provider. Interaction data 202 may include data from interactions performed by the interaction processing system and may be stored in an interaction database.
[0084] Aggregated data can include data related to the entities, devices, and accounts involved in the interaction. Aggregated data can include data that is peripheral to the interaction but related to it. Aggregated data can include data related to historical interactions. Aggregated data can include statistics from previous interactions. For example, aggregated data can include interaction speed, interaction volume, average interaction amount, etc.
[0085] Unified data can be created as Figure 2 The feature vector 200 shown includes an interactive data portion 220 and an aggregated data portion 230. A computer can generate the feature vector 200 from the interactive data 202 and the aggregated data associated with the interactive data 202. For example, a computer can use an interactive identifier from the interactive data 202 to obtain aggregated data from an interactive database.
[0086] Interaction data section 220 may include data obtained from interaction data 202. Interaction data section 220 may include one or more segments related to different types of interactions. For example, interaction data section 220 may include interaction authorization data 222 and interaction settlement data 224. Interaction data section 220 may include additional segments (not shown) related to other types of interactions.
[0087] When generating feature vector 200, the computer can determine the interaction type of interaction data 202. For example, interaction data 202 may include an indicator that indicates the interaction type represented by interaction data 202. The computer can input interaction data 202 into a segment of interaction data portion 220 that matches the interaction type. If the format is different, the computer can also format interaction data 202 into the format indicated by the segment of interaction data portion 220.
[0088] For example, interaction data 202 may include an indicator that indicates the type of interaction used for authorization integration. The computer may input interaction data 202 into the interaction authorization data 222 segment of interaction data section 220.
[0089] For fragments of interaction data 220 that do not correspond to the interaction type of interaction data 202, the computer can assign a default value to each field indicating a missing value. Whenever a value exists in the feature vector from the interaction data, that value can be retained in the feature vector; otherwise, a default missing value token can be inserted for the field. Default values can also be used for pre-training tasks for masking purposes when the computer trains a machine learning model. For example, an authorization interaction will not have a liquidation interaction value, so the liquidation interaction value can be set to its corresponding default value, but if some values in the authorization are the value "none," then the value can remain "none." Therefore, default values can act as identifiers as well as null values, which 1) identify the data type and 2) identify that the data value is empty.
[0090] Interaction authorization data 222 may include an indicator indicating that the interaction is an authorized interaction. Interaction authorization data 222 may include interaction data related to the authorized interaction. For example, interaction authorization data 222 may include authorization request message data and / or other data related to requesting and / or obtaining authorization for the interaction. Authorization request message data may include date, time, amount, user identifier, resource provider identifier, and user credentials. User credentials may be a password, passcode, or secret message. As another example, user credentials may include any information associated with and / or identifying an account (e.g., a payment account and / or a payment device associated with said account). Examples of account information may include account identifiers such as primary account number (PAN), token, sub-token, gift card number or code, prepaid card number or code, username, expiration date, CVV (card verification value), dCVV (dynamic card verification value), CVV2 (card verification value 2), etc. An example of a PAN is a 16-digit number, such as "41470900 0000 1234".
[0091] Interactive clearing data 224 may include indicators that indicate the clearing interaction is a clearing interaction. Interactive clearing data 224 may include interactive data related to the clearing interaction. For example, interactive clearing data 224 may include clearing-related data, settlement-related data, batch data, etc.
[0092] Instances of interactive authorization data and interactive settlement data are provided as fragments of interactive data section 220, but it should be understood that interactive data section 220 may include fields for other interactive types.
[0093] Aggregated data section 230 may include data obtained in some way from aggregated data related to interactive data 202. Aggregated data section 230 may include one or more segments related to different categories of aggregated data. For example, aggregated data section 230 may include aggregated account data 232, aggregated resource provider data 234, aggregated transfer data 236, and aggregated authorizing entity data 238. Aggregated data section 230 may include additional segments (not shown) related to other categories of aggregated data.
[0094] When generating feature vector 200, the computer can obtain aggregated data from an interactive database, which may include aggregated data of different categories of aggregated data in aggregated data section 230.
[0095] Each category of aggregated data may include information related to a specific aspect of interaction data 202. For example, aggregated data section 230 may include aggregated account data 232, which may include data related to the accounts included in interaction data 202. Account-related data may include interaction history data, account balance data, account number, account age, account type, account holder data, account manager data, etc. Aggregated account data 232 may include historical data associated with the accounts involved in the interaction (e.g., primary account (PAN)). Aggregated account data 232 may include, for example, the total amount spent by the user associated with the account from the account last week.
[0096] The aggregated data portion 230 may also include aggregated resource provider data 234. Aggregated resource provider data 234 may include data associated with the resource providers involved in the interaction. For example, aggregated resource provider data 234 may include the number of interactions performed by the resource providers in the past day.
[0097] The aggregated data portion 230 may further include aggregated transport data 236. Aggregated transport data 236 may include data associated with a transport computer, which may be operated by an acquirer associated with a resource provider. Aggregated transport data 236 may include, for example, the number of interactions processed by the transport computer over the past month.
[0098] The aggregated data portion 230 may further include aggregated authorization entity data 238. Aggregated authorization entity data 238 may include data associated with an authorization entity computer, which can determine whether to authorize an interaction between a user and a resource provider. Aggregated authorization entity data 238 may include, for example, data used to determine whether to authorize an interaction.
[0099] As an example, interactions may involve multiple parties making final decisions (e.g., the authorizing entity may need to decide whether to authorize the interaction request or publish the authorized interaction, and the acquiring party may need to decide whether to clear the interaction). These decisions may also require additional aggregated data, such as historical data, related to the master account, the resource provider's computer, the transport computer, and the authorizing entity's computer. The computer can unify each interaction of different interaction types into a single feature vector format containing all the information.
[0100] III. Machine Learning Models
[0101] Predictions can be generated using interaction data and aggregated data. Computers can use interaction data and aggregated data to generate predictions about a specific task (e.g., predicting whether an interaction will be authorized for use in authorizing the interaction). Machine learning models can generate task-specific predictions based on interaction data and aggregated data.
[0102] A. Model Overview
[0103] Machine learning models can be neural network models, large language models, or other suitable types of models capable of generating predictions based on input. For example, a machine learning model can accept prompts and queries as input to determine task-specific predictions.
[0104] Figure 3 A block diagram illustrating the model inputs of a machine learning model according to an embodiment is shown. Figure 3 This includes model input 300, which contains prompts 302 and queries 304. Figure 3 It also includes model 306 and task-specific prediction 308.
[0105] Model input 300 may include data input to model 306 to generate task-specific predictions 308. A computer, such as evaluation computer 102, may generate model input 300. Model input 300 may include data in a uniform format, such as reference data. Figure 2 It is described by the eigenvectors.
[0106] Hint 302 may include a set of feature-label pairs. Each feature-label pair may include a task-specific label and a feature vector. The task-specific label may be a label for an associated feature vector that is specific to a particular task (e.g., a task in a machine learning model). Each feature vector in the set of feature-label pairs may include a feature vector containing an interactive data portion and an aggregated data portion.
[0107] For example, prompt 302 may include multiple interactive data portions, such as a first interactive data portion 310, a second interactive data portion 312, and a third interactive data portion 314 for different historical interactions. Prompt 302 may also include multiple aggregated data portions, such as a first aggregated data portion 316, a second aggregated data portion 318, and a third aggregated data portion 320. Each interactive data portion may be combined with an aggregated data portion to form a feature vector. For example, the first interactive data portion 310 may be combined with the first aggregated data portion 316 to form a first feature vector. Each feature vector in prompt 302 may correspond to a task-specific label. Prompt 302 may include a first task-specific label 322, a second task-specific label 324, and a third task-specific label 326. Prompt 302 may include any suitable number of feature vector entries and labels.
[0108] The first interactive data portion 310 and the first aggregated data portion 316 can correspond to the first task-specific label 322. The second interactive data portion 312 and the second aggregated data portion 318 can correspond to the second task-specific label 324. The third interactive data portion 314 can correspond to the third task-specific label 326.
[0109] Model input 300 also includes query 304. Query 304 may include data about the predictions the computer is generating. Query 304 may include a fourth interactive data portion 328, which combines with a fourth aggregated data portion 330 to form a fourth feature vector. The fourth feature vector may correspond to an unknown task-specific label 332. Query 304 does not include the task-specific label of the feature vector. The unknown task-specific label 332 may be a task-specific label that the computer is using to make predictions with model 306.
[0110] The interaction data in query 304 can include data involved in the target interaction. The interaction data in hint 302 can correspond to multiple interactions different from the target interaction. These multiple interactions can be interactions of the same interaction type as the target interaction. For example, if the target interaction is an authorization interaction, then the multiple interactions can be multiple authorization interactions. If the target interaction is a settlement interaction, then the multiple interactions can be multiple settlement interactions.
[0111] In some embodiments, prompt 302 may include data relating to the same user as the data in query 304. In other embodiments, prompt 302 may include data relating to the same resource as the resource indicated in query 304.
[0112] Model 306 can be a machine learning model (e.g., a neural network, a large language model, etc.) trained to generate inference such as task-specific prediction 308. Model 306 can generate a prediction for query 304 based on prompt 302 and query 304. Model 306 provides a context-based learning framework. The context of the model input 300 of model 306 in the context-based learning framework can be multiple interactions, which can be different types of interactions and have labels in prompt 302. Model 306 can also accept unlabeled query 304 as input.
[0113] Model 306 can learn from cue 302, which can be a labeled context, to generate a label for query 304. Based on cue 302 and query 304, model 306 can generate a task-specific prediction 308. (Reference) Figure 4 and Figure 5 More details about model 306 are described in more detail.
[0114] In some embodiments, task-specific prediction 308 may be a prediction of a smarter STIP (e.g., predicting whether an interaction will be approved or accepted), a prediction of a smarter posting (e.g., determining whether to post an interaction to an account), a prediction of a smarter card verification (e.g., predicting whether account verification will be successful), or a prediction of other interactions.
[0115] This type of model framework is applicable to a variety of new tasks without requiring adaptation. In contrast, previous methods required training different models for different tasks to adapt to them.
[0116] B. Training and Reasoning
[0117] Machine learning models can be trained and used for inference to generate task-specific predictions. Computers such as the Evaluation Computer 102 can train, maintain, and utilize machine learning models.
[0118] Figure 4 A flowchart of the training and inference method according to an embodiment is shown. Figure 4 It includes two phases for processing the machine learning model 418, namely the training phase 400 and the inference phase 410.
[0119] For training phase 400, at step 402, the computer can obtain data corresponding to multiple interactions. The computer can obtain multiple interaction data from a database. Each interaction data can be associated with a task-specific label. The interaction data can be stored in the database in association with the task-specific label.
[0120] As an illustrative example, multiple interaction data may include four interaction data instances, each corresponding to a task-specific label. Each interaction data may correspond to an interaction of a specific interaction category. For example, each of the four interaction data may correspond to an authorization interaction. The first interaction data may include data for an interaction between a user and a resource provider, allowing the resource provider to grant the user access to a resource. Each task-specific label may relate to the response status (e.g., result) of the corresponding interaction. For example, each task-specific label for an authorization interaction may be a value "authorized" or a value "unauthorized". As an example, two of the four task-specific labels may indicate "authorized", while the other two task-specific labels may indicate "unauthorized".
[0121] As another example, the interaction may include a user requesting access to a secure webpage. Interaction data may include data related to the user's access to the secure webpage (e.g., username, password, IP address, machine identifier code, etc.). Each of the four pieces of interaction data may correspond to an interaction category for the webpage access interaction. The first piece of interaction data may include data related to the interaction in which the user attempts to access the secure webpage. Each task-specific label may relate to the response status (e.g., result) of the corresponding interaction. For example, each task-specific label for the webpage access interaction may be a value of "Authorized Access" or a value of "Access Denied". As an example, three of the four task-specific labels may indicate "Authorized Access", while the fourth task-specific label may indicate "Access Denied". For a webpage access request interaction, the task may be to determine whether to authorize the user to access the secure webpage.
[0122] As another example, images can be classified during an interaction or during any other image classification task. Given an image, a set of task-specific labels could be "there's a dog in it" and "there's no dog in it," indicating whether a dog is in the image. Further, for the same image data, another set of task-specific labels could include "it's in color" and "it's black and white." For example, in some embodiments, the interaction could include a user requesting access to an image, transferring an image from one computer to another, uploading an image to a website, etc. The computer can utilize machine learning models to evaluate image data to determine labels to be applied to the image, text descriptions to be applied to the image, text titles to be applied to the image, whether the image violates one or more rules, whether the image has accessibility issues, and so on.
[0123] At step 404, after obtaining multiple interaction data instances and corresponding task-specific labels, the computer can combine each interaction data instance with aggregated data to create a feature vector. These multiple feature vectors can correspond to multiple task-specific labels used for training.
[0124] The computer may obtain aggregated data from an interactive database and / or from other computers and / or databases. Aggregated data may include data that is related to the interactive data in some way. For example, for the first interactive data, aggregated data may include information related to user accounts and resource providers and / or resource provider computers.
[0125] For a webpage access instance, aggregated data may include data representing historical webpage access attempts. Aggregated data may include the total number of all user access attempts to access the webpage, the number of user access attempts to access the webpage, an indication of where the webpage is hosted on the Internet, the number of access denial attempts, the number of access authorization attempts, and / or any other data related to the webpage, accessing the webpage, the host of the webpage (e.g., it may be a resource provider's computer), and / or the user's attempts to access the webpage.
[0126] Computers can combine webpage access interaction data with webpage access aggregate data to form webpage access feature vectors that can be used in machine learning model 418.
[0127] At step 406, after combining the interaction data and aggregated data using a uniform format of feature vectors, the computer can create training samples that include a cue 407, a query 408, and a target label 409 (e.g., the label to be predicted) as input to the current training step. The cue 407 may include multiple feature vectors and multiple task-specific labels for training. The query 408 may include interaction data for the interactions the computer is using to train the machine learning model 418 to predict task-specific labels.
[0128] As an example, prompt 407 may include four feature vectors created from interaction data, aggregated data, and task-specific tags related to webpage access interactions. Query 408 may include feature vectors created from webpage access interaction data and aggregated data, and may correspond to the target tag 409. Input to the model in a data sample may include five feature vectors.
[0129] In some embodiments, training data can be processed in batches. For example, if the batch size is 100, there may be 500 (5 × 100) feature vectors representing the total interactions of the batches.
[0130] The computer can input training samples into the model to train the machine learning model 418. The machine learning model 418 can accept training samples as input, where the training samples include prompts 407, queries 408, and target labels 409.
[0131] A computer can iteratively train a machine learning model 418 using a training set of training samples. The computer can create a first training set containing one or more training samples. The computer can use the first training set to train the machine learning model 418 in a first stage. The computer can then create a second training set for a second training stage. The computer can use the second training set to train the neural network in the second stage.
[0132] In some embodiments, model 418 may accept both training samples and inference samples as input and output a prediction (e.g., a few-shot training process), instead of training machine learning model 418 on some data samples and performing inference on some other data samples from the same task. Few-shot learning is a machine learning framework in which machine learning model 418 can learn to make accurate predictions by being trained on a small number of labeled instances.
[0133] For the inference phase 410, the computer may use a machine learning model 418, which may be trained during the training phase 400 to generate predictions (e.g., perform inference).
[0134] At step 412, the computer can obtain the current interaction data and the current aggregate data. The computer can obtain the current interaction data and the current aggregate data from the interaction database. The current interaction data can correspond to the current interaction for which the computer can generate predictions.
[0135] In some embodiments, the current interaction may be a recent interaction stored in an interaction database. In other embodiments, the current interaction may correspond to the interaction indicated by an interaction identifier in a prediction request message received from another device (e.g., a downstream device, a network processing computer, etc.).
[0136] As an illustrative example, current interaction data could be current secure webpage access interaction data, where a user is attempting to access a secure webpage. Current secure webpage access interaction data could include a username and password, which, for security reasons, could be hashed or otherwise obfuscated. Current aggregate data could be current secure webpage access aggregate data, which could indicate aggregated data about historical access requests to the webpage.
[0137] At step 414, the computer can use the current interaction data and the current aggregated data to create the current feature vector. The current feature vector can be a query of machine learning model 418.
[0138] For example, a computer can combine current secure webpage access interaction data and current secure webpage access aggregate data to form a query.
[0139] At step 416, the computer may determine to obtain and / or generate one or more additional feature vectors to use as a prompt. In some embodiments, the computer may use a context window to determine which feature vector(s) to use as the prompt. The computer may obtain a prompt containing multiple obtained feature vectors.
[0140] For example, a computer can obtain additional secure webpage access interaction data and additional secure webpage access aggregate data. This additional data may relate to other secure webpage access interactions. The computer can use this additional secure webpage access interaction data and the additional secure webpage access aggregate data to form additional feature vectors. The computer can use these additional feature vectors related to the additional secure webpage access interactions to generate prompts.
[0141] After receiving the prompts and queries, the computer can provide them to the machine learning model 418 to determine task-specific predictions in response to the prompts and queries. Task-specific predictions can specify the response state to the current interaction. For example, a task-specific prediction could indicate "authorized access" or "access denied" for the current secure webpage access interaction.
[0142] C. Context Window
[0143] The computer can obtain additional feature vectors from additional interaction data. Additional feature vectors can be obtained from the interactions included in the context window. The context window can instruct the computer how to determine which additional interaction data to utilize when making predictions about the current interaction. For example, the context window can indicate the length of time the computer can utilize when determining additional interaction data. The context window can indicate time lengths such as 1 hour, 1 day, 3 days, 1 week, 1 month, 4 months, 1 year, etc. The context window can indicate the number of interactions to include (e.g., 2 interactions, 3 interactions, 4 interactions, 7 interactions, 10 interactions, 20 interactions, etc.).
[0144] During model training and inference, the computer can use a context window to determine and obtain data for further interactions based on the corresponding interaction data. This additional interaction data, along with other aggregated data, can be input into the machine learning model.
[0145] In some embodiments, the context window may be small compared to the dataset size (e.g., 100 interactions within the context window size compared to 100 million total interactions stored in a database). The computer can select more informational interactions to put into the context window. However, it may be difficult or impossible to search all interaction data, or even 1% of the interaction data, to select several other comparable interactions similar to the current interaction. Embodiments may provide an efficient and effective method for selecting interactions for a context window.
[0146] Computers can obtain additional interaction data and additional aggregated data based on the criteria of context windows. For example, a computer can obtain k samples of interactions for each of n categories to best provide clues for learning a model in context, thereby making predictions about target samples (e.g., interaction data).
[0147] The sampling method can have two steps: 1) determining the feature-label rule set and 2) creating and using the context window.
[0148] During the determination of the feature-label rule set, for each task, the computer can compute feature-label rule pairs. To generate rules, the computer can: 1) evaluate the relationship between features and labels, 2) select the most relevant features, and 3) generate rules for each relevant feature.
[0149] To assess the relationship between features and labels, computers can perform feature-label correlation analysis to determine any correlation between feature vectors and labels. For example, a computer can assess the strength of the relationship (e.g., a linear relationship) between two variables and calculate their association. During correlation analysis, a computer can assess the change in one variable caused by a change in the other.
[0150] After determining the strength of the relationship between features and labels, the computer can select a threshold number of features for the rules. The computer can choose the most relevant features and labels. For example, the computer can select the top 25% or other suitable threshold percentages of feature-label relationships to generate rules for its rule set.
[0151] After selecting the features for which rules are to be generated, the computer can generate rules for the selected features. The selected features can belong to different types. For example, features can be numerical features or nominal features. The computer can generate rules for numerical features (e.g., greater than the mean plus standard deviation, less than the mean minus standard deviation, etc.). The computer can generate rules for nominal features (e.g., feature_type_set, which only includes types with a frequency of less than 30%).
[0152] Table 1 below shows an example task feature-label rule set that includes multiple rules. For example, a computer can identify five most relevant feature-label pairs. The most relevant features could be numeric feature 1, nominal feature 2, numeric feature 3, numeric feature 4, and numeric feature 5.
[0153] The computer can generate rules for each selected feature, indicating the allowed features to be selected during the use of a context window to obtain additional interactive data. For example, numeric feature 1 can have rules with values less than 0.9 or greater than 1.4. Nominal feature 2 can have rules with values equal to the "red" or "green" category. Numeric feature 3 can have rules with values less than 100 or greater than 400. Numeric feature 4 can have rules with values less than 9 or greater than 14. Numeric feature 5 can have rules with values equal to 2, 5, or 6.
[0154] As an illustrative example, the rules in Table 1 can relate to, for example, flower classification: feature 1 relates to petal size, feature 2 relates to color, feature 3 relates to petal weight in milligrams, feature 4 relates to the number of petals on each stem, and feature 5 relates to the number of petals on each flower.
[0155] Feature 1: <0.9 or >1.4 Feature 2: {'Red', 'Green'} Feature 3: <100 or >400 Feature 4: <9 or >14 Feature 5: {2, 5, 6}
[0156] Table 1: Task Feature-Label Rule Set
[0157] After generating the rule set, the computer can continue using a context window based on the rule set. To use the context window, the computer can sample data (e.g., interaction data and / or aggregated data) from the database according to the rule set. The computer can randomly sample interaction data from the database and determine whether to retain the interaction data based on the rule set.
[0158] The computer can continue to sample the interaction data until it obtains a predetermined number (k) of interaction data samples (e.g., 50 interaction data samples, 100 interaction data samples, 200 interaction data samples, etc.).
[0159] D. Model Components
[0160] Machine learning models can accept input data including prompts and queries, and can determine outputs including task-specific predictions. The internal processing of a machine learning model can include encoders and decoders that can modify the input data to determine the output. A machine learning model can be a neural network or a large language model that includes multiple encoders and decoders and a self-attention mechanism.
[0161] Machine learning models can receive input from a computer. Input can include prompts and queries. Output can be the predicted label of the queried interaction. Because machine learning models learn from previous tasks, it can be advantageous to have multiple tasks to back up and retrain the model. For example, a machine learning model can include a classification head and a regression head to help perform different tasks. Both the classification head and the regression head can be used to train the machine learning model.
[0162] Figure 5 A block diagram illustrating a machine learning model according to an embodiment is shown. Figure 5 The input data 502, the machine learning model 504, and the output 506 are shown.
[0163] Machine learning model 504 can accept input data 502 as input. Input data 502 includes prompts 508 and queries 510. Prompts 508 can include interactions with multiple entities (e.g., in...). Figure 5The described example includes data related to two interactions and corresponding labels for those interactions. Hint 508 may include feature vectors and corresponding task-specific labels, whereby the feature vectors include interaction data and aggregated data.
[0164] Query 510 may include data related to the interaction being queried against it. For example, a computer may use machine learning model 504 to generate predictive labels related to the interaction. Query 510 may include a feature vector that includes interaction data and aggregated data.
[0165] Output 506 can be a predicted label for the queried interaction. Output 506 can be a value indicating a prediction associated with the queried interaction. Output 506 can be a task-specific prediction. Task-specific predictions can specify the response status to the queried interaction. Response status can be, for example, "Authorized," "Unauthorized," "Cleared," "Not Cleared," "Fraudulent," "Non-Fraudulent," etc., depending on the interaction's category.
[0166] Machine learning model 504 may include multiple interactive encoders, including a first interactive encoder 512, a second interactive encoder 514, and a third interactive encoder 516. In some embodiments, each encoder may be pre-trained. In other embodiments, each encoder may not be pre-trained. The interactive encoders can process interactive data. Machine learning model 504 also includes a first label encoder 518 and a second label encoder 520. The label encoders can process task-specific labels. Machine learning model 504 may include any number of interactive encoders and label encoders. For example, machine learning model 504 may include 2 interactive encoders, 4 interactive encoders, 8 interactive encoders, 15 interactive encoders, etc. Machine learning model 504 may include the same number of label encoders as the number of interactive encoders. The interactive encoders can generate interactive embeddings, and the label encoders can generate label embeddings.
[0167] Each of multiple encoders can generate vectors (e.g., embeddings) from an input sequence (e.g., feature vectors, labels, etc.). Each encoder can include fully connected layers and can be a feedforward neural network. Each encoder can be pre-trained to generate embeddings based on the input vectors.
[0168] Machine learning model 504 may also include a first mixer encoder 522 and a second mixer encoder 524. Machine learning model 504 may include any number of mixer encoders. Machine learning model 504 may include the same number of mixer encoders as the number of interactive encoders. Each mixer encoder may accept interactive embeddings and label embeddings corresponding to feature-label pairs from input data 502. The mixer encoder may generate a hybrid embedding based on the interactive embeddings and label embeddings. In some embodiments, interactive embeddings and label embeddings may be cascaded together or otherwise combined to form a single input vector that can be input into the mixer encoder.
[0169] The machine learning model 504 also includes a self-attention module 526 and a decoder 528.
[0170] Machine learning model 504 also includes a classification head 530. The classification head 530 can be a classification machine learning model trained to determine a classification based on an input vector (e.g., an output vector from decoder 528). The classification machine learning model of classification head 530 can perform any classification process to classify the input data (e.g., decision tree, support vector machine, random forest, k-nearest neighbor, etc.).
[0171] Machine learning model 504 also includes a regression head 532. The regression head 532 can be a regression machine learning model trained to perform a regression task based on an input vector (e.g., the output vector from decoder 528). The regression machine learning model of the regression head 532 can perform any regression process to determine an output value based on the input data (e.g., linear regression, multinomial regression, ridge regression, elastic network regression, etc.).
[0172] As an illustrative example, the computer may obtain a prompt 508 and a query 510. The query 510 may include feature vectors for the target interaction. The prompt 508 may include a first additional feature vector and a second additional feature vector. The computer may input the feature vectors from the input data 502 into the interactive encoder in the machine learning model 504. For example, the computer may input the first additional feature vector from the prompt 508 into a first interactive encoder 512. The computer may input the second additional feature vector from the prompt 508 into a second interactive encoder 514. The computer may input the feature vector from the query 510 into a third interactive encoder 516.
[0173] Each interaction encoder can encode both interactive and aggregated data from the input feature vector. Each interaction encoder can generate an interaction embedding. Each interaction encoder can generate vectors representing the interactive and aggregated data. Each interaction encoder can be trained during the training phase to generate interaction embeddings representing the feature vectors used for the interaction.
[0174] The first interaction encoder 512 can generate a first interaction embedding from a first additional feature vector. The second interaction encoder 514 can generate a second interaction embedding from a second additional feature vector. The third interaction encoder 516 can generate a third interaction embedding from a feature vector corresponding to the target interaction.
[0175] The computer can also input task-specific labels from cue 508 into the label encoder in machine learning model 504. For example, the computer can input a first task-specific label corresponding to a first additional feature vector into a first label encoder 518. The computer can input a second task-specific label corresponding to a second additional feature vector into a second label encoder 520. The label encoder can encode the task-specific labels. The label encoder can generate label embeddings. Each label encoder can be trained during the training phase to generate label embeddings representing task-specific labels used for interaction.
[0176] Each interactive encoder in machine learning model 504 can be paired with a label encoder. For example, the first interactive encoder 512 can be paired with the first label encoder 518. The first interactive encoder 512 can encode the feature vector of the first feature-label pair from the cue 508. The first label encoder 518 can encode the label of the first feature-label pair from the cue 508.
[0177] After generating the interaction embedding and label embedding for each feature-label pair in cue 508, machine learning model 504 can combine the interaction embedding and label embedding for each feature-label pair. For example, machine learning model 504 can provide the first interaction embedding from the first interaction encoder 512 and the first label embedding from the first label encoder 518 to the first mixer encoder 522. Machine learning model 504 can also provide the second interaction embedding from the second interaction encoder 514 and the second label embedding from the second label encoder 520 to the second mixer encoder 524.
[0178] Each mixer encoder can accept two input vectors (e.g., an interaction embedding and a label embedding) and can generate a vector representing the combination of the two input vectors. Each mixer encoder can generate a mixed embedding from the received interaction embedding and the received label embedding. Each mixer encoder can be trained during the training phase to generate a mixed embedding representing the combination of the interaction embedding and the label embedding.
[0179] After obtaining the mixer embeddings (e.g., output encodings) from each mixer encoder (e.g., first mixer encoder 522 and second mixer encoder 524), the machine learning model 504 can provide the mixer embeddings to the self-attention module 526. The machine learning model 504 can provide the self-attention module 526 with the same number of mixer embeddings as the feature-label pairs present in the cue 508. A self-attention process can be performed on each mixer embedding.
[0180] The self-attention module 526 may include a self-attention mechanism for the neural network, such that each processed mixer embedding focuses on itself on each element in the mixer embedding vector. The self-attention module 526 may examine the relevance of each element in the mixer embedding to the other elements in the mixer embedding.
[0181] Typically, attention is a machine learning method that determines the relative importance of each component in a sequence relative to other components in the sequence. The sequence provided to the self-attention module 526 may include elements of each mixer embedding. The self-attention module 526 may determine a query matrix, a key matrix, and a value matrix based on the input mixer embedding. The self-attention module 526 may utilize the query matrix, key matrix, and value matrix to determine an attention score for each element in the input mixer embedding. For each input mixer embedding, the self-attention module 526 may output a vector representing the input mixer embedding modified according to the attention score (e.g., a weighted element in the input mixer embedding). The output of the self-attention module 526 may be a self-attention mixer embedding.
[0182] After generating the self-focused mixer embedding, the machine learning model 504 can provide the self-focused mixer embedding and a third interaction embedding corresponding to the target interaction to the decoder 528. The decoder 528 can accept multiple vectors as input and can generate a single output vector representing the input. The decoder 528 can generate an output vector representing the self-focused mixer embedding and interaction embedding (e.g., the third interaction embedding) used for the query. In some embodiments, the decoder 528 can be a flexible decoder.
[0183] The output vector can be fitted to either the classification head 530 or the regression head 532 for final prediction. The machine learning model 504 can determine which head to use based on the input data 502. If the task-specific label from the cue 508 is a binary label (e.g., authorized or unauthorized), the output vector can be fed into the classification head 530. If the task-specific label from the cue 508 is a quantity or other value (e.g., probability of something, number of interactions, risk score, etc.), the output vector can be provided to the regression head 532 for the final label.
[0184] The classification head 530 can be trained to determine classification based on vectors. The classification head 530 obtains an output vector from the decoder 528. The classification head 530 can determine the classification based on the output vector. For example, the classification head 530 can determine the probability that the output vector corresponds to one or more categories. One or more categories can include, for example, authorized, unauthorized, fraudulent, non-fraudulent, cleared, not cleared, posted, not posted, etc. The classification head 530 can output the category with the highest probability. The classification can be a task-specific prediction. A task-specific prediction can specify the response state to the current interaction.
[0185] The regression head 532 can be trained to determine values based on vectors. The regression head 532 can obtain an output vector from the decoder 528. The regression head 530 can determine the output value representing the output vector. The output value can be a value representing something related to the current interaction. The output value can be a task-specific prediction. The output value can be, for example, a risk score, a probability score, or other predicted values related to the current interaction and / or the response state of the current interaction.
[0186] This implementation provides the advantage of using the same machine learning model for different tasks, rather than needing to create custom machine learning models for each different task. If a classification task exists, the computer can reuse the entire model structure with the same regression head to perform any task, whether classification or regression. The computer can reuse or retrain the model, but can still utilize the overall model structure with minor differences for each head to perform each task.
[0187] E. Prediction Determination
[0188] Computers can use machine learning models to determine predictions for task-specific labels of interactions. These predictions can forecast the response state of the current interaction. For example, the current interaction could be an authorization interaction between a user and a resource provider, where the user is requesting authorization to access resources offered by the provider. The response state can indicate whether the authorization interaction has been authorized.
[0189] Figure 6 A flowchart illustrating a prediction determination method according to an embodiment is shown. Figure 6The methods shown can be performed by a computer such as the evaluation computer 102.
[0190] In some embodiments, prior to step 602, the evaluation computer 102 may receive a prediction request from the downstream device 106. The prediction request may be a task-specific request for a particular interaction. The prediction request may include an interaction identifier that uniquely identifies the interaction. The evaluation computer 102 may use the interaction identifier to identify interaction data and related data in the interaction database 104.
[0191] At step 602, the evaluation computer 102 can obtain interaction data and aggregated data for the current interaction. The evaluation computer 102 can obtain the interaction data and aggregated data from the interaction database 104. The interaction data and aggregated data can be used for the current interaction. The current interaction may include a device (e.g., a user device) requesting access to resources from a resource provider's computer.
[0192] The evaluation computer 102 can obtain interaction data associated with a specific interaction identifier. The evaluation computer 102 can also search the interaction database 104 for data related to the devices, entities, and / or accounts indicated in the interaction data to obtain aggregated data of devices, entities, and / or accounts.
[0193] At step 604, after obtaining the interaction data and aggregated data, the evaluation computer 102 can use the interaction data and aggregated data to generate feature vectors. The feature vectors can adopt a uniform format, enabling the generation of feature vectors for any type of interaction. The evaluation computer 102 can use the interaction data and aggregated data to generate feature vectors, as shown in reference... Figure 2 A more detailed description.
[0194] At step 606, after generating the feature vector, the evaluation computer 102 can generate a cue containing a set of feature-label pairs. The evaluation computer 102 can obtain multiple additional interactive data and multiple additional aggregated data from the interactive database 104.
[0195] Each feature-label pair may include a task-specific label and an additional feature vector. The additional feature vector may include additional interaction data and additional aggregate data. The additional interaction data and the additional aggregate data are associated with other interactions that have the same classification type as the current interaction.
[0196] At step 608, after generating the prompt, the evaluation computer 102 can use a machine learning model trained to determine predictions for a specific task (e.g., Figure 5The machine learning model (504) described herein is loaded into memory. The machine learning model can be a neural network or a large language model used to generate predictions based on input prompts and queries. The machine learning model can be a pre-trained machine learning model.
[0197] In some embodiments, the evaluation computer 102 may store the machine learning model in a data storage device and load the machine learning model from the data storage device. In other embodiments, the machine learning model may be stored in a model database. The evaluation computer 102 may retrieve the machine learning model from the model database.
[0198] At step 610, after the machine learning model is loaded into memory, the evaluation computer 102 can input prompts and queries, including feature vectors, into the machine learning model.
[0199] At step 612, the evaluation computer 102 can use a machine learning model to determine task-specific predictions for the query and prompt. The task-specific prediction can specify the response state to the interaction. For example, the current interaction could be an authorized interaction. The interaction in the prompt can also be an authorized interaction. The response state to the interaction can indicate whether the current interaction is authorized. Therefore, the evaluation computer 102 can generate a prediction of whether the current interaction is authorized based on the current interaction and other interactions.
[0200] F. Training Mask
[0201] Masks can be applied to training data. Training data can involve interactions of a specific interaction category (e.g., authorization interactions, liquidation interactions, etc.) and can involve one of many different learning tasks, which can be regression or classification tasks. Instances in the training data can represent the learning task of a machine learning model. Masks help the machine learning model understand the current task during the training phase.
[0202] Figure 7 A block diagram illustrating masking uniform data (e.g., masking feature vectors) according to an embodiment is shown. The learning objective and pre-trained machine learning model will be described in the context of reference. Figure 7 .
[0203] Figure 7This demonstrates that uniform data 702 (e.g., feature vectors) can be masked in two different ways. Uniform data 702 can be masked into first masked uniform data 704 and second masked uniform data 706. Specific fields in uniform data 702 can be masked to create first masked uniform data 704 and second masked uniform data 706. A first field 708 can be masked in uniform data 702 to obtain first masked uniform data 704. A second field 710 can be masked in uniform data 702 to obtain second masked uniform data 706. A computer can mask fields in the interactive data portion of the feature vector.
[0204] Before being fed into a machine learning model for training, a computer can mask fields of a feature vector. Prior to training, the computer may have masked interactive data, aggregated data, and task-specific labels for specific instances used in the training data.
[0205] Computers can mask specific interactive data in various ways (e.g., masking different parts of the interactive data). For example, a computer can generate first masked unified data 704 and second masked unified data 706 from the same interactive data. The computer can use the two instances of masked data to train a machine learning model, allowing the model to optimize a loss function that takes the mask into account. The computer can verify that the output of the machine learning model for the first masked unified data 704 is similar to that for the second masked unified data 706. In some embodiments, the computer can also train a machine learning model on unified data 702 and can compare the output from the machine learning model with the output of the masked input based on unified data 702. Therefore, the computer can determine a loss value from unified data 702 and can compare loss values determined from the masked data to help optimize the loss function.
[0206] As an illustrative example, a computer can train a machine learning model using a mask based on the pseudocode shown in Table 2.
[0207]
[0208] Table 2: Pre-trained pseudocode
[0209] It is the standard prediction loss, and This is a contrast loss. r includes interactive data x and labels y. r can be extended to a unified data format u, which includes interactive data x and aggregated data x. expand And label y. m are multiple masks. CE represents cross-entropy loss. Indicates the model's predicted label for feature vector u, y u The predicted label indicating the feature vector u. The predicted label of the first masking feature vector u is indicated, and The predicted label indicates the second masking feature vector u.
[0210] Labels can correspond to specific tasks and can be associated with interactive data. Tasks can be divided into two different groups: classification tasks and regression tasks.
[0211] Classification tasks can include tasks such as the Smarter VisaNet series of tasks (e.g., Smarter Posting, Smarter Account Verification, Smarter STIP, etc.). For classification tasks, the computer can select a set of nominal fields, use those fields as labels, and change those fields to default values in the uniform data section.
[0212] For regression tasks, the computer can select a set of numeric fields, use those fields as labels, and change those fields to default values in the uniform data section.
[0213] Referring to Table 2, for each training step in the training phase, the computer can use multiple batches of training data to train the machine learning model. During each training batch, the computer can sample the specific task t to be trained. The computer can randomly sample task t from multiple tasks.
[0214] A computer can obtain multiple data samples of a specific interaction (r = (x, y)) from an interaction database. A computer can obtain interaction data x and task-specific labels y for x. A computer can obtain any number of pairs of interaction data x and task-specific labels y from an interaction database associated with task t.
[0215] After obtaining the data for the sampled interactions, the computer can generate a feature vector u for each sampled interaction. The computer can use the interaction data x, the task-specific label y, and the aggregated data xx related to the sampled interactions. expand This generates a feature vector u. The computer can then generate a feature vector dataset U from the feature vectors.
[0216] After forming the feature vector dataset U, the computer can generate one or more masking feature vectors for each feature vector u in the feature vector dataset U. For example, the computer can generate two masking feature vectors from each feature vector u (e.g., and ).
[0217] After generating one or more masking feature vectors for each feature vector u in the feature vector dataset U to form a set of masking feature vectors, the computer can use said set of masking feature vectors to train a machine learning model. The computer can use the machine learning model to generate predicted labels for feature vector u (e.g., ), the predicted label of the first masking feature vector u (e.g., ) and the predicted label of the second masking feature vector u (e.g., ).
[0218] A computer can optimize a loss function using different predicted labels, which instructs the machine learning model on the accuracy of generating predicted labels. Based on the loss function, the computer can update the weights of the entire machine learning model to predict the next predicted label more accurately. The computer can iteratively train the machine learning model using feature vectors and masked feature vectors.
[0219] IV. Advantages
[0220] The embodiments disclosed herein offer several advantages. For example, a machine learning model can learn from several instances of the same task and make predictions about that task, rather than learning from data on the same type of task (e.g., learning only from authorized interactions). By doing so, the model can be seamlessly used for a variety of downstream tasks, which may include existing or new tasks.
[0221] The embodiments of this disclosure have several additional advantages. For example, machine learning models can be trained by generating uniform data and using it as feature vectors at different interaction stages, thereby allowing the use of data from more different interactions in the machine learning model. This also eliminates the need to train many different machine learning models for each different interaction stage.
[0222] Furthermore, current methods require different machine learning models to perform different types of predictions (e.g., predicting whether an interaction is fraudulent, whether an interaction will be approved, whether an interaction will be posted to the user's account as a pending event or as another state, etc.). To improve the performance of all these different machine learning models, more data and larger models are often used. However, more data and larger models may not improve model performance due to the interaction data used, as predicting arbitrary numbers is unrealistic. As an illustrative example, in US restaurant interactions, regardless of what information the model may have about the restaurant so far, even after authorization, it may be difficult to predict what the settlement amount might be, because the model may not have any information about the restaurant's tip. Therefore, even if a large model is trained using a large number of authorization-settlement interactions, it is difficult to obtain a model with good performance in this case. Therefore, simply using more data and larger models may not improve the performance of interaction data. The embodiments can provide a single machine learning model that can perform multiple different tasks, rather than requiring the training of many different models on a large amount of data. The embodiments do not simply rely on using a large amount of data to improve prediction accuracy as in previous methods.
[0223] Machine learning models can also adjust for data distribution shifts, especially label distribution shifts. For example, when several instances in a prompt can reveal a pattern, a machine learning model can make good predictions even if the labels shift over time.
[0224] V. Computer Systems
[0225] The processing and methods described herein can be executed by a computer system.
[0226] Figure 8 A block diagram of an evaluation computer 102 according to an embodiment is shown. The exemplary evaluation computer 102 may include a processor 804. The processor 804 may be coupled to a memory 802, a network interface 806, and a computer-readable medium 808. The computer-readable medium 808 may include any number of modules. The computer-readable medium 808 may include a feature vector module 808A, a prompting module 808B, a query module 808C, and a machine learning module 808D.
[0227] Memory 802 can be used to store data and code. For example, memory 802 can store training data, machine learning models, machine learning model weights, inference data, etc. Memory 802 can be coupled internally or externally to processor 804 (e.g., a cloud-based data storage device) and can contain any combination of volatile and / or non-volatile memory (such as RAM, DRAM, ROM, flash memory, or any other suitable memory device).
[0228] Computer-readable medium 808 may contain code executable by processor 804 for performing a method. For example, a method may include an evaluation computer 102 obtaining interaction data and aggregate data for a current interaction. The current interaction may involve means of requesting access to a resource. The evaluation computer 102 may use the interaction data and aggregate data to generate feature vectors. The evaluation computer 102 may then generate a cue containing a set of feature-label pairs, each feature-label pair containing a task-specific label and an additional feature vector. The additional feature vectors may include additional interaction data and additional aggregate data. The additional interaction data and additional aggregate data may be associated with other interactions having the same classification type as the current interaction. The evaluation computer 102 may load a pre-trained machine learning model into memory, the pre-trained machine learning model being trained to determine predictions for a specific task. The evaluation computer 102 may input a cue and a query including the feature vectors into the pre-trained machine learning model. The evaluation computer 102 may use the pre-trained machine learning model to determine task-specific predictions for the query and cue. Task-specific predictions can specify the response status to the current interaction (e.g., authorized or unauthorized, liquidated or unliquidated, fraudulent or non-fraudulent, etc.).
[0229] The feature vector module 808A may contain code or software executable by the processor 804 for creating feature vectors. The feature vector module 808A, in conjunction with the processor 804, can generate feature vectors from interactive and aggregated data.
[0230] The prompt module 808B may include code or software that can be executed by the processor 804 to create prompts. The prompt module 808B, in conjunction with the processor 804, can create prompts from additional interactive data and additional aggregated data to form a set of feature-label pairs.
[0231] The query module 808C may include code or software that can be executed by the processor 804 to create queries. The query module 808C, in conjunction with the processor 804, can generate queries from feature vectors.
[0232] The machine learning module 808D may include code or software executable by the processor 804 for training, maintaining, and utilizing machine learning models. The machine learning module 808D, in conjunction with the processor 804, can train and utilize machine learning models, which may be neural networks or large language models, capable of generating task-specific predictions for the current interaction based on queries and prompts.
[0233] Network interface 806 may include an interface that allows evaluation computer 102 to communicate with external computers. Network interface 806 enables evaluation computer 102 to transfer data to and from another device (e.g., interactive database 104, downstream device 106, network processing computer 108, etc.). Some examples of network interface 806 may include a modem, a physical network interface (such as an Ethernet card or other network interface card (NIC)), a virtual network interface, a communication port, a Personal Computer Memory Card International Association (PCMCIA) slot, and cards, etc. Wireless protocols enabled by network interface 806 may include Wi-Fi. TM Data transmitted through network interface 806 may take the form of signals, which may be electrical signals, electromagnetic signals, optical signals, or any other signals that can be received by an external communication interface (collectively, "electronic signals" or "electronic messages"). These electronic messages, which may contain data or instructions, may be provided between network interface 806 and other devices via a communication path or channel. As described above, any suitable communication path or channel may be used, such as wires or cables, optical fibers, telephone lines, cellular links, radio frequency (RF) links, WAN or LAN networks, the Internet, or any other suitable medium.
[0234] Any computer system mentioned in this article can utilize any suitable number of subsystems. Figure 9 An example of such a subsystem in a computer system 900 is illustrated. In some embodiments, the computer system includes a single computer device, wherein the subsystem may be a component of the computer device. In other embodiments, the computer system may include multiple computer devices having internal components, each of which is a subsystem. The computer system may include desktop and laptop computers, tablet computers, mobile phones, and other mobile devices.
[0235] Figure 9 The subsystems shown are interconnected via system bus 924. Other subsystems are shown, such as printer 908, keyboard 916, storage device 918, monitor 922 (e.g., display screen, such as LED) coupled to display adapter 912, etc. Peripheral devices and input / output (I / O) devices coupled to I / O controller 902 can be connected via any number of input / output (I / O) ports 914 (e.g., USB, etc.). Devices known in the art, such as I / O port 914 or external interface 920 (e.g., Ethernet, Wi-Fi, etc.), can be used to connect computer system 900 to a wide area network such as the Internet, a mouse input device, or a scanner. Interconnection via system bus 924 allows central processing unit 906 to communicate with each subsystem and control the execution of multiple instructions from system memory 904 or storage device 918 (e.g., a fixed disk, such as a hard disk drive or optical disk) and the exchange of information between subsystems. System memory 904 and / or storage device 918 may be embodied in a computer-readable medium. Another subsystem is a data collection device 910, such as a camera, microphone, accelerometer, etc. Any data mentioned herein can be output from one component to another and can be output to a user.
[0236] A computer system may include multiple identical components or subsystems connected together, for example, via an external interface 920, an internal interface, or a removable storage device that can be connected and removed from one component to another. In some embodiments, the computer system, subsystem, or device may communicate via a network. In such cases, one computer may be considered a client, and another computer may be considered a server, where each computer may be part of the same computer system. Clients and servers may each include multiple systems, subsystems, or components. In various embodiments, the method may involve a variety of numbers of clients and / or servers, including at least 10, 20, 50, 100, 200, 500, 1,000, or 10,000 devices. The method may include a variety of numbers of communication messages between devices, including at least 100, 200, 500, 1,000, 10,000, 50,000, 100,000, 500,000, or one million communication messages. Such communication can involve at least 1MB, 10MB, 100MB, 1GB, 10GB, or 100GB of data.
[0237] Although the steps in the flowcharts and process flows above are shown or described in a specific order, it should be understood that embodiments of the invention may include methods with steps in a different order. Furthermore, steps may be omitted or added, and may still be included in embodiments of the invention.
[0238] Various aspects of the embodiments may be implemented using hardware circuitry (e.g., application-specific integrated circuits or field-programmable gate arrays) and / or using computer software stored in a modular or integrated manner in memory with a general-purpose programmable processor in the form of control logic, and thus the processor may include memory storing software instructions for configuring the hardware circuitry, and an FPGA or ASIC having configuration instructions. As used herein, the processor may include a single-core processor, a multi-core processor on the same integrated chip, or multiple processing units on a single circuit board or networked, as well as dedicated hardware. Based on the disclosure and teachings provided herein, those skilled in the art will know and understand other ways and / or methods of implementing the embodiments of this disclosure using hardware and combinations of hardware and software.
[0239] Any software component or function described in this application may be implemented as software code executed by a processor using any suitable computer language such as Java, C, C++, C#, Objective-C, Swift, or a scripting language such as Perl or Python, employing techniques such as conventional or object-oriented methods. The software code may be stored as a series of instructions or commands on a computer-readable medium for storage and / or transmission. Suitable non-transitory computer-readable media may include random access memory (RAM), read-only memory (ROM), magnetic media such as hard disk drives or floppy disks, or optical media such as optical discs (CDs) or DVDs (Digital Universal Optical Discs) or Blu-ray discs, flash memory, etc. The computer-readable medium may be any combination of such devices. Furthermore, the order of operations may be rearranged. A process may be terminated upon completion of its operations, but may include additional steps not included in the figures. A process may correspond to a method, function, program, subroutine, subroutines, etc. When a process corresponds to a function, its termination may correspond to the function returning to the calling function or the main function.
[0240] Such programs can also be encoded and transmitted using carrier signals suitable for transmission over wired, optical, and / or wireless networks conforming to various protocols, including the Internet. Therefore, computer-readable media can be created using data signals encoded with such programs. Computer-readable media encoded with program code can be packaged with a compatible device (e.g., as firmware) or provided separately from other devices (e.g., for download via the Internet). Any such computer-readable media can reside on or within a single computer product (e.g., a hard disk drive, CD, or an entire computer system) and can exist on or within different computer products within a system or network. A computer system may include a monitor, printer, or other suitable display for providing a user with any of the results mentioned herein.
[0241] Any method described herein can be performed wholly or partially by a computer system including one or more processors that can be configured to perform the steps. Any operation performed by a processor can be performed in real time. The term "real time" can refer to a computational operation or process completed within a specific time constraint. As an example, the time constraint can be 30 seconds, 1 minute, 10 minutes, 30 minutes, 1 hour, 4 hours, 1 day, or 7 days. Therefore, embodiments can relate to a computer system configured to perform the steps of any method described herein, possibly having different components that perform the respective steps or groups of respective steps. Although presented as numbered steps, the steps of the methods herein can be performed simultaneously or at different times or in different orders. Additionally, portions of these steps can be used in conjunction with portions of other steps of other methods. Furthermore, all or part of the steps can be optional. Additionally, any step of any method can be performed using other means of modules, units, circuits, or systems for performing these steps.
[0242] Specific details of particular embodiments may be combined in any suitable manner without departing from the spirit and scope of embodiments of this disclosure. However, other embodiments of this disclosure may relate to specific embodiments relating to each individual aspect, or specific combinations of such individual aspects.
[0243] The foregoing description of exemplary embodiments of this disclosure has been presented for purposes of illustration and description. It is not intended to be exhaustive or to limit this disclosure to the precise forms described, and many modifications and variations are possible in accordance with the teachings above.
[0244] Unless explicitly stated otherwise, the use of “a / an” or “the” is intended to mean “one or more”. Unless explicitly stated otherwise, the use of “or” is intended to mean “inclusive or” rather than “exclusive or”. A reference to a “first” component does not necessarily require the provision of a second component. Furthermore, unless explicitly stated otherwise, a reference to a “first” or “second” component does not limit the referenced component to a particular location. The term “based on” is intended to mean “at least partially based on”.
[0245] Claims can be drafted to exclude any element that may be optional. Therefore, this statement is intended to serve as a basis for the use of exclusive terms such as “only,” “merely,” etc., or the use of “negative” in relation to the description of the elements of the claims.
[0246] All patents, patent applications, publications, and descriptions mentioned herein are incorporated herein by reference in their entirety for all purposes. None of them are considered prior art. In the event of any conflict between this application and the references provided herein, this application shall prevail.
Claims
1. A method comprising: The computer obtains interaction data and aggregated data for the current interaction, wherein the current interaction relates to means of requesting access to a resource; The computer generates feature vectors using the interactive data and the aggregated data; The computer generates a prompt containing a set of feature-label pairs, each feature-label pair containing: Task-specific tags, and The additional feature vectors include additional interaction data and additional aggregate data, wherein the additional interaction data and the additional aggregate data are related to additional interactions that have the same classification type as the current interaction. A pre-trained machine learning model is loaded into the computer's memory, and the pre-trained machine learning model is trained to determine predictions for a specific task. The prompt and the query including the feature vector are input into the pre-trained machine learning model; as well as The computer uses the pre-trained machine learning model to determine task-specific predictions for the query and the prompt, the task-specific predictions specifying the response state to the current interaction.
2. The method of claim 1, wherein the pre-trained machine learning model comprises a neural network model.
3. The method of claim 1, further comprising: The computer generates the query using the feature vector.
4. The method of claim 1, further comprising: The computer obtains multiple historical interaction data, multiple historical aggregate data, and multiple historical task-specific tags from multiple historical interactions. The computer uses the multiple historical interaction data, the multiple historical aggregate data, and the multiple historical task-specific labels to generate multiple historical feature vectors; The computer uses the multiple historical feature vectors to generate historical prompts and historical queries; and The computer uses the historical prompts and historical queries to train the pre-trained machine learning model.
5. The method of claim 1, wherein the pre-trained machine learning model comprises a plurality of interactive encoders, a plurality of task-specific label encoders, and a plurality of mixer encoders.
6. The method of claim 5, wherein each of the plurality of mixer encoders combines encodings created from a subset of the plurality of interactive encoders and a subset of the plurality of task-specific tag encoders.
7. The method of claim 5, wherein determining the task-specific prediction comprises: The computer uses the pre-trained machine learning model to generate the current interaction code from the feature vector; The computer uses the pre-trained machine learning model to generate a set of interactive codes from the additional feature vectors of each feature-label pair in the set of feature-label pairs; The computer uses the pre-trained machine learning model to generate a set of label codes from the task-specific labels of each feature-label pair in the set of feature-label pairs; and The computer uses the pre-trained machine learning model to generate a set of hybrid codes from the set of interactive codes and the set of label codes corresponding to the same feature-label pairs in the set of feature-label pairs.
8. The method of claim 7, further comprising: The computer uses the pre-trained machine learning model to determine a set of attention scores based on the set of mixed codes using a self-attention process; The computer uses the pre-trained machine learning model to modify the set of hybrid codes using the self-attention process and the set of attention scores; The computer uses the pre-trained machine learning model to determine the output vector based on the current interactive encoding and the set of modified hybrid encodings; and The computer uses the pre-trained machine learning model to determine task-specific predictions based on the current task using a regression head or a classification head.
9. The method of claim 8, wherein the current task is indicated by the task-specific label of each feature-label pair in the set of feature-label pairs.
10. The method of claim 1, wherein the response status is authorized or unauthorized.
11. The method of claim 1, wherein the interaction data includes an indication of an interaction type among a plurality of interaction types.
12. The method according to claim 1, wherein the aggregated data includes aggregated account data and aggregated resource provider data.
13. The method of claim 1, wherein the interaction data includes an indication of an interaction type among a plurality of interaction types, wherein generating the feature vector using the interaction data and the aggregated data comprises: The computer generates the interactive data portion of the feature vector, wherein the interactive data portion includes multiple interaction type entries; The computer sets the value of the interaction type entry corresponding to the interaction type from the plurality of interaction type entries to the value of the interaction data; and The computer sets the values of other interaction type entries that do not correspond to the interaction type among the plurality of interaction type entries to default values.
14. The method of claim 1, wherein the pre-trained machine learning model comprises a plurality of interactive encoders, a plurality of task-specific label encoders, and a plurality of mixer encoders, wherein the number of mixer encoders in the plurality of mixer encoders is the same as the number of feature-label pairs in the set of feature-label pairs.
15. The method of claim 1, wherein the method further comprises: The computer iteratively trains the pre-trained machine learning model using historical interaction data, historical aggregated data, historical task-specific labels, and training masks.
16. The method of claim 1, wherein the current interaction belongs to the category of authorized interactions, and wherein the other interaction belongs to the category of authorized interactions.
17. The method of claim 1, wherein obtaining the interaction data and the aggregated data for the current interaction comprises: The computer obtains the interactive data and the aggregated data from the interactive database, wherein the network processing computer stores the interactive data in the interactive database.
18. The method of claim 1, wherein after determining the task-specific prediction, the method further comprises: The evaluation computer generates a prediction message containing task-specific predictions; and The evaluation computer provides the prediction message to downstream devices.
19. A computer product comprising a computer-readable medium storing a plurality of instructions for controlling a computer system to perform any of the methods described above.
20. A system comprising: The computer product according to claim 19; and One or more processors, the one or more processors being configured to execute instructions stored on the computer-readable medium.