Machine learning system
The machine learning system with a category history module addresses the challenges of fraudulent transaction detection by processing transaction data through a category history module that updates state data based on time differences, achieving improved detection performance and accuracy.
Patent Information
- Application Number
- JP2024563395
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-04-28
- Publication Date
- 2025-05-27
AI Technical Summary
Existing systems for detecting fraudulent transactions in digital payment systems face challenges such as high false positive rates, low accuracy, and the inability to process large volumes of transactions in real-time, especially in environments with siloed data systems.
A machine learning system with a category history module that processes transaction data by storing state data for categories associated with entities, applying a decay function to update this data based on time differences, and using neural networks to generate scalar values indicating the likelihood of transaction anomalies.
The system improves detection performance and reduces false positives by learning subtle interaction patterns, enabling faster and more accurate fraud detection in high-volume transaction processing environments.
Smart Images

Figure 2025516199000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates in particular to a categorical history module for a machine learning system for detecting anomalies in data patterns, for example for use in a machine learning system for detecting fraudulent transactions. Some examples relate to machine learning systems for use in real-time transaction processing.
Background Art
[0002] Digital payments have increased rapidly over the past 20 years, and more than three-quarters of all payments worldwide are made using some form of payment card or e-wallet. Point-of-sale systems are becoming increasingly digital rather than cash-based. Put simply, the world's system of commercial transactions now relies heavily on electronic data processing platforms. This presents many engineering challenges, mainly hidden from the average user. For example, digital transactions need to be completed in real time, i.e., with the minimum level of delay received by a computer device at the time of purchase. Digital transactions also need to be secure and resistant to attacks and exploitation. The processing of digital transactions is also constrained by the historical development of the world's electronic systems for payments. For example, many infrastructures are still configured around models designed for mainframe architectures that were used more than 50 years ago.
[0003] As digital transactions increase, new security risks are emerging. Digital transactions present new opportunities for fraud and malicious activities. In 2015, it was estimated that 7% of digital transactions were fraudulent and that this figure would only increase with the shift to more online economic activities. Fraud losses are estimated to be four times the world's population (e.g., in US dollars) and are increasing.
[0004] Traditional methods of preventing fraud, such as the authentication of identifying information (e.g., passwords, digital biometrics, national IDs, etc.), have proven ineffective in preventing fraud vectors such as synthetic identity information and fraud, so financial service institutions are increasingly subject to more regulatory scrutiny. These far more complex threat vectors for fraud require significantly more analysis in extremely short (50 sub-millisecond) times and are often based on much smaller data sampling sizes for the fraud or fraud itself. This poses significant technical challenges.
[0005] Risks such as fraud are economic issues for businesses involved in commerce, while the implementation of technical systems for processing transactions is an engineering challenge. Conventionally, banks, merchants, and card issuers have established "paper" rules or procedures that were manually implemented by clerks to flag or block some transactions. Since transactions have become digital, one approach to building a technical system for processing transactions is to supply these sets of criteria established by computer engineers and have them implemented using the digital representation of the transaction, i.e., convert the handwritten rules into codified logical statements that can be applied to electronic transaction data, and rely on computer engineers to do so. This conventional approach has run into several problems as the volume of digital transactions has grown. First, any applied processing needs to be done "in real time," e.g., with a latency of milliseconds. Second, thousands of transactions per second need to be processed (e.g., a typical "load" can be 1000 - 2000 per second), and the load varies unexpectedly over time (e.g., the launch of a new product or set of tickets can easily multiply the average load level several times). Third, transaction processors and banks' digital storage systems are often siloed or segmented for security reasons, and digital transactions often involve an interconnected web of merchant systems. Fourth, large-scale analysis of actual reported and predicted fraud is now possible. This indicates that conventional methods of fraud detection have not reached the level, have low accuracy, and have a high false positive rate. This then has a physical impact on digital transaction processing, more legitimate points of sale and online purchases are rejected, and those trying to exploit the new digital system often get away with it.
[0006] In recent years, more machine learning-based approaches have been incorporated into the processing of transaction data. As machine learning models mature in the academic world, engineers have begun to attempt to apply them to the processing of transaction data. However, this again runs into problems. Even when engineers are provided with and asked to implement academic or theoretical machine learning models, this is not easy. For example, problems arise in large-scale transaction processing systems. Machine learning models do not have the luxury of unlimited inference time like in a laboratory. This means that it is simply not practical to implement some models in a real-time setting, or that they require significant adaptation to enable real-time processing at the levels of volume that real-world servers encounter. Moreover, engineers need to address the problem of implementing machine learning models in siloed or partitioned data, based on access security and in situations where the rate of data updates is extreme. Therefore, the problems faced by engineers building transaction processing systems may be seen as similar to those faced by network or database engineers, and machine learning models need to be applied while meeting the system throughput and query response time constraints set by the processing infrastructure. There is no easy solution to these problems. In fact, the fact that many transaction processing systems are confidential, proprietary, and based on old technologies means that engineers often do not have the body of knowledge developed in these neighboring fields and often face challenges specific to the field of transaction processing. Moreover, the field of large-scale and practical machine learning is still shallow, and there are few established design patterns or textbooks that engineers can rely on.
[0007] The fraud detection system creates risk scores for new transactions in real time (i.e., latency ≤ 100 milliseconds), where a high risk score indicates the likelihood of fraud, error, or unauthorized use of an account. Financial institutions use these scores within their decision logic when determining whether to approve a transaction (typically, this will result in rejecting the transaction if the risk score is higher than some threshold). Financial institutions generally desire that a set of transactions above a certain risk threshold contains the highest possible proportion of fraud (i.e., to provide the “maximum detection rate”) with the lowest possible rate of false positives classification (i.e., to provide “maximum accuracy”).
[0008] One approach known in the art per se involves using machine learning techniques such as recurrent neural networks (RNNs) to detect anomalies in patterns of behavior. It will be understood by those skilled in the art that the output of the RNN is the next state, which depends on the previous state and the input given to the RNN. In this way, the state of the RNN is based not only on the instantaneous information related to a given transaction but also on the history of previous transactions and thus retains a memory of the observed patterns. This can be used to detect abnormal behavior that may indicate a fraudulent transaction. The terms “behavior” and “behavioral” are used herein to refer to patterns of activity or actions.
[0009] This detection of abnormal activity can provide an improvement in the security of the system compared to “conventional” methods of fraud prevention (e.g., passwords, digital biometrics, national IDs, etc.). It will of course be understood that behavioral pattern analysis techniques can be used in combination with one or more of these “conventional” methods as appropriate.
[0010] Generally, a given machine learning system can be configured to handle transactions related to a number of different entities, such as different cardholders. The pattern of behavior for each entity may be different, and each history generally must be maintained independently to detect behavior that is abnormal for that entity. Thus, a machine learning system for a high-volume transaction processing pipeline can generally store state data for each entity (i.e., cardholder).
[0011] The applicant understands that the duration between consecutive transactions for a given entity can often vary significantly. Thus, when determining the validity of a new transaction, it is the degree of influence that the previous state should have. Previous patent applications by the applicant, published as WO / 2022 / 008130 and WO / 2022 / 008131, each of which is incorporated herein by reference, describe a configuration in which state data for a given entity is decayed based on the time period elapsed since that state data was stored (i.e., since the last transaction for that entity).
[0012] Data scientists building machine learning models typically struggle to design those models and extract machine learning features for them that result in the best trade-off between detection rate and accuracy.
[0013] However, the applicant understands that further improvements can be made to the technical system for performing related analyses. Specifically, the applicant understands that the novelty of the interaction of entities, for example, the novelty of the relationship between an account holder and a specific merchant, is useful when detecting abnormal behavior. The novelty of the buyer-seller interaction is an important signal for maximizing the performance of these scores. An established relationship is a good indicator that a new activity is genuine (i.e., honestly executed by the buyer), but a new relationship indicates greater suspicion regarding new activities (which can typically be taken together with other risk signals). SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM
[0014] The present invention aims to provide, for example, an improvement to a machine learning system for automatically extracting the novelty of interactions between a buyer and a seller.
[0015] According to a first aspect, an embodiment of the present invention is a machine learning system for processing data corresponding to incoming transactions associated with an entity, the machine learning system comprising a category history module configured to process at least some of the data, the category history module comprising i) a memory configured to store state data for a plurality of categories indexed by respective category identifiers, wherein the state data stored in the memory for each category identifier corresponds to the entity and previous transactions associated with each category, and the memory; ii) a decay logic stage configured to change the state data stored in the memory based on the time difference between the time of the incoming transaction and the time of the previous transaction, the decay logic stage comprising a) When the memory contains state data for entity and category identifier pairs associated with an incoming transaction, retrieve the state data, apply a decaying function to the state data to generate a decayed version of the state data, and output the decayed version of the state data, where the decaying function depends on the time difference, a decay logic stage, and iii) Using an input tensor generated from data corresponding to the incoming transaction, update the state data output from the decay logic stage to generate updated state data; for the entity and category identifier pair associated with the incoming transaction, store the updated state data in the memory; and output an output tensor comprising the updated state data, an update logic stage configured to perform comprising iv) The machine learning system is configured to map the output tensor from the category history module to a scalar value representing the likelihood that the incoming transaction presents an anomaly within a sequence of actions, v) A machine learning system is provided, wherein the scalar value is used to determine whether to approve or reject the incoming transaction.
[0016] A first aspect of the present invention is also a method of processing data associated with a proposed transaction associated with an entity, the method comprising: i) storing in a memory state data for a plurality of categories indexed by respective category identifiers, wherein the state data stored in the memory for each category identifier corresponds to an entity and previous transactions associated with each category; ii) changing the state data stored in the memory based on a time difference between the time of the incoming transaction and the time of the previous transaction, the step of changing the state data comprising a) When the memory contains state data for entity and category identifier pairs associated with an incoming transaction, retrieving the state data and applying a decay function to the state data to generate a decayed version of the state data, wherein the decay function depends on the time difference; and outputting the decayed version of the state data including the step of iii) Using an input tensor generated from data corresponding to the incoming transaction to update the output state data to generate updated state data, storing the updated state data in the memory for the entity and category identifier pair associated with the incoming transaction, and outputting an output tensor comprising the updated state data iv) Mapping the output tensor to a scalar value representing the likelihood that the incoming transaction presents an anomaly within a sequence of actions v) Using the scalar value to determine whether to approve or reject the incoming transaction The method includes the steps above
[0017] A first aspect of the present invention is a non - transitory computer - readable medium comprising instructions, which, when executed by a processor, cause the processor to perform a method of processing data associated with a proposed transaction associated with an entity, the method including i) Storing in a memory state data for a plurality of categories indexed by respective category identifiers, wherein the state data stored in the memory for each category identifier corresponds to previous transactions associated with the entity and the respective category ii) Modifying the state data stored in the memory based on the time difference between the time of the incoming transaction and the time of the previous transaction, wherein the step of modifying the state data a) When the memory contains state data for entity and category identifier pairs associated with an incoming transaction, retrieving the state data and applying an attenuation function to the state data to generate an attenuated version of the state data, the attenuation function being dependent on a time difference, and outputting the attenuated version of the state data including the step of iii) Using an input tensor generated from data corresponding to the incoming transaction to update the output state data to generate updated state data, storing the updated state data in the memory for the entity and category identifier pair associated with the incoming transaction, and outputting an output tensor comprising the updated state data iv) Mapping the output tensor to a scalar value representing the likelihood that the incoming transaction presents an anomaly within a sequence of actions v) Using the scalar value to determine whether to approve or reject the incoming transaction further comprising a non - transitory computer - readable medium including the above steps
[0018] A first aspect of the present invention is a computer software product comprising instructions which, when executed by a processor, cause the processor to perform a method of processing data associated with a proposed transaction associated with an entity, the method comprising i) Storing in a memory state data for a plurality of categories indexed by respective category identifiers, wherein the state data stored in the memory for each category identifier corresponds to a previous transaction associated with the entity and the respective category ii) Modifying the state data stored in the memory based on a time difference between the time of the incoming transaction and the time of the previous transaction, the step of modifying the state data comprising a) When the memory contains state data for entity and category identifier pairs associated with an incoming transaction, retrieving the state data and applying an attenuation function to the state data to generate an attenuated version of the state data, the attenuation function being dependent on the time difference; and outputting the attenuated version of the state data including the step of iii) using an input tensor generated from data corresponding to the incoming transaction to update the output state data to generate updated state data, storing the updated state data in the memory for the entity and category identifier pair associated with the incoming transaction, and outputting an output tensor comprising the updated state data iv) mapping the output tensor to a scalar value representing the likelihood that the incoming transaction presents an anomaly within a sequence of actions v) using the scalar value to determine whether to approve or reject the incoming transaction further comprising a computer software product including
[0019] Accordingly, it will be appreciated that embodiments of the present invention provide an improved configuration in which state data relating to previous transactions of a matching category is stored in memory and, when there is a transaction in that category for the same entity, i.e., for the entity and category identifier pair associated with the new incoming transaction, the stored state data for that category is retrieved. The categories that can be used will be described in more detail below, but generally, a category can include, for example, a merchant, merchant type, transaction amount, time, or some combination thereof.
[0020] Accordingly, the category history module provides a mechanism for tracking patterns across transactions of a particular category. In general, unlike conventional mechanisms such as time-decay cells that extract state from the last transaction (i.e., the immediately preceding transaction for that entity), the category history module selectively extracts state data corresponding to the same category as the current transaction, which may not be the most recent transaction, and a significant period of time may have elapsed since the last transaction for that category was processed, and any number of intervening transactions corresponding to other categories may have occurred in between. This enables the extraction of signals that provide information about the novelty of the interaction between the entity and the particular category, where such novelty - or the behavioral patterns related to such novelty - may indicate anomalous behavior (e.g., fraud).
[0021] Advantageously, embodiments of the present invention may provide a significant improvement in the speed and / or confidentiality of implementation. Specifically, data scientists no longer need to manually design this type of feature, which means that this type of feature does not need to be disclosed for model management as an input to the classifier, and the implementation time for the model can be reduced.
[0022] Embodiments of the present invention may also advantageously provide an improvement in detection performance. The category history module can learn more subtle interaction logic than can be manually designed by a human in a finite amount of time. Specifically, the category history module can learn that only a particular type of activity should be eligible as trust building for scoring the current type of activity. The applicant has discovered that the use of the present invention typically leads to a significant increase in detection as compared to techniques that may be known per se in the art.
[0023] Embodiments of the present invention can be applied to a wide variety of digital transactions, including but not limited to card payments, so-called "telegraphic" transfers, peer-to-peer payments, Bankers' Automated Clearing System (BACS) payments, and Automated Clearing House (ACH) payments. The output of the machine learning system can be used to prevent a wide variety of fraudulent and criminal behaviors, such as card fraud, application fraud, payment fraud, merchant fraud, gaming fraud, and money laundering.
[0024] As outlined above, the category history module retrieves state data corresponding to the same category as the arriving transaction. Generally, the memory can store state data for every possible category (e.g., for each merchant, or for each merchant type), but in practice, this may not be the case. For example, there may be a limit imposed on the number of categories for which state data can be stored in the memory. This can be due to hardware or software limitations for a given implementation, or it can be a design decision to limit the amount of state data held to some specific amount. Thus, there is a possibility that the arriving transaction may correspond to a category for which there is no existing state data for the corresponding category identifier stored in the memory at that time. In some embodiments, the decay logic stage is further configured to generate new state data and output the new state data when the memory does not contain state data for an entity and category identifier pair associated with the arriving transaction The new state data can, for example, typically comprise a zero tensor. However, it will be appreciated that any other default tensor can be used for the new state data where appropriate.
[0025] Generally, data from arriving transactions can be converted into an input tensor suitable for input to the rest of the machine learning system. The input tensor can be regarded as a feature tensor or vector, i.e., a tensor of values each corresponding to the extent to which the arriving transaction has some features, which will be understood by those skilled in the art. In some embodiments, the machine learning system is configured to apply a neural network layer having a plurality of learned weights respectively to the data corresponding to the arriving transaction to generate an input tensor and to provide the input tensor to an update logic stage of the category history module. In other words, the first neural network can be represented as an embedding layer. This first neural network can comprise a first multi-layer perceptron (MLP) in some embodiments.
[0026] Similarly, the output tensor produced by the update logic stage is mapped to a scalar value as outlined above. In some embodiments, the machine learning system is further provided with a second neural network stage configured to apply a second neural network layer having a plurality of learned weights respectively to the output tensor generated by the update logic stage of the category history module to generate a scalar value representing the likelihood that the arriving transaction presents an anomaly within the sequence of actions. This second neural network can comprise a second multi-layer perceptron (MLP) in some embodiments.
[0027] An MLP is generally a type of feed-forward neural network constructed from at least an input layer, a hidden layer, and an output layer, and those skilled in the art will understand that each layer comprises several nodes. The hidden layer and output layer nodes have an activation function (e.g., a non-linear activation function). Nodes from each layer are coupled, with a certain weight, to nodes in subsequent layers (generally, each node in a layer is coupled to every node in the next layer). These weights can be learned during supervised learning, for example, via backpropagation, etc., during a training phase.
[0028] In some embodiments, the update logic stage comprises a third neural network stage configured to generate updated state data using decayed state data and an input tensor derived from arriving transactions. In other words, the updated state data can also be generated by a neural network layer. This neural network layer can also use weights that are trained. The third neural network stage can comprise a recurrent neural network in some embodiments.
[0029] The amount of time elapsed between consecutive transactions for a given entity (e.g., a cardholder) belonging to a particular category can vary significantly. For example, if the category includes merchant identification information and state data is stored for each merchant (i.e., state data from previous buyer-seller interactions involving that cardholder and merchant), the card may be used several times a week for one merchant, but the same card may be weeks or months apart between transactions for another merchant. Similarly, different entities may have transactions at quite different rates, and the pattern of behavior for any given entity may vary over time.
[0030] It is understood that the amount of time elapsed since the relevant state data retrieved from memory for a particular entity and category combination of arrival transactions is stored can have a significant impact on its relevance when determining the validity of new arrival transactions. The attenuation logic stage modifies the state data to reduce that impact.
[0031] The attenuation function is used to reduce the influence of previous state data depending on how old the previous state data is. In a set of embodiments, the attenuation logic stage is configured to modify the state data stored in memory for an entity associated with an arrival transaction based on the time difference between the time of the arrival transaction and the time of the most recent transaction for that entity. In other words, in such embodiments, the entire state data table for an entity is attenuated each time there is a transaction for that entity. Advantageously, this approach does not require any record-keeping of when each category was last seen for that entity. Rather, a single timestamp is used for each entity, where that timestamp corresponds to the last transaction for that entity (regardless of which category that transaction is associated with). In some such embodiments, the memory is configured to store a timestamp for the state data stored in the memory for each entity. By attenuating the state table for a particular entity, latency can be advantageously reduced (as compared to attenuating the table for all entities) without making it unnecessary to store a timestamp for each entity-category pair (as would be the case if only the state data for a particular entity-category pair were attenuated, as outlined for an alternative set of embodiments below).
[0032] In an alternative set of embodiments, the attenuation logic stage is configured to modify all state data stored in memory for each entity based on the time difference between the time of the arriving transaction and the time of the immediately preceding transaction. In other words, in such embodiments, every time there is a transaction, the entire state data table for all entities is attenuated using the timestamp associated with the last transaction (regardless of who was involved in that transaction), regardless of which entity the transaction was for. However, it will be appreciated that this can impose significant computing requirements in order to reduce an unnecessary increase in latency.
[0033] In a further alternative set of embodiments, the attenuation logic stage is configured to modify only the state data stored in memory for an entity and category identifier pair associated with the arriving transaction, and the attenuation logic stage modifies the state data based on the time difference between the time of the arriving transaction and the time of the previous transaction for the entity and category identifier pair associated with the arriving transaction. Under this configuration, only the state data for a particular entity and category combination is attenuated each time there is a transaction for that same combination. In some such embodiments, the memory is configured to store a timestamp for each state data stored in the memory. The timestamp may form part of the state data or may be stored separately. Advantageously, this approach does not require attenuating the entire table for each transaction, thereby enabling faster processing in the trade-off of increasing memory usage for storing timestamps.
[0034] There are several options for implementing the storage of state data in memory. However, in some embodiments, the memory is configured as a stack, such that the retrieved state data is removed from the stack and the updated state data is stored at the top of the stack. According to such embodiments, the most recent entity-category interaction can be found at the top of the stack, with the oldest interaction at the bottom of the stack.
[0035] In some embodiments, the category history module is configured such that when the memory is full, the state data stored for at least one category identifier is erased. In some such embodiments, the category history module is configured such that when the memory is full, the state data stored for the most recently seen category identifier is erased. Under such a configuration, the oldest data is “dropped” from the memory to make space for new data. Effectively, this provides an additional “forgetting” function in that, assuming an appropriate choice of memory size, data that is too old is removed. Generally, the size of the memory can be selected to ensure that data that it is desirable to retain is not routinely erased from the memory.
[0036] In some embodiments, the state data stored for each entity and category identifier pair comprises a tensor of values. Each value in the tensor may correspond to some feature or combination of features that a transaction can have, and the magnitude of the value provides a measure of how strongly the transaction (or the pattern of previous transactions) exhibits those features.
[0037] In some such embodiments, the decay function applies a respective decay multiplier (e.g., exponential decay) to each of the values. It will be appreciated that different multipliers (e.g., different exponential decays) may be applied to each value. By using different decay multipliers for different values, it becomes possible to set the strength of the memory to be different for different features. It may be desirable for the history of some features to be retained for a longer time period than the history of other features. For example, state data may be related to (i.e., dependent on) the average transaction amount for transactions of entity and category identifier pairs, and it may be useful to retain the influence of this average for several months at a time (e.g., 6 or 12 months). On the other hand, state data that depends on the transaction speed (i.e., the rate of transactions per unit time) may be important over a shorter time period (e.g., over 1 or 2 hours), and thus it may be preferable to limit its influence beyond such a period. In practice, the values stored in the state data may not have a direct (i.e., one-to-one) correspondence with any particular one or more features, but rather, it will be appreciated that each is a function of one or more features. In other words, the numerical values produced (e.g., via a neural network layer) are stored as state data, and each numerical value may have a complex dependence on multiple features. The relationship between those features and the corresponding state data values (i.e., numerical values) associated with them can be learned during training. Thus, there is not necessarily any user-specified decay of the state data related to any particular feature, but rather, it should be understood that the machine learning system can learn that the historical state data derived from some features (such as the examples described previously) - or combinations of such features - should be retained longer than others.
[0038] The decay multipliers can be optionally selected to adjust, as appropriate, the degree of memory retention for each feature. These decay multipliers may, in at least some embodiments, be determined in advance, i.e., may be preset. According to some such embodiments, the machine learning system should receive values derived from some features or combinations of features to different extents by one or more of these decay factors, i.e., may learn that the decay factors affect the end-to-end training process.
[0039] However, in some embodiments, at least some (and potentially all) of the decay multipliers can be learned. Instead of providing a preselection of the decay factors, the machine learning system can learn for itself the best decay factors to use and can learn for itself for which state data those learned decay factors should be used.
[0040] Generally, the decay multipliers can be configured in a tensor (such as a vector) such that time decay is multiplied element by element using the state data to be decayed. If categorical state data is stored in a table and each column can be indexed by a category identifier, the tensor (such as a vector) of time decay can be multiplied element by element using each column of the table. It will of course be understood that the table can be equivalently configured such that the state data is stored in rows indexed by category identifiers instead.
[0041] As outlined above, the categorical history module provides that patterns over transactions for a particular entity and category pair are retained. In some embodiments, the machine learning system further comprises a time decay cell module configured to process at least some of the data corresponding to the arriving transactions, the time decay cell module a second memory configured to store second state data corresponding to the immediately preceding transaction, a second decay logic stage configured to modify second state data stored in a second memory based on a time difference between a time of an arrival transaction and a time of a previous transaction and a time decay cell module. Unlike the category history module, the time decay cell module works on a per transaction basis for a particular entity regardless of the category to which those transactions pertain, as will be appreciated by those skilled in the art.
[0042] According to such an embodiment, the category history module and the time decay cell module may be arranged in parallel with each other and may provide complementary analysis of arrival transactions. The category history module helps to detect anomalies (e.g., fraud) based on the novelty of the interaction between an entity (e.g., a cardholder) and a category (e.g., a merchant, merchant type, time, transaction amount, etc.), while the time decay cell module helps to detect anomalies based on the recent transaction behavior for that entity across all categories.
[0043] In some such embodiments, the time decay cell module a fourth neural network stage configured to determine next state data using the previous state data modified by the second decay logic stage and a second input tensor derived from the arrival transaction, wherein the fourth neural network stage is configured to store the next state data in a second memory is further provided. The fourth neural network stage may comprise a recurrent neural network in some embodiments.
[0044] In embodiments where a category history module and a time decay cell are provided, data corresponding to arrival transactions is supplied to each of both of them. Each of these modules may receive different data, or there may be some (or potentially complete) overlap in the data received by each module. Generally, these two modules come to act in different characteristics, and thus it is expected that the data provided to each module may be unique to each module.
[0045] When a neural network layer (e.g., MLP) with learned weights is used to generate an input tensor from an arrival transaction, the same neural network layer may generate the input tensor such that a first portion of the input tensor is supplied to the category history module and a second portion of the input tensor is supplied to the time decay cell module.
[0046] In some embodiments where a time decay cell module is used, the machine learning system is configured to map output data from both the time decay cell module and the category history module to a scalar value representing the likelihood that the arrival transaction presents an anomaly within a sequence of actions.
[0047] In some embodiments, a memory is configured to store state data for a plurality of entities, and for each entity, a plurality of categories indexed by respective category identifiers are stored, and the state data stored in the memory for each category identifier corresponds to previous transactions associated with each entity and each category.
[0048] As used herein, "category" relates to some aspect of a transaction, i.e., it should be understood that a category can be regarded as a group or label that can be applied to a transaction that meets certain criteria. For example, a category can represent different secondary entities (e.g., merchants) with which an entity (or a "primary" entity, which can be, for example, a cardholder) is involved in a transaction. There are several different categories that can be used. In some embodiments, each of a plurality of categories comprises one or more of a secondary entity, a secondary entity type, a transaction value, a transaction value band, a day of the week, a time, an hour of the day, and a time window. In a particular set of embodiments, a plurality of categories can comprise a composite of one or more of these categories.
[0049] Viewed from a second aspect, an embodiment of the present invention is a method of processing data associated with a proposed transaction, comprising: receiving an arrival event from a client transaction processing system, the arrival event being associated with a request for an approval determination for a proposed transaction; analyzing the arrival event to extract data for the proposed transaction, the analyzing step including determining a time difference between the proposed transaction and a previous transaction; applying a machine learning system of an embodiment of the first aspect of the present invention to output a scalar value representing the likelihood that the proposed transaction presents an anomaly within a sequence of actions, the applying step including accessing a memory to retrieve at least state data for entities and category identifiers associated with the transaction; A step of determining a binary output based on a scalar value output by a machine learning system, wherein the binary output indicates whether a proposed transaction is approved or rejected. A step of returning the binary output to a client transaction processing system. Provided is a method including the above.
[0050] A second aspect of the present invention is a non-transitory computer-readable medium storing instructions, which, when executed by a processor, cause the processor to perform a method of processing data associated with a proposed transaction. The method includes: Receiving an arrival event from a client transaction processing system, wherein the arrival event is associated with a request for an approval determination for a proposed transaction. Analyzing the arrival event to extract data for the proposed transaction, including determining a time difference between the proposed transaction and a previous transaction. Applying a machine learning system according to an embodiment of the first aspect of the present invention to output a scalar value representing the likelihood that the proposed transaction presents an anomaly within a sequence of actions. The applying step includes accessing a memory to retrieve at least state data for entities and category identifiers associated with the transaction. Determining a binary output based on the scalar value output by the machine learning system, wherein the binary output indicates whether the proposed transaction is approved or rejected. Returning the binary output to the client transaction processing system. The non-transitory computer-readable medium includes the above steps.
[0051] A second aspect of the present invention is a computer software product comprising instructions which, when executed by a processor, cause the processor to perform a method of processing data associated with a proposed transaction, the method comprising receiving an arrival event from a client transaction processing system, the arrival event being associated with a request for an approval determination for a proposed transaction; analyzing the arrival event to extract data for the proposed transaction, the step including determining a time difference between the proposed transaction and a previous transaction; applying a machine learning system of an embodiment of the first aspect of the present invention to output a scalar value representing the likelihood that the proposed transaction presents an anomaly within a sequence of actions, the applying step including accessing a memory to retrieve at least state data for entities and category identifiers associated with the transaction; determining a binary output based on the scalar value output by the machine learning system, the binary output indicating whether the proposed transaction is approved or rejected; returning the binary output to the client transaction processing system and a computer software product.
[0052] The category history module is novel and inventive in itself, and thus, viewed from a third aspect, an embodiment of the present invention is a category history module for use in a machine learning system for processing data corresponding to arrival transactions associated with entities, the category history module being configured to process at least some of the data, the category history module i) A memory configured to store state data for a plurality of categories indexed by respective category identifiers, wherein for each category identifier, the state data stored in the memory corresponds to an entity and a previous transaction associated with each category, the memory and, ii) A decay logic stage configured to modify the state data stored in the memory based on a time difference between the time of an arriving transaction and the time of the previous transaction, wherein the decay logic stage, a) When the memory contains state data for an entity and category identifier pair associated with an arriving transaction, retrieves the state data, applies a decay function to the state data to generate a decayed version of the state data, and is configured to output the decayed version of the state data, the decay function being dependent on the time difference, the decay logic stage, iii) Using an input tensor generated from data corresponding to an arriving transaction, updating the state data output from the decay logic stage to generate updated state data; storing the updated state data in the memory for an entity and category identifier pair associated with the arriving transaction; and outputting an output tensor comprising the updated state data, an update logic stage configured to perform Providing a category history module comprising.
[0053] A third aspect of the present invention is a method of processing data associated with a proposed transaction associated with an entity, comprising: i) Storing in a memory state data for a plurality of categories indexed by respective category identifiers, wherein for each category identifier, the state data stored in the memory corresponds to an entity and a previous transaction associated with each category, the step and, ii) modifying the state data stored in the memory based on a time difference between the time of the arrival transaction and the time of the previous transaction, wherein the step of modifying the state data a) when the memory contains state data for an entity and category identifier pair associated with the arrival transaction, retrieving the state data, applying an attenuation function to the state data to generate an attenuated version of the state data, wherein the attenuation function depends on the time difference, and outputting the attenuated version of the state data; and including; iii) using an input tensor generated from data corresponding to the arrival transaction to update the output state data to generate updated state data, storing the updated state data in the memory for an entity and category identifier pair associated with the arrival transaction, and outputting an output tensor comprising the updated state data. The method further extends to.
[0054] A third aspect of the present invention is a non - transitory computer - readable medium comprising instructions which, when executed by a processor, cause the processor to perform a method of processing data associated with a proposed transaction associated with an entity, the method comprising: i) storing in the memory state data for a plurality of categories indexed by respective category identifiers, wherein the state data stored in the memory for each category identifier corresponds to the previous transaction associated with the entity and the respective category; and ii) modifying the state data stored in the memory based on a time difference between the time of the arrival transaction and the time of the previous transaction, wherein the step of modifying the state data a) When the memory contains state data for entity and category identifier pairs associated with an arriving transaction, retrieving the state data and applying a decay function to the state data to generate a decayed version of the state data, wherein the decay function depends on the time difference; and outputting the decayed version of the state data including the step of iii) Using an input tensor generated from data corresponding to the arriving transaction to update the output state data to generate updated state data, storing the updated state data in the memory for the entity and category identifier pair associated with the arriving transaction, and outputting an output tensor comprising the updated state data further comprising a non - transitory computer - readable medium including the step of
[0055] A third aspect of the present invention is a computer software product comprising instructions which, when executed by a processor, cause the processor to perform a method of processing data associated with a proposed transaction associated with an entity, the method comprising i) Storing in a memory state data for a plurality of categories indexed by respective category identifiers, wherein the state data stored in the memory for each category identifier corresponds to a previous transaction associated with the entity and the respective category ii) Modifying the state data stored in the memory based on a time difference between the time of the arriving transaction and the time of the previous transaction, wherein the step of modifying the state data comprises a) When the memory contains state data for entity and category identifier pairs associated with an arriving transaction, retrieving the state data and applying a decay function to the state data to generate a decayed version of the state data, wherein the decay function depends on the time difference; and outputting the decayed version of the state data including steps of iii) using an input tensor generated from data corresponding to an arrival transaction to update the output state data to generate updated state data, storing the updated state data in memory for entity and category identifier pairs associated with the arrival transaction, and outputting an output tensor comprising the updated state data which further extends to a computer software product including
[0056] It will be understood that the optional features described above for embodiments of the first aspect of the invention are equally applicable to the second and third aspects of the invention where technically appropriate.
[0057] Where technically appropriate, embodiments of the invention may be combined. Embodiments are described herein as comprising a number of features / elements. The present disclosure also extends to separate embodiments consisting of or consisting essentially of said features / elements.
[0058] Technical references such as patents and applications are incorporated herein by reference.
[0059] Any embodiment specifically and explicitly described herein may form the basis of a disclaimer, alone or in combination with one or more further embodiments.
[0060] In the context of this specification, "comprising" should be construed as "including". Aspects of the invention comprising a number of elements are also to be taken to extend to alternative embodiments "consisting" or "consisting essentially" of the relevant elements.
[0061] The term "memory" should be understood to mean any means suitable for storing data, and includes both volatile and non-volatile memory as appropriate for the intended application. This includes, but is not limited to, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), magnetic storage, solid state storage, and flash memory, as will be understood by those skilled in the art. When technically appropriate, one or more combinations of these may also be used for storing data (for example, using a volatile memory with faster access for frequently accessed data).
[0062] The term "data" is used in different contexts herein to refer to digital information such as that represented by known bit structures within one or more programming languages. In use, data may refer to digital information stored as a sequence of bits in a computer memory. Some machine learning models may operate on a structured array of data in a predefined bit format. It will be readily understood by those skilled in the art that these may be referred to as arrays, multi-dimensional arrays, matrices, vectors, tensors, or other such similar terms. In the case of machine learning methods, for example, a multi-dimensional array or tensor having a defined range in multiple dimensions may be "flattened" so as to be represented as a sequence or vector of values stored according to a predefined format (for example, n-bit integers or floating point numbers, signed or unsigned) (for example, in memory). Accordingly, the term "tensor" as used herein covers multi-dimensional arrays having one or more dimensions (for example, vectors, matrices, volumetric arrays, etc.).
[0063] The principles of the present invention apply regardless of the particular data format selected. The data may be represented as an array, vector, tensor, or any other suitable format. For ease of reference, these terms are used interchangeably herein, and it should be understood that a reference to a "vector" or "vectors" of values extends, as appropriate, to any one or more n-dimensional tensors of values. Similarly, a reference to a "tensor" or "tensors" of values should be understood to extend to vectors, which are understood by those skilled in the art to be simply one-dimensional tensors. The principles of the present invention may apply regardless of the formatting of the data structure used for these arrays of values. For example, state data may be stored in memory as a one-dimensional tensor (i.e., a vector) or as a tensor having two or more dimensions (i.e., multiple tensors), and those skilled in the art will readily understand that suitable modifications may be made to the data processing elements to handle the selected data format. The relative positions between various state values, e.g., how they are ordered within a vector or tensor, are generally not a problem, and the scope of the present invention is not limited to any particular data format or structure.
[0064] The term "structured numerical representation" is used to refer to numerical data in a structured form, such as an array of one or more dimensions that stores numerical values having a general data type, such as integer or floating-point values. A structured numerical representation may comprise a vector or tensor (as used within the terms of machine learning). A structured numerical representation is typically stored as a set of indexed and / or contiguous memory locations. For example, a one-dimensional array of 64-bit floating-point numbers may be represented in computer memory as a contiguous sequence of 64-bit memory locations in a 64-bit computing system.
[0065] The term "transaction data" is used herein to refer to electronic data associated with a transaction. A transaction includes a series of communications between different electronic systems for effecting a payment or a currency exchange. Generally, transaction data may comprise data indicating events (e.g., actions taken in a timely manner) that are related to and may be beneficial for transaction processing. Transaction data may include structured data, unstructured data, and semi-structured data, or any combination thereof. Transaction data may also include data associated with a transaction, such as data used to process the transaction. In some cases, transaction data may be widely used to refer to actions taken with respect to one or more electronic devices. Transaction data may take various forms depending on the exact implementation. However, different data types and formats may be appropriately converted by pre-processing or post-processing.
[0066] The term "interface" is used herein to refer to any physical and / or logical interface that enables one or more of data input and data output. The interface may be implemented by a network interface adapted to send and / or receive data, or by retrieving data from one or more memory locations such that it is implemented by a processor that executes a set of instructions. The interface may also comprise a physical (network) connection through which data is received, such as hardware to enable wired or wireless communication over a particular medium. The interface may comprise an application programming interface, and / or method calls or method returns. For example, in a software implementation, the interface may comprise data passing and / or memory references to a function initiated via a method call, where the function comprises computer program code executed by one or more processors, and in a hardware implementation, the interface may comprise a wired interconnect between different chips, chip sets, or portions of a chip. In the drawings, an interface may be indicated by the boundary of a processing block having inward and / or outward arrows representing data transfer.
[0067] The terms "component" and "module" are used interchangeably to refer to either a hardware structure having a particular function (e.g., in the form of mapping input data to output data), or a combination of general hardware and particular software (e.g., particular computer program code executed on one or more general purpose processors). A component or module may be implemented as a particular packaged chip set, such as an application specific integrated circuit (ASIC) or a programmed field programmable gate array (FPGA), and / or as software objects, classes, class instances, scripts, code portions, etc. that are executed in use by a processor.
[0068] The term "machine learning model" is used herein to refer to an implementation form that is executed at least by hardware of a machine learning model or a machine learning function. Known models within the field of machine learning include logistic regression models, naive Bayes models, random forests, support vector machines, and artificial neural networks. Implementations of classifiers can be provided within one or more machine learning programming libraries including, but not limited to, scikit-learn, TensorFlow, and PyTorch.
[0069] The term "mapping" is used herein to refer to a transformation or conversion from a first set of data values to a second set of data values. These two sets of data values can be arrays of different sizes where the output array has a smaller number of dimensions than the input array. The input array and the output array can have the same or different data types. In some examples, the mapping is a one-way mapping to scalar values.
[0070] The term "neural network architecture" refers to a set of one or more artificial neural networks configured to perform a specific data processing task. For example, a "neural network architecture" can comprise a specific configuration of one or more neural network layers of one or more neural network types. Neural network types include convolutional neural networks, recurrent neural networks, and feedforward neural networks. Convolutional neural networks involve the application of one or more convolution operations. Recurrent neural networks involve an internal state that is updated during a sequence of inputs. Thus, recurrent neural networks are considered to include forms of recurrent or feedback connections, whereby the state of a recurrent neural network at a given time or iteration (e.g., t) is updated using the state of the recurrent neural network at a previous time or iteration (e.g., t-1). Feedforward neural networks involve a transformation operation without feedback, e.g., the operation is applied in a one-way sequence from input to output. Feedforward neural networks include simple "neural networks" and "fully-connected" neural networks. The term "multi-layer perceptron" is used to denote a fully-connected layer and is understood by those skilled in the art to be a special case of a feedforward neural network.
[0071] The term "deep" neural network is used to indicate that a neural network comprises multiple neural network layers in series (this term "deep" is used with both feedforward neural networks and recurrent neural networks). Some of the examples described herein utilize recurrent neural networks and fully-connected neural networks.
[0072] A "neural network layer", typically as defined within machine learning programming tools and libraries, can be regarded as an operation that maps input data to output data. A "neural network layer" can apply one or more parameters, such as weights, to map input data to output data. One or more bias terms can also be applied. The weights and biases of a neural network layer can be applied using one or more multi-dimensional arrays or matrices. Generally, a neural network layer has multiple parameters whose values affect how the input data is mapped to output data by that layer. These parameters can be trained in a supervised manner by optimizing an objective function. This typically involves minimizing a loss function. Some parameters can also be pre-trained or fixed in another manner. Fixed parameters can be regarded as configuration data that controls the operations of a neural network layer. A neural network layer or neural network architecture can comprise a mixture of fixed parameters and learnable parameters. A recurrent neural network layer can apply a series of operations to update a recurrent state and transform input data. The update of the recurrent state and the transformation of input data can involve one or more transformations of the previous recurrent state and input data. A recurrent neural network layer can be trained by unfolding modeled recurrent units so as to be applicable within machine learning programming tools and libraries. A recurrent neural network may appear to comprise several (sub)layers for applying different gating operations, but most machine learning programming tools and libraries refer to the application of a recurrent neural network as a "neural network layer" as a whole, and here, this convention is followed. Finally, a feedforward neural network layer can apply one or more of a set of weights and biases to input data to generate output data. This operation can be represented as a matrix operation (e.g., when a bias term can be included by adding a value such as 1 onto the input data).Alternatively, the bias can be applied through a separate addition operation. As described above, the term "tensor" is used to refer to an array that can have multiple dimensions, such as in a machine learning library. For example, a tensor can comprise a vector, a matrix, or a data structure with a greater number of dimensions. In a preferred example, the tensor being described can comprise a vector having a predefined number of elements.
[0073] To model complex non-linear functions, a neural network layer as described above may be followed by a non-linear activation function. Common activation functions include the sigmoid function, the hyperbolic tangent function, and the rectified linear unit (RELU). Many other activation functions exist and can be applied. The activation function can be selected based on testing and preference. The activation function may be omitted in some situations and / or can form part of the internal structure of the neural network layer.
[0074] The exemplary neural network architectures described herein can be configured through training. In some cases, "learnable" or "trainable" parameters can be trained using a technique called backpropagation. During backpropagation, the neural network layers that make up each neural network architecture are initialized (e.g., using randomized weights) and then used to make predictions using a set of input data from a training set (e.g., a so-called "forward" pass). The predictions are used to evaluate a loss function. For example, a "ground truth" output may be compared to the predicted output, and the difference can form part of the loss function. In some examples, the loss function can be based on the absolute difference between a predicted scalar value and a binary ground truth label. The training set can comprise a set of transactions. When gradient descent is used, the loss function is used to determine the gradient of the loss function with respect to the parameters of the neural network architecture, where the gradient is then used to backpropagate updates to the parameter values of the neural network architecture. Typically, the updates are propagated according to the derivative of the weights of the neural network layer. For example, the gradient of the loss function with respect to the weights of a neural network layer can be determined and used to determine updates to the weights that minimize the loss function. In this case, optimization techniques such as gradient descent, stochastic gradient descent, Adam, etc. can be used to adjust the weights. The chain rule and an auto-differentiation function can be applied to efficiently compute the gradient of the loss function by working backward sequentially through the neural network layers.
[0075] Next, some embodiments of the present invention will be described with reference to the accompanying drawings.
Brief Description of the Drawings
[0076]
Figure 1A
Figure 1B
Figure 1C
Figure 2A
Figure 2B
Figure 3A
Figure 3B
Figure 4
Figure 5A
Figure 5B
Figure 6
Figure 7
Mode for Carrying Out the Invention
[0077] Some exemplary embodiments relating to a machine learning system for use in transaction processing are described herein. In some embodiments, a machine learning system is applied in a real-time high-volume transaction processing pipeline to provide an indication as to whether a transaction or entity matches a previously observed and / or predicted pattern of activities or actions, e.g., an indication as to whether the transaction or entity is "normal" or "abnormal". The term "behavioral" is used herein to refer to this pattern of activities or actions. The indication may comprise a scalar value that is normalized within a predefined range (e.g., 0 to 1) and can then be used to prevent fraud and other misuse in a payment system. The machine learning system may apply a machine learning model that is updated as more transaction data is obtained, e.g., continuously trained based on new data, to reduce false positives and maintain the accuracy of the output metric. This example may be particularly useful for preventing fraud in cases where the physical presence of a payment card cannot be verified (e.g., an online transaction referred to as "card-not-present") or for commercial transaction processing where expensive transactions may be routine and it may be difficult to classify patterns of behavior as "unexpected". Thus, this example facilitates the processing of transactions when these transactions are primarily "online", i.e., digitally performed via one or more public communication networks.
[0078] Some of the embodiments described herein enable a machine learning model to be adjusted to be specific to patterns of behavior between entities (such as account holders) and categories (such as merchants, transaction amounts, times, etc.). For example, a machine learning model can model entity-category pair-specific patterns of behavior. The machine learning systems described herein can provide a machine learning model that dynamically updates despite large-scale transaction flows and / or despite the need to isolate different data sources.
[0079] As outlined previously, embodiments of the present invention can be applied to a wide variety of digital transactions including, but not limited to, card payments, so-called "wire" transfers, peer-to-peer payments, bank automated clearing system (BACS) payments, and automated clearing house (ACH) payments. The output of the machine learning system can be used to prevent a wide variety of fraudulent and criminal behaviors such as card fraud, application fraud, payment fraud, merchant fraud, gaming fraud, and money laundering.
[0080] For example, this exemplary machine learning system, configured and / or trained as outlined below, is configured to detect novelty in entity-category interactions.
[0081] Non-limiting exemplary embodiments of a machine learning system according to an embodiment of the present invention are described below. FIGS. 1A-5B provide context for the machine learning system. A machine learning system in the form of a particular neural network architecture is shown as in FIGS. 6 and 7.
[0082] Exemplary Transaction Processing System Figures 1A - 1C show a set of exemplary transaction processing systems 100, 102, 104. These exemplary transaction processing systems are described to provide a context for the invention described herein, but should not be considered limiting, as the configuration of any one implementation may vary based on the specific requirements of that implementation. However, the exemplary transaction processing systems described enable one of ordinary skill in the art to identify some high - level technical features that are appropriate for the following description. These three exemplary transaction processing systems 100, 102, 104 represent different areas where variations can occur.
[0083] Figures 1A - 1C show a set of client devices 110 configured to initiate a transaction. In this example, the set of client devices 110 includes a smartphone 110 - A, a computer 110 - B, a point - of - sale (POS) system 110 - C, and a portable merchant device 110 - D. These client devices 110 provide a non - exhaustive set of examples. Generally, any electronic device or set of devices can be used to initiate a transaction. In some cases, a transaction includes a purchase or payment. For example, a purchase or payment can be an online or mobile purchase or payment made through the smartphone 110 - A or the computer 110 - B, or a purchase or payment made at a merchant store via the POS system 110 - C or the portable merchant device 110 - D, etc. A purchase or payment can be for goods and / or services.
[0084] In FIGS. 1A - 1C, client device 110 is communicatively coupled to one or more computer networks 120. Client device 110 can be communicatively coupled in various ways, including by one or more wired and / or wireless networks, including a telecommunications network. In a preferred example, all communications across one or more computer networks are secured, for example, using the Transport Layer Security (TLS) protocol. In FIG. 1A, two computer networks are shown as 120 - A and 120 - B. These can be separate networks or different portions of a common network. The first computer network 120 - A communicatively couples client device 110 to merchant server 130. Merchant server 130 can execute a computer process that implements a process flow for a transaction. For example, merchant server 130 can be a backend server that handles transaction requests received from POS system 110 - C or portable merchant device 110 - D, or can be used by an online merchant to implement a website where purchases can be made. It will be understood that the examples of FIGS. 1A - 1C are a necessary simplification of an actual architecture, and that there can be several interacting server devices implementing an online merchant, including separate server devices for providing hypertext markup language (HTML) pages that detail products and / or services and for handling the payment process.
[0085] In FIG. 1A, the merchant server 130 is communicatively coupled to a further set of backend server devices to process transactions. In FIG. 1A, the merchant server 130 is communicatively coupled to the payment processor server 140 via the second network 120-B. The payment processor server 140 is communicatively coupled to a first data storage device 142 that stores transaction data 146 and a second data storage device 144 that stores auxiliary data 148. The transaction data 146 may comprise a batch of transaction data related to different transactions initiated over a period of time. The auxiliary data 148 may comprise data associated with the transaction, such as records storing merchant data and / or end-user data. In FIG. 1A, the payment processor server 140 is communicatively coupled to the machine learning server 150 via the second network 120-B. The machine learning server 150 implements a machine learning system 160 for processing transaction data. The machine learning system 160 is configured to receive input data 162 and map it to output data 164 used by the payment processor server 140 to process specific transactions, such as transactions originating from the client device 110. In some cases, the machine learning system 160 receives at least transaction data associated with a specific transaction and provides an alert or numerical output used by the payment processor server 140 to determine whether the transaction should be permitted (i.e., approved) or rejected. Thus, the output of the machine learning system 160 may comprise labels, alerts, or other indications of fraud, or general malicious or abnormal activity. The output may comprise probabilistic indications such as scores or probabilities. In some cases, the output data 164 may comprise scalar numerical values. The input data 162 may further comprise data derived from one or more of the transaction data 146 and the auxiliary data 148.In some cases, the output data 164 indicates a level of deviation from a particular expected pattern of behavior based on past observations or measurements. For example, when this is significantly and particularly large-scale different from the observed pattern of behavior, this may indicate fraudulent or criminal behavior. The output data 164 may form behavioral measurements. The expected pattern of behavior may be defined either explicitly or implicitly based on observed interactions between different entities within the transaction process flow, such as end users or customers, merchants (including points of sale and back-end locations or entities if these can be different), and banks.
[0086] The machine learning system 160 can be implemented as part of a transaction processing pipeline. An exemplary transaction processing pipeline will be described later with respect to FIGS. 5A and 5B. The transaction processing pipeline can include electronic communication between the client device 110, the merchant server 130, the payment processor server 140, and the machine learning server 150. Other server devices, such as a banking server that provides authorization from an issuing bank, may also be involved. In some cases, the client device 110 can communicate directly with the payment processor server 140. In use, the transaction processing pipeline typically needs to be completed within 100 or 200 milliseconds. Generally, a processing time of less than 1 second can be considered real-time (for example, a human typically perceives an event in a time span of 400 ms). Further, 100 - 200 ms can be the desired maximum latency for the total round-trip time for transaction processing, and within this time span, the time allocated for the machine learning system 160 can be a very small portion of this total amount, such as 10 ms (i.e., less than 5 - 10% of the target processing time), when most of that time can be reserved for other operations in the transaction processing flow. This presents a technical constraint for the implementation of the machine learning system 160. Further, in a real-world implementation, the average throughput can be on the order of 1000 - 2000 per second. This means that most "off-the-shelf" machine learning systems are not suitable for implementing the machine learning system 160. This further means that most of the machine learning techniques described in academic papers cannot be implemented within the above-described transaction processing pipeline without non-trivial adaptation.
[0087] Moreover, existing machine learning systems do not extract the novelty of entity-category interactions. As will be described later with reference to FIGS. 6 and 7, the machine learning system according to an embodiment of the present invention can be used to provide several technical benefits.
[0088] Figure 1B shows a variant form 102 of the exemplary transaction processing system 100 of Figure 1A. In this variant form 102, the machine learning system 160 is implemented within the payment processor computer infrastructure, for example, executed by the payment processor server 140 and / or executed on a server locally coupled within the same local network as the payment processor server 140. The variant form 102 of Figure 1B can be preferable for larger payment processors because it allows for faster response times, greater control, and improved security. However, functionally, the transaction processing pipeline can be similar to the transaction processing pipeline of Figure 1A. For example, in the example of Figure 1A, the machine learning system 160 can be initiated by a secure external application programming interface (API) call, such as a Representational State Transfer (REST) API call using the Hypertext Transfer Protocol Secure (HTTPS), while in Figure 1B, the machine learning system 160 can be initiated by an internal API call, provided that a general end API can handle both requests (for example, a REST HTTPS API can provide an external wrapper for the internal API).
[0089] FIG. 1C shows another variant 104 of the exemplary transaction processing system 100 of FIG. 1A. In this variant 104, the machine learning system 160 is communicatively coupled to the local data storage device 170. For example, the data storage device 170 may be on the same local network as the machine learning server 150 or may comprise a local storage network accessible to the machine learning server 150. In this case, there are a plurality of local data storage devices 170-A to 170-N, where each data storage device stores the segmented auxiliary data 172. The segmented auxiliary data 172 may comprise parameters for one or more machine learning models. In some cases, the auxiliary data 172 may comprise a state for the machine learning model, where the state may be related to a particular entity such as a user or merchant. The segmentation of the auxiliary data 172 may need to be applied to meet security requirements set by third parties such as a payment processor, one or more banks, and / or one or more merchants. In use, the machine learning system 160 accesses the auxiliary data 172-A to 172-N via the plurality of local data storage devices 170-A to 170-N based on the input data 162. For example, the input data 162 may be received through an API request from a particular source and / or may comprise data identifying that a particular segmentation should be used to handle the API request. More details of different storage systems that may be applied to meet security requirements are described in FIGS. 2A and 2B.
[0090] Exemplary Data Storage Configuration Figures 2A and 2B show two exemplary data storage configurations 200 and 202 that may be used by an exemplary machine learning system 210 for processing transaction data. The examples of FIGS. 2A and 2B are two non-limiting examples showing different options available for implementation, and a particular configuration may be selected according to individual circumstances. The machine learning system 210 may comprise an implementation of the machine learning system 160 described in the previous example of FIGS. 1A-1C. The examples of FIGS. 2A and 2B enable the processing of transaction data that is secured using different cryptographic parameters, for example, for the machine learning system 210 to securely process transaction data for heterogeneous entities. It will be appreciated that the configurations of FIGS. 2A and 2B may not be used if the machine learning system 160 is implemented, for example, within an internal transaction processing system, for a single set of secure transaction data and auxiliary data, or as a hosted system for use by a single payment processor.
[0091] FIG. 2A shows a machine learning system 210 communicatively coupled to a data bus 220. The data bus 220 may comprise an internal data bus of the machine learning server 150 or may form part of a storage area network. The data bus 220 communicatively couples the machine learning system 210 to a plurality of data storage devices 230, 231, 232. The data storage devices 230, 231, 232 may comprise any known data storage device such as a magnetic hard disk and a solid state device. The data storage devices 230, 231, 232 are shown as different devices in FIG. 2A, but alternatively they may form different physical regions or portions of storage within a common data storage device. In FIG. 2A, the plurality of data storage devices 230, 231, 232 store entity transaction data 240, category history transaction data 241 (for example, a table of stored state data for entity-category pairs, as will be outlined in more detail later), and auxiliary data 242.
[0092] In FIG. 2A, a first set of data storage devices 230 stores entity transaction data 240, a second set of data storage devices 231 stores category history transaction data 241, and a third set of data storage devices 232 stores auxiliary data 242.
[0093] The auxiliary data 242 may comprise one or more of model parameters for a set of machine learning models (such as trained parameters for a neural network architecture and / or configuration parameters for a random forest model), and state data for those models. In some cases, different sets of entity transaction data 240-A to N, category history transaction data 241-A to N, and auxiliary data 242-A to N are associated with different entities that securely and collectively use the services provided by the machine learning system 210. For example, these may represent data for different banks that need to be kept separately as part of the conditions for providing machine learning services to those entities.
[0094] FIG. 2B shows another way in which different sets of entity transaction data 240-A to N, category history transaction data 241-A to N, and auxiliary data 242-A to N can be stored. In FIG. 2B, the machine learning system 210 is communicatively coupled to at least one data storage device 260 via a data transfer channel 250. The data transfer channel 250 can comprise a local storage bus, a local storage area network, and / or a remote secure storage connection (such as overlaid on an unsecure network such as the Internet). In FIG. 2B, a secure logical storage layer 270 is provided using the physical data storage device 260. The secure logical storage layer 270 can be a virtualized system that is actually implemented independently on top of at least one data storage device 260 but appears to the machine learning system 210 as a separate physical storage device. The logical storage layer 270 may provide separate encrypted partitions 280 for data related to a group of entities (such as related to different issuing banks), and different entity transaction data 240-A to N, category history transaction data 241-A to N, and auxiliary data 242-A to N can be stored in the corresponding partitions 280-A to N. In some cases, an entity can be dynamically created when a transaction is received for processing based on data stored by one or more of the server systems shown in FIGS. 1A-1C.
[0095] Exemplary Transaction Data Figures 3A and 3B show examples of transaction data that can be processed by a machine learning system such as 160 or 210. Figure 3A shows how transaction data can comprise a set of chronological records 300, where each record has a timestamp and comprises a plurality of transaction fields. Optionally, the transaction data can be grouped and / or filtered based on the timestamp. For example, Figure 3A shows a breakdown of the transaction data into current transaction data 310 associated with the current transaction and "older" or historical transaction data 320 that is within a predefined time range of the current transaction. The time range can be set as a hyperparameter of any machine learning system. Alternatively, the "older" or historical transaction data 320 can be set as some number of transactions. A mixture of the two approaches is also possible.
[0096] FIG. 3B shows how transaction data 330 for a particular transaction can be stored in numerical form for processing by one or more machine learning models. For example, in FIG. 3B, the transaction data has at least fields, namely, transaction amount, a timestamp (e.g., as Unix epoch), transaction type (e.g., card payment or account debit), product description or product identifier (i.e., related to the item being purchased), merchant identifier, issuing bank identifier, a set of characters (e.g., Unicode characters within a field of a predefined character length), country identifier, and the like. Note that a wide variety of data types and formats can be received and preprocessed into an appropriate numerical representation. In some cases, the transaction data that occurs, such as that generated by a client device and sent to the merchant server 130, is preprocessed to convert alphanumeric data types to numerical data types for application of one or more machine learning models. Other fields that may be present in the transaction data include, but are not limited to, account number (e.g., credit card number), the location where the transaction is taking place, and the manner in which the transaction is executed (e.g., in person, via phone, on a website).
[0097] As will be outlined in more detail below, the machine learning system 600 of FIGS. 6 and 7 seeks to identify novelty in the interaction between an entity (e.g., a cardholder) and one or more categories (e.g., merchant, merchant type, transaction amount, time, day of week, etc.). Fields in the transaction data can be used directly to identify one or more relevant categories. Additionally or alternatively, one or more categories associated with the transaction can be inferred or derived from the transaction data. It will be appreciated that such derived categories can be obtained in a number of ways. By way of non-limiting example, this can include simple preprocessing such as binning of numerical features, preprocessing that combines multiple categories, or a machine learning model that learns to construct useful categories and is trained with the main machine learning system 600.
[0098] Exemplary Machine Learning System Figure 4 shows an example 400 of a machine learning system 402 that can be used to process transaction data. The machine learning system 402 can implement one or more of the machine learning systems 160 and 210. The machine learning system 402 receives input data 410. The format of the input data 410 can depend on which machine learning model is being applied by the machine learning system 402. If the machine learning system 402 is configured to perform fraud or anomaly detection on a transaction, such as a transaction in progress as described above, the input data 410 can include transaction data such as 330 (i.e., data that forms part of a data package for a transaction), and data derived from entity transaction data (such as 300 in FIG. 3A), and / or data derived from category history data (as outlined in more detail below), and / or data derived from auxiliary data (such as 148 in FIGS. 1A and 1B, or 242 in FIGS. 2A and 2B). The auxiliary data can include secondary data linked to one or more entities identified in the primary data associated with the transaction. For example, if the transaction data for a transaction in progress identifies a user, merchant, and one or more banks (such as the issuing bank for the user and the merchant bank) associated with the transaction via a unique identifier present in the transaction data, the auxiliary data can include data related to these transaction entities. The auxiliary data can be used to enrich the information available to the neural network layer. The auxiliary data can also include data derived from records of activities, such as conversation logs and / or authentication records. In some cases, the auxiliary data is stored in one or more static data records and retrieved from these records based on the received transaction data. Additionally or alternatively, the auxiliary data can include machine learning model parameters retrieved based on the content of the transaction data.For example, a machine learning model may have parameters that are specific to one or more of a user, a merchant, and an issuing bank, and these parameters can be retrieved based on which of these are identified in the transaction data. For example, one or more of the user, the merchant, and the issuing bank may have corresponding embeddings that may include extractable or mappable tensor representations for the entity. For example, each user or merchant may have a tensor representation (e.g., a floating point vector of size 128 to 1024) that can either be retrieved from a database or other data storage based on, for example, a user index or merchant index, or generated by an embedding layer.
[0099] Input data 410 is received at an input data interface 412. The input data interface 412 may comprise an API interface, such as an internal or external API interface as described above. In some cases, a payment processor server 140 as shown in FIGS. 1A - 1C makes a request to this interface, where the request payload contains transaction data. The API interface may be defined to be agnostic with respect to the format or source of the transaction data. The input data interface 412 is communicatively coupled to a machine learning model platform 414. In some cases, a request made to the input data interface 412 triggers the execution of the machine learning model platform 414 that uses the transaction data supplied to the interface. The machine learning model platform 414 is configured as an execution environment for the application of one or more machine learning models to the input data 410. In some cases, the machine learning model platform 414 is configured as an execution wrapper for a plurality of different selectable machine learning models. For example, the machine learning model may be defined using a model definition language (similar to or using markup languages such as, for example, Extensible Markup Language - XML). The model definition language may include SQL, TensorFlow, Caffe, Thinc, and PyTorch (in particular, independently or in combination). In some cases, the model definition language comprises computer program code executable to perform one or more of training and inference of the defined machine learning model. The machine learning model may comprise, for example, in particular, an artificial neural network architecture, an ensemble model, a regression model, a decision tree such as a random forest, a graph model, and a Bayesian network. One exemplary machine learning model based on an artificial neural network will be described later with reference to FIG. 6.The machine learning model platform 414 can define a common (i.e., shared) input definition and output definition such that different machine learning models are applied in a common (i.e., shared) manner.
[0100] In this example, the machine learning model platform 414 is configured to provide at least a single scalar output 416. This can be normalized within a predefined range such as 0 to 1. When normalized, the scalar output 416 can be regarded as the probability that a transaction associated with the input data 410 is fraudulent or abnormal. In this case, a value of "0" can represent a transaction that matches the normal pattern of activity for one or more of the user, merchant, and issuing bank, whereas a value of "1" can indicate that the transaction is fraudulent or abnormal, i.e., does not match the expected pattern of activity (it should be noted by those skilled in the art that the normalized range can be inverted or within different boundaries, etc., and can have the same functional effect). Although the value range can be defined as 0 to 1, it should be noted that the output values may not be uniformly distributed within this range. For example, a value of "0.2" may be a common output for a "normal" event, and a value of "0.8" may be considered to exceed the threshold for a typical "abnormal" or fraudulent event. Thus, the machine learning model implemented by the machine learning model platform 414 can implement a certain form of mapping between high-dimensional input data (e.g., transaction data and any withdrawal assistance data) and a single-value output. In some cases, for example, the machine learning model platform 414 may be configured to receive input data for the machine learning model in a numerical format, and each defined machine learning model is configured to map input data defined in a similar manner. The exact machine learning model applied by the machine learning model platform 414 and the parameters for that model can be determined based on configuration data. The configuration data may be included within the input data 410 and / or may be identified using the input data 410 and / or may be set based on one or more configuration files parsed by the machine learning model platform 414.
[0101] In some cases, the machine learning model platform 414 may provide additional output depending on the context. In some implementations, the machine learning model platform 414 may be configured to return "reason codes" that capture a human-friendly explanation of the output of the machine learning model with respect to suspect input attributes. For example, the machine learning model platform 414 may indicate which of one or more input elements or input units within the input representation contributed to the model output, such as the "amount" channel exceeding a learned threshold, and the set of "merchant" elements or units (such as embeddings or indices) being outside of a given cluster. If the machine learning model platform 414 implements a decision tree, these additional outputs may include the path through the decision tree, or aggregate feature importance based on an ensemble of trees. In the case of a neural network architecture, this may include layer output activations and / or layer filters with positive activations.
[0102] In FIG. 4, some implementations may include an optional alarm system 418 that receives a scalar output 416. In other implementations, the scalar output 416 may be passed directly to an output data interface 420 without further processing. In this latter case, the scalar output 416 may be packaged in response to the original request to the input data interface 412. In either case, output data 422 derived from the scalar output 416 is provided as the output of the machine learning system 402. The output data 422 is returned to enable final processing of the transaction data. For example, the output data 422 may be returned to a payment processor server 140 and used as a basis for a decision to approve or reject the transaction. Depending on the implementation requirements, in some cases, the alarm system 418 may process the scalar output 416 and return a binary value indicating whether the transaction should be approved or rejected (e.g., "1" equals reject). In some cases, the decision may be made by applying a threshold to the scalar output 416. This threshold may be context-dependent. In some cases, the alarm system 418 and / or the output data interface 420 may also receive additional inputs, such as explanatory data (e.g., the "reason code" described above) and / or the original input data. The output data interface 420 may generate an output data package for the output data 422 that combines these inputs with the scalar output 416 (e.g., for at least logging and / or later consideration). Similarly, an alarm generated by the alarm system 418 may include, for example, in addition to the scalar output 416, the above-described additional inputs and / or be based thereon additionally.
[0103] The machine learning system 402 can be used in an "online" mode to process a large number of transactions within a narrowly defined time range. For example, under normal processing conditions, the machine learning system 402 can process requests within 7 - 12 ms and manage 1000 - 2000 requests per second (these are median constraints from real-world operating conditions). However, the machine learning system 402 can also be used in an "offline" mode, for example, by providing selected historical transactions to the input data interface 412. In the offline mode, the input data can be passed to the input data interface in batches (i.e., groups). The machine learning system 402 can also implement a machine learning model that provides scalar outputs for entities as well as transactions, or for entities instead of transactions. For example, the machine learning system 402 may receive requests associated with an identified user (e.g., cardholder or payment account holder) or an identified merchant, and be configured to provide a scalar output 416 indicating the likelihood that the user or merchant is fraudulent, malicious, or abnormal (i.e., a general threat or risk). For example, this can form part of an ongoing or periodic monitoring process, or a one-time request (e.g., as part of an application for a service). The provision of scalar outputs for specific entities can be based on a set of transaction data up to the last approved transaction containing it within a sequence of transaction data (e.g., transaction data for entities similar to those shown in FIG. 3A).
[0104] Exemplary Transaction Process Flow Figures 5A and 5B show two possible transaction process flows 500 and 550, respectively. These process flows can occur in the context of the exemplary transaction process systems 100, 102, 104 shown in FIGS. 1A - 1C, as well as other systems. Process flows 500 and 550 are provided as an example of a context in which a machine learning transaction processing system can be applied, but not all transaction process flows necessarily follow the processes shown in FIGS. 5A and 5B, and the process flows may vary between implementations, between systems, and over time. The exemplary transaction process flows 500 and 550 reflect two possible cases, namely, a first case represented by transaction process flow 500 where the transaction is approved, and a second case represented by transaction process flow 550 where the transaction is rejected. Each transaction process flow 500, 550 involves the same set of five interacting systems and devices, namely, a POS or user device 502, a merchant system 504, a payment processor (PP) system 506, a machine learning (ML) system 508, and an issuing bank system 510. The POS or user device 502 may comprise one of the client devices 110, the merchant system 504 may comprise a merchant server 130, the payment processor system 506 may comprise a payment processor server 140, and the machine learning system 508 may comprise an implementation of machine learning systems 160, 210, and / or 402. The issuing bank system 510 may comprise one or more server devices that perform transaction functions on behalf of the issuing bank. The five interacting systems and devices 502 - 510 can be communicatively coupled by one or more internal or external communication channels, such as network 120. In some cases, some of these systems may be combined; for example, the issuing bank may also perform the function of the payment processor, and thus, systems 506 and 510 can be implemented together with a common system.In other cases, a similar process flow may be executed, particularly for merchants (e.g., without involving a payment processor or issuing bank). In this case, the machine learning system 508 may communicate directly with the merchant system 504. In these variations, the general functional transaction process flow may remain similar to that described below.
[0105] The transaction process flows in both FIGS. 5A and 5B include several common (i.e., shared) processes 512-528. In block 512, the POS or user device 502 initiates a transaction. In the case of a POS device, this may include the cashier attempting to conduct an electronic payment using a front-end device, and in the case of the user device 502, this may include the user making an online purchase (e.g., clicking "complete" in an online basket) using a credit or debit card, or an online payment account. In block 514, payment details are received as electronic data by the merchant system 504. In block 516, the transaction is processed by the merchant system 504 and a request is made to the payment processor system 506 to authorize the payment. In block 518, the payment processor system 506 receives the request from the merchant system 504. The request may be made via a proprietary communication channel or as a secure request (e.g., an HTTPS request over the Internet) over a public network. The payment processor system 506 then requests a score or probability for use in processing the transaction from the machine learning system 508. Block 518 may further include retrieving auxiliary data for combination with the transaction data sent to the machine learning system 508 as part of the request. In other cases, the machine learning system 508 may have access to a data storage device storing auxiliary data (e.g., similar to the configurations of FIGS. 2A and 2B), and thus may retrieve this data as part of its internal operations (e.g., based on an identifier such as provided and / or implemented within the transaction data and defined as part of a machine learning model).
[0106] Block 520 represents the model initialization operation that occurs prior to any request from the payment processor system 506. For example, the model initialization operation may include loading a defined machine learning model and the parameters that instantiate the defined machine learning model. In block 522, the machine learning system 508 receives a request from the payment processor system 506 (e.g., via a data input interface such as 412 in FIG. 4). In block 522, the machine learning system 508 may perform any defined preprocessing prior to the application of the machine learning model initialized in block 520. For example, if the transaction data still holds character data, such as a merchant identified by a string or character transaction description, this may be converted to suitable structured numerical data (e.g., by converting string categorical data to identifiers via a lookup operation or other mapping and / or by mapping characters or groups of characters to vector embeddings). Next, in block 524, the machine learning system 508 supplies the input data derived from the received request to the model to apply the instantiated machine learning model. This may comprise applying a machine learning model platform 414 as described with reference to FIG. 4. In block 526, a scalar output is generated by the instantiated machine learning model. This may be processed in the machine learning system 508 to determine an "approve" or "deny" binary decision or, preferably, returned to the payment processor system 506 as a response to the request made in block 518.
[0107] In block 528, the output of the machine learning system 508 is received by the payment processor system 506 and used to approve or reject a transaction. FIG. 5A shows the process by which a transaction is approved based on the output of the machine learning system 508, and FIG. 5B shows the process by which a transaction is rejected based on the output of the machine learning system 508. In FIG. 5A, in block 528, the transaction is approved. Then, in block 530, a request is made to the issuing bank system 510. In block 534, the issuing bank system 510 approves or rejects the request. For example, the issuing bank system 510 may approve the request if the end user or cardholder has sufficient funds and authorization to cover the transaction cost. In some cases, the issuing bank system 510 may apply a second level of security, although this may not be necessary if the issuing bank relies on the fraud detection performed by the payment processor using the machine learning system 508. In block 536, permission from the issuing bank system 510 is returned to the payment processor system 506, and the payment processor system 506 sends a response to the merchant system 504 in block 538, and the merchant system 504 responds to the POS or user device 502 in block 540. If the issuing bank system 510 approves the transaction in block 534, the transaction may be completed, and a positive response may be returned to the POS or user device 502 via the merchant system 504. The end user may receive this as a "permitted" message on the screen of the POS or user device 502. The merchant system 504 may then complete the purchase (e.g., initiate internal processing to perform the purchase).
[0108] At a later point in time, one or more of the payment processor system 506 and the machine learning system 508 may store data related to a transaction, for example, as part of the transaction data 146, 240, or 300, and / or the category history data 241 (described in more detail below). This is shown in dashed blocks 542 and 544. The transaction data and / or the category history data may be stored together with one or more of the output of the machine learning system 508 (e.g., a scalar fraud or anomaly probability) and the final result of the transaction (e.g., whether the transaction was approved or rejected). The stored data may be stored for use as training data (e.g., as a basis for training data) for a machine learning model implemented by the machine learning system 508. The stored data may also be accessed as part of a future iteration of block 524, for example, and may form part of future auxiliary data. In some cases, the final result or outcome of a transaction may not be known at the time of the transaction. For example, a transaction may only be labeled as an anomaly through subsequent review by an analyst and / or an automated system, or based on feedback from a user (e.g., when the user reports fraud or indicates that a payment card or account has been at risk on a certain date). In these cases, ground truth labels for the purpose of training the machine learning system 508 may be collected over time following the transaction itself.
[0109] Next, referring to the alternative process flow of FIG. 5B, in this case, one or more of the machine learning system 508 and the payment processor system 506 rejects the transaction based on the output of the machine learning system 508. For example, if the scalar output of the machine learning system 508 exceeds the extracted threshold, the transaction can be rejected. At block 552, the payment processor system 506 issues a response to the merchant system 504, and that response is received at block 554. At block 554, the merchant system 504 proceeds with steps to prevent the transaction from completing and returns an appropriate response to the POS or user device 502. This response is received at block 556, and the end user or customer can be notified that their payment has been rejected, for example, via a "rejected" message on the screen. The end user or customer can be prompted to use a different payment method. Although not shown in FIG. 5B, in some cases, the issuing bank system 510 can be notified that a transaction related to a particular account holder has been rejected. The issuing bank system 510 can be notified as part of the process shown in FIG. 5B or as part of a periodic (e.g., daily) update. The transaction may not become part of the transaction data 146, 240, or 300 (when not approved), but as shown by block 544, can still be logged at least by the machine learning system 508. For example, with respect to FIG. 5A, the transaction data can be stored along with the output of the machine learning system 508 (e.g., scalar fraud or anomaly probability) and the final result of the transaction (e.g., that the transaction was rejected).
[0110] Machine learning system with a category history module Some examples described herein, such as machine learning systems 160, 210, 402, and 508 in FIGS. 1A - 1C, FIGS. 2A and 2B, FIG. 4, and FIGS. 5A and 5B, can be implemented as a modular platform that allows different machine learning models and configurations to be used to provide the transaction processing described herein. This modular platform can allow different machine learning models and configurations to be used as technology improves and / or based on certain characteristics of the available data.
[0111] A machine learning system 600 according to an embodiment of the present invention is shown in FIGS. 6 and 7, as will be described in more detail below.
[0112] As previously explained, in certain embodiments described below, the state data is stored in a table of vectors, but note that the state data can be readily structured as an n - dimensional tensor and / or stored in any other data structure or format, such as in a multi - dimensional array. Similarly, the various vectors (e.g., input vectors, output vectors, etc.) described for this embodiment can also be n - dimensional tensors and can be stored using any suitable data structure.
[0113] The machine learning system 600 receives input data 601 and maps it to a scalar output 602. This general - purpose processing follows the same framework as previously described. The machine learning system 600 comprises a first processing stage (or "lower model") 603 and a second processing stage (or "upper model") 604.
[0114] The machine learning system 600 is applied to at least the data associated with the proposed transaction to generate a scalar output 602 for the proposed transaction. The scalar output 602 represents the likelihood that the proposed transaction presents a behavioral anomaly, e.g., the proposed transaction implements a pattern of actions or events that is different from the expected or typical pattern of actions or events. In some cases, the scalar output 602 represents the probability that the proposed transaction presents an anomaly in a sequence of actions, where the actions include at least the previous transaction for that entity and entity-category pair, as outlined in more detail below. The scalar output 602 may be used to complete an approval determination for the proposed transaction, e.g., to determine whether to approve or reject the proposed transaction, as described with reference to FIGS. 5A and 5B.
[0115] The input data may include transaction time data and transaction feature data. The transaction time data may comprise data derived from a timestamp such as that shown in FIG. 3B, or any other data format representing the date and / or time of the transaction. The date and / or time of the transaction may be set as the time at which the transaction was initiated on a client computing device such as 110 or 502, or as the time at which the request was received in the machine learning system 600, similar to the time of the request as received at block 522 in FIGS. 5A and 5B. The transaction time data may comprise time data for a plurality of transactions, such as the current proposed transaction and a historical set of one or more previous transactions. The time data for the historical set of one or more previous transactions may be received with the request and / or retrieved from a storage device communicatively coupled to the machine learning system 600.
[0116] In the configuration of FIG. 6, the transaction time data is converted into relative time data for the application of one or more neural network architectures. Specifically, the transaction time data is converted into a set of time difference values, where the time difference values represent the time difference Δt between the current proposed transaction and a previous transaction (generally, the most recent previous transaction). For example, the time difference may comprise a normalized time difference in seconds, minutes, or hours. The time difference Δt can be calculated by subtracting the timestamp of the previous transaction (or, if multiple differences are being calculated and used, multiple transactions) from the timestamp for the proposed transaction. The time difference values can be normalized by dividing by the maximum predefined time difference and / or clipped at the maximum time difference value.
[0117] In inference mode, the machine learning system 600 uses a top model 604 and a bottom model 603, as outlined in more detail below, to detect a set of features from the input data 601. The output from the top model 604 is a scalar value 602, which is used when determining whether to block a transaction (e.g., by comparing the scalar value 602 to a threshold).
[0118] As shown in FIG. 6, the machine learning system 600 comprises a feed-forward neural network-based embedding layer within the bottom model 603. In this particular embodiment, the embedding layer is a first multi-layer perceptron 610, although it will be understood that this could alternatively be any other type of feed-forward neural network, such as a transformer layer or a neural arithmetic logic unit.
[0119] The first multi-layer perceptron 610 comprises a fully-connected neural network architecture with a plurality of neural network layers (e.g., 1 to 10 layers) for preprocessing data for a proposed transaction from input 601. Although multi-layer perceptrons are described herein, in some implementations, preprocessing may be omitted (e.g., if the received transaction data in input data 601 is already in a suitable format) and / or only a single layer of linear mapping may be provided. In some examples, preprocessing may be considered in the form of an "embedding" layer or an "initial mapping" layer for the input transaction data. Optionally, one or more of the neural network layers of the first multi-layer perceptron 610 may provide learned scaling and / or normalization of the received transaction data. Generally, the fully-connected neural network architecture of the first multi-layer perceptron 610 represents a first learned feature preprocessing stage that converts input data associated with a proposed transaction into a feature vector for further processing. Optionally, the fully-connected neural network architecture may learn some relationships, e.g., some correlations, between elements of the transaction feature data and output an efficient representation that takes these correlations into account. In some cases, the input transaction data may comprise integers and / or floating-point numbers, and the output 612 of the first multi-layer perceptron 610 may comprise a vector of values between 0 and 1. The number of elements or units of the first multi-layer perceptron 610 (i.e., the output vector size) may be set as a configurable hyperparameter. Optionally, the number of elements or units may be between 32 and 2048.
[0120] The output 612 created by the first multi-layer perceptron 610 is divided into two parts 612a, 612b, and the two parts 612a, 612b are provided to the time decay cell 614 and the category history module 616 respectively. However, although they are shown as being completely separable, it will be understood that a configuration with at least some overlap is envisioned in the parts of the output 612 supplied to each of the time decay cell 614 and the category history module 616.
[0121] Note that the input transaction data of the input 601 may comprise data from a proposed transaction (such as received via one or more of, for example, a client device, a POS device, a merchant server device, and a payment processor server device), and data associated with the proposed transaction not contained within the data packet of the proposed transaction. For example, as described with reference to FIGS. 1A - 1C and FIGS. 2A and 2B, the input 601 may further comprise auxiliary data (such as 148 and 242), where the auxiliary data is retrieved by the machine learning system 600 (as shown in FIG. 1C) and / or by the payment processor server 140 (as shown in FIG. 1A or FIG. 1B). The exact content contained within the input data 601 may vary between implementations. This example relates to the general technical architecture for the processing of such data rather than the exact format of the transaction data. Generally, when the machine learning system 600 comprises a set of neural network layers, the parameters can be learned based on what input data configuration is desired or based on what input data is available. This example relates to the engineering design of such a technical architecture to enable transaction processing at the speeds and scales described herein.
[0122] In the example of FIG. 6, the output of the first multi-layer perceptron 610 is received by the time decay cell 614. The time decay cell 614 is based on the previous state data h t-1 , the time difference Δt, and a portion of the output 612a supplied by the first multi-layer perceptron 610, and is configured to generate the output data o t . The time decay cell 614 may implement the techniques described in WO / 2022 / 008130 and WO / 2022 / 008131, each of which is incorporated herein by reference.
[0123] The previous state data h t-1 is a vector of values corresponding to the state of the system from the previous transaction. Generally, those skilled in the art will understand that the previous state data h t-1 is stored in memory and provides a history of the patterns of transactions that have been performed. The state data is stored for each primary entity (e.g., cardholder), and thus the state data for a given primary entity provides information regarding the previous chain of transactions for that primary entity.
[0124] The degree of influence that the historical state data should have on the determination related to the proposed transaction may depend on the time difference Δt. If a long period of time has elapsed since the previous transaction, some of the state data may be less relevant when assessing whether the proposed transaction is abnormal or fraudulent. The time decay cell 614 applies a decay function to the previous state data h t-1 based on the time difference Δt.
[0125] When the previous state data h t-1 is decayed based on the time difference Δt, the time decay cell 614 performs a mapping to generate the output data o t . The time decay cell 614 also stores new state data h t in memory for the next transaction.
[0126] The mapping process performed by the time decay cell 614 can be machine learning-based. For example, the time decay cell 614 can use neural network mapping (e.g., one or more parameterized functions) to map the previous state data h t-1 and the output 612a from the first multi-layer perceptron 610 to a fixed-size vector output of a predefined size. However, it will be understood that other methods may be used that do not require any such neural network or other machine learning technique.
[0127] A portion 612b of the output from the first multi-layer perceptron 610 is input to the category history module 616 and is specifically used by the update layer 619. Unlike the time decay cell 614, which operates using a single state data vector for each entity, the category history module 616 constructs a table of multiple states (e.g., a lookup table), one for each secondary category that a primary entity has interacted with it, as can be better understood with reference to FIG. 7.
[0128] In a typical fraud application example, the primary entity represents the payer or account holder, while the secondary category may represent the beneficiary of the funds, e.g., the merchant. However, it should be understood that the secondary category can be any categorical field in the data. It can be, for example, a numerical or date-time field bucketed into different categories, e.g., "amount" can be bucketed into "low", "medium", and "high" categories, and date and time can be bucketed into time zones or weekdays, etc. The category can also be a composite of multiple categories (e.g., merchant identification information concatenated with the day of the week). The category does not need to indicate an order. Hereinafter, in this specific example, for clarity of notation, we will refer to the secondary category field as "merchant". Note that state data will be stored and retrieved for a specific entity-category pair so that all states remain associated with the primary entity. Thus, generally, it should be understood that there will be such a table for each primary entity.
[0129] In one exemplary software-based implementation, this table is a 2D tensor of shape (merchant_state_size, max_num_merchant_state) whose columns represent the individual states for different merchants (of course, it should be understood that in other examples, different secondary categories can be used). "max_num_merchant_state" is understood to be the maximum number of merchants (or equivalently, other categorical values such as day of the week, transaction amount, or price range) for which state data is stored, i.e., the number of columns in the table. On the other hand, "merchant_state_size" is the number of independent values stored in each column, i.e., the number of rows in the table. As described above, generally, there will be such a state data table for each primary entity (i.e., for each account holder), so it should be understood that the stored merchant (i.e., secondary category) data becomes unique to that primary entity.
[0130] The table in this example has a fixed size and thus has a limit on how many merchant states it can store. As input, it takes the lookup table G (for this entity) from the previous time step t - 1 t-1 and an update vector of size merchant_state_size. As output, it returns a vector containing the updated merchant state. In some embodiments, this vector provides, for example, the concatenation of the decayed merchant state (for the current merchant) and the updated merchant state, i.e., a vector of size 2 * merchant_state_size, and may also include the decayed merchant state by providing
[0131] As shown in FIG. 7, for each transaction, the category history module 616 executes decay and update logic, similar to the standard time decay cell 614. The category history module 616 first applies exponential time decay to all stored states in the table G t-1 where the decay is based on the time difference Δt from the last transaction.
[0132] The decay function enables selective and tuned use of the previous transaction data (i.e., can apply a certain degree of "forgetting") based on the time difference Δt between the proposed transaction and the previous transaction, where the previous transaction data is an aggregated function of the transaction feature vectors for a plurality of previous transactions.
[0133] The decay function has the form f(s)=[e (-s / w_1) ,e (-s / w_2) ,...,e (-s / w_d)As a vectorized exponential time decay function, i.e., as a series of parameterized time decay calculations using [w_1, w_2... w_d] with a set of weights representing the "decay length" in seconds, where d is the size of the recursive state vector (which can generally be equal to the length of the state data vector in which the decay function operates). In this case, Δt is passed as a function of s to output a vector of length d and can be passed, for example, as a scalar floating-point value representing the number of seconds elapsed between a previous event (e.g., a transaction) and the current event (e.g., a proposed transaction) for the same entity. The decay coefficients (i.e., the weights) can be learned during the training process.
[0134] The decay weight can be a fixed (e.g., manually configurable) parameter and / or can have trainable parameters. In the case of an exponential function, the decay length can include (or be trained to include) a mixture of different values representing decay from a few minutes to several weeks so that both short-term and long-term behavior are captured. The time decay function is applied to the state data table from the previous iteration using an element-wise multiplication (e.g., an Hadamard product is calculated) between the time decay vector and each state data entry (each category, i.e., the state data stored for a column of the table). In this example, the state data comprises a two-dimensional table of vectors, each of which is the same length as the transaction feature vector output. This effectively controls how much of the state data from the previous iteration is remembered and how much is forgotten.
[0135] Various time decay constants can be used to decay the table. For example, decay on a short time scale serves to suppress the memory contribution from old activities, while decay on a long time scale serves to retain the memory contribution from old activities. This enables the network to compare new and old activities to determine novelty. Further, state updates are a function of the input attributes of the event. This means that some state elements can be updated only for some types of events. This is graphically shown in FIG. 7, such as different groupings of rows in a table, where the top three rows are retained for a relatively long time period, the middle three rows are retained for an intermediate time period, and the bottom three rows are retained for a relatively short time period. However, this is just an example, and it will be understood that different time periods can be appropriately applied to each part of the state data.
[0136] The network can learn that only a particular type of activity (e.g., low-price transactions) with a given merchant has been seen before, and thus, even if the merchant was present before, the current event is new. For example, the network can learn to ignore the existence of past low-price transactions with this merchant when risk scoring the current activity as being high-price. The actual decay applied can be learned or pre-determined (and can be defined by the user, for example, during the initial setup or optimization of the system). In either case, the machine learning system can learn which decay weights should be applied to which combinations of behavioral features.
[0137] In this exemplary embodiment, while the entire table for a given entity is decayed, other embodiments are envisioned where only the currently selected column is decayed. To achieve this, a timestamp may be stored for each entity-category pair, and the time difference Δt is taken as the time between the current transaction and the time when data for the most recent instance of that entity-category pair was stored in the table.
[0138] After the decay step, the category history module 616 extracts the merchant state column that matches the current merchantId (or, if a new merchant as outlined below, a zero vector). The category history module 616 queries the lookup table using the current merchantId to ascertain whether the table contains state data for this merchant (i.e., whether the column exists). If so, the state data stored for that merchantId is retrieved from the table. Conversely, if it does not exist in the table, a new "default" state (e.g., a zero vector) can be added to the table.
[0139] Next, the (appropriately retrieved or new) state data is updated by the update layer 619 using the output portion 612b from the multi-layer perceptron 610 (derived from the input vector x as previously outlined) and inserted back into the table. Different approaches can be taken for the update step (i.e., the steps performed by the update layer 619), but in one implementation, the decayed state data vector for the entity-category pair and the input vector from the first multi-layer perceptron 610 (i.e., the output portion 612b) are added together using element-wise addition. This results in the output vector y t being produced. t
[0140] The table is managed as a stack, and updated columns are inserted at the front of the table, with the most recently updated column found at the back of the table. Due to the table size limit, when the table is full, the state for the most recently seen merchant is discarded, and a new merchant state is added. As seen in Figure 7, the updated state data for this transaction is inserted at the front of the table, and the "gap" left behind (as indicated by the dashed line) is closed.
[0141] The outputs o of the time decay cell 614 and the category history module 616 t , y t are received by the upper model 604 and, in particular, by the output layer of the second feedforward neural network. In this particular embodiment, the output layer is the second multilayer perceptron 620, although it will be understood that this could alternatively be any other type of feedforward neural network, such as a transformer layer or a neural arithmetic logic unit. In this particular embodiment, the input data 601 is also supplied to the second multilayer perceptron 620. In this example, the outputs o of the time decay cell 614 and the category history module 616 t , y t are passed directly to the second multilayer perceptron 620. However, other embodiments (not shown) are envisioned where this passing is done indirectly through one or more additional intervening neural network layers.
[0142] The second multilayer perceptron 620 is used to generate a scalar output 602. The second multilayer perceptron 620 receives the outputs o of the time decay cell 614 and the category history module 616 t , y t , as well as the input data 601 in order to receive and their outputs o t , y tTo map to scalar value 602, it comprises a fully-connected neural network architecture with a plurality of neural network layers. This may comprise a series of dimensionality reduction layers. The last activation function in the plurality of neural network layers may comprise a sigmoid activation function for mapping the output to the range of 0 to 1. The second multi-layer perceptron 620 may be trained to extract the correlation between features output by an attention mechanism (e.g., using a plurality of attention heads in a manner known per se in the art) and to apply one or more non-linear functions to finally output the scalar value 602.
[0143] The scalar value 602 may then be used by another part of the transaction processing pipeline (not shown) when determining whether the transaction should be allowed to continue as normal or whether the transaction should instead be stopped and marked as fraudulent or abnormal. For example, the system may compare the scalar value 602 to a suitable threshold such that scores above the threshold are flagged as abnormal.
[0144] In a preferred implementation form, the state data is entity - dependent, i.e., specific to a particular user, account holder, or merchant account. In this way, the appropriate entity may be identified as part of the transaction data pre - processing, and the machine learning system 600 may be configured for that entity. In some cases, the machine learning system 600 may apply the same parameters for the neural network architecture for each entity, but may separately store the state data and retrieve only the historical transaction data and / or auxiliary data associated with (e.g., indexed by) that entity. Thus, the machine learning system 600 may include, for example, an entity state store as described with reference to FIGS. 2A and 2B, whereby data for different entities can be effectively segregated. In other cases, the parameters for the decay function may be shared across multiple entities. This can be advantageous for reducing the number of learnable parameters for each entity of the overall machine learning system. For example, the exponential time - decay weights described above may be fixed (or learned) and shared across multiple entities, while the state data may be specific to each entity.
[0145] Training of a machine learning system for processing transaction data In some examples, a machine learning system, such as those described herein, can be trained using labeled training data. For example, a training set can be provided that includes data associated with transactions labeled as "normal" or "fraudulent". In some cases, these labels can be provided based on reported fraudulent activities, i.e., using past reports of fraudulent behavior. For example, data associated with transactions that have been approved and processed without subsequent reports of fraud can be labeled as "0" or "normal", whereas data associated with transactions that have been rejected and / or later reported as fraudulent and marked or blocked in another way can be labeled as "1" or abnormal.
[0146] The end-to-end training process can be performed so that the machine learning system can learn appropriate values for any learnable parameters. This can include, for example, training the parameters of any multi-layer perceptron and / or the decay weights for a decay function. Any suitable training algorithm known in the art can be used, and it should be understood that the principles of the present invention are not limited to any particular training algorithm. However, by way of example, a suitable training algorithm can be based on stochastic gradient descent and can iteratively improve the model parameters using backpropagation of the error in the model predictions to find the model parameters that minimize the objective function. In some cases, the objective function (which measures the error in the model predictions) can be defined to be binary cross-entropy between the predicted value 602 and the ground truth label ("0" or "1").
[0147] Those skilled in the art will appreciate that when references are made to the various rows and columns of the state data table, the respective roles of these may be interchanged, such that the state data is stored in rows indexed by category identifiers, while state data values for the various features are instead stored across columns.
[0148] It will be appreciated that embodiments of the present invention may provide a configuration in which a category history module brings signal extraction from interactive monitoring to the architecture. Advantageously, this may lead to an improvement in the speed and / or confidentiality of implementation. Specifically, the data scientist no longer needs to manually design this type of feature, which means that this type of feature does not need to be disclosed for model management as an input to the classifier, reducing the implementation time for the model.
[0149] In addition, embodiments of the present invention may provide an improvement in detection performance. The category history module can learn more subtle interaction logics than can be manually designed by humans in a finite time. Specifically, the category history module can learn the patterns of user activities specific to each category and detect anomalies based on the history data for the relevant entity-category pairs. This approach may provide a significant increase in detection compared to other approaches known per se in the art.
[0150] Although specific embodiments of the present invention have been described in detail, those skilled in the art will appreciate that the embodiments described in detail are not a limitation of the scope of the claimed invention.
Explanation of Reference Numerals
[0151] 100 Transaction processing system, transaction process system 102, 104 Transaction processing system, variant form, transaction process system 110 Client device 110-A Smartphone 110-B Computer 110-C Point-of-Sale (POS) System, POS System 110-D Portable Merchant Device 120 Computer Network, Network 120-A First Computer Network 120-B Second Network 130 Merchant Server 140 Payment Processor Server 142 First Data Storage Device 144 Second Data Storage Device 146, 330 Transaction Data 148, 172-A~172-N, 242, 242-A~N Auxiliary Data 150 Machine Learning Server 160, 210, 402, 600 Machine Learning System 162, 410 Input Data 164, 422 Output Data 170 Local Data Storage Device, Data Storage Device 170-A~170-N Local Data Storage Devices 172 Categorized Auxiliary Data, Auxiliary Data 200, 202 Data Storage Configuration 220 Data Bus 230, 231, 232, 260 Data Storage Devices 240 Entity Transaction Data, Transaction Data 240-A~N Entity Transaction Data 241 Category History Transaction Data, Category History Data 241-A~N Category History Transaction Data 250 Data Transfer Channel 270 Secure Logical Memory Layer, Logical Memory Layer 280, 280-A~N Categories 300 Chronological Records, Transaction Data 310 Current transaction data 320 "Older" or historical transaction data 400 An example 412 Input data interface 414 Machine learning model platform 416 Scalar output 418 Alarm system 420 Output data interface 502 POS or user device, user device 504 Merchant system 506 Payment processor (PP) system, payment processor system, system 508 Machine learning (ML) system, machine learning system 510 Issuing bank system, system 601 Input data, input 602 Scalar output, scalar value, predicted value 603 First processing stage (or "lower model"), lower model 604 Second processing stage (or "upper model"), upper model 610 First multi-layer perceptron, multi-layer perceptron 612 Output 612a Part, output 612b Part, output part 614 Time decay cell 616 Category history module 619 Update layer 620 Second multi-layer perceptron
Claims
1. A machine learning system for processing data corresponding to arrival transactions associated with an entity, the machine learning system comprising a category history module configured to process at least some of the data, the category history module comprising: i) A memory configured to store state data for a plurality of categories indexed by respective category identifiers, wherein for each category identifier, the state data stored in the memory corresponds to the entity and previous transactions associated with each category; a memory; ii) A decay logic stage configured to modify the state data stored in the memory based on a time difference between the time of the arrival transaction and the time of the previous transaction, the decay logic stage comprising: a) When the memory contains state data for an entity and category identifier pair associated with the arrival transaction, the decay logic stage is configured to retrieve the state data, apply a decay function to the state data to generate a decayed version of the state data, and output the decayed version of the state data, the decay function depending on the time difference; a decay logic stage; iii) An update logic stage configured to use an input tensor generated from the data corresponding to the arrival transaction to update the state data output from the decay logic stage to generate updated state data, store the updated state data in the memory for the entity and category identifier pair associated with the arrival transaction, and output an output tensor containing the updated state data; Comprising; iv) The machine learning system is configured to map the output tensor from the category history module to a scalar value representing the likelihood that the arrival transaction presents an anomaly within a sequence of actions; v) A machine learning system, wherein the scalar value is used to determine whether to approve or reject the arrival transaction.
2. The decay logic stage is: b) The machine learning system according to claim 1, further configured to generate new state data and output the new state data when the memory does not contain state data for entity and category identifier pairs associated with the arriving transaction.
3. The machine learning system according to claim 1 or 2, further comprising a first neural network stage configured to apply a neural network layer having respective plural learned weights to the data corresponding to the arriving transaction to generate the input tensor and provide the input tensor to the update logic stage of the category history module.
4. The machine learning system according to any one of claims 1 to 3, further comprising a second neural network stage configured to apply a second neural network layer having respective plural learned weights to the output tensor generated by the update logic stage of the category history module to generate the scalar value representing the likelihood that the arriving transaction presents an anomaly within a sequence of actions.
5. The update logic stage comprises a third neural network stage configured to use the decayed state data and the input tensor derived from the arriving transaction to generate the updated state data and, optionally, the third neural network stage comprises a recurrent neural network. The machine learning system according to any one of claims 1 to 4.
6. The machine learning system according to any one of claims 1 to 5, wherein the decay logic stage is configured to change the state data stored in the memory for the entity associated with the arriving transaction based on a time difference between the time of the arriving transaction and the time of the most recent transaction for that entity.
7. The machine learning system according to any one of claims 1 to 5, wherein the decay logic stage is configured to change all state data stored in the memory for each entity based on a time difference between the time of the arriving transaction and the time of the immediately preceding transaction.
8. The decay logic stage is configured to modify only the state data stored in the memory for entity and category identifier pairs associated with the arrival transaction, and the decay logic stage modifies the state data based on a time difference between the time of the arrival transaction and the time of the previous transaction for the entity and category identifier pair associated with the arrival transaction. The machine learning system according to any one of claims 1 to 5.
9. The machine learning system according to any one of claims 1 to 8, wherein the memory is configured as a stack, the retrieved state data is removed from the stack, and the updated state data is stored at the top of the stack.
10. The machine learning system according to any one of claims 1 to 9, wherein the category history module is configured such that when the memory is full, the state data stored for at least one category identifier is erased.
11. The machine learning system according to claim 10, wherein the category history module is configured such that when the memory is full, the state data stored for the most recently seen category identifier is erased.
12. The machine learning system according to any one of claims 1 to 11, wherein the state data stored for each entity and category identifier pair includes a tensor of values.
13. The machine learning system according to claim 12, wherein the decay function applies a respective decay multiplier to each of the values.
14. The machine learning system according to claim 13, wherein at least some of the decay multipliers are learned.
15. The machine learning system further comprises a time decay cell module configured to process at least some of the data corresponding to the arrival transaction, and the time decay cell module a second memory configured to store second state data corresponding to the previous transaction, and a second decay logic stage configured to modify the second state data stored in the second memory based on a time difference between the time of the arrival transaction and the time of the previous transaction The machine learning system according to any one of claims 1 to 14, comprising **Claim 16** wherein the time decay cell module is a fourth neural network stage configured to determine next state data using the previous state data changed by the second decay logic stage and a second input tensor derived from the arrival transaction, and the fourth neural network stage is further configured to store the next state data in the second memory. Optionally, the machine learning system according to claim 15, wherein the fourth neural network stage comprises a recurrent neural network. **Claim 17** The machine learning system according to claim 15 or 16, wherein the machine learning system is configured to map output data from both the time decay cell module and the category history module to the scalar value representing the likelihood that the arrival transaction presents an anomaly within a sequence of actions. **Claim 18** The memory is configured to store state data for a plurality of entities, and for each entity, a plurality of categories indexed by respective category identifiers are stored, and the state data stored in the memory for each category identifier corresponds to previous transactions associated with each entity and each category. The machine learning system according to any one of claims 1 to 17. **Claim 19** Each of the plurality of categories comprises one or more of a secondary entity, a secondary entity type, a transaction price, a transaction price band, a day of the week, a time, a time zone, and a time window. The machine learning system according to any one of claims 1 to 18. **Claim 20** One or more of the plurality of categories includes a composite of one or more of a secondary entity, a secondary entity type, a transaction price, a transaction price band, a day of the week, a time, a time zone, and a time window. The machine learning system according to claim 19. **Claim 21** A method of processing data associated with a proposed transaction associated with an entity, comprising i) Storing in memory state data for a plurality of categories indexed by respective category identifiers, wherein for each category identifier, the state data stored in the memory corresponds to the entity and the previous transaction associated with each category; ii) Modifying the state data stored in the memory based on a time difference between the time of the arriving transaction and the time of the previous transaction, wherein the step of modifying the state data a) When the memory contains state data for the entity and category identifier pair associated with the arriving transaction, retrieving the state data, applying an attenuation function to the state data to generate an attenuated version of the state data, and outputting the attenuated version of the state data, wherein the attenuation function depends on the time difference; iii) Updating the state data output using the input tensor generated from the data corresponding to the arriving transaction to generate updated state data, storing the updated state data in the memory for the entity and category identifier pair associated with the arriving transaction, and outputting an output tensor containing the updated state data; iv) Mapping the output tensor to a scalar value representing the likelihood that the arriving transaction presents an anomaly within a sequence of actions; v) Determining whether to approve or reject the arriving transaction using the scalar value A method comprising the above steps. **Claim 22** The method according to claim 21, further comprising applying a neural network layer having respective plural learned weights to the data corresponding to the arriving transaction to generate the input tensor. **Claim 23** The method according to claim 21 or 22, further comprising applying a second neural network layer having respective plural learned weights to the output tensor to generate the scalar value representing the likelihood that the arriving transaction presents an anomaly within a sequence of actions. **Claim 24** changing, based on a time difference between the time of the arrival transaction and the time of the most recent transaction for the entity, the state data stored in the memory for the entity associated with the arrival transaction, the method according to any one of claims 21 to 23.
25. changing all state data stored in the memory for each entity based on a time difference between the time of the arrival transaction and the time of the immediately preceding transaction, the method according to any one of claims 21 to 23.
26. changing only the state data stored in the memory for the entity and category identifier pair associated with the arrival transaction, the attenuation logic stage changing the state data based on a time difference between the time of the arrival transaction and the time of the previous transaction for the entity and category identifier pair associated with the arrival transaction, the method according to any one of claims 21 to 23.
27. A method of processing data associated with a proposed transaction, receiving an arrival event from a client transaction processing system, the arrival event being associated with a request for an approval determination for the proposed transaction, analyzing the arrival event to extract data for the proposed transaction, including determining a time difference between the proposed transaction and a previous transaction, applying a machine learning system according to any one of claims 1 to 17 to output a scalar value representing the likelihood that the proposed transaction presents an anomaly within a sequence of actions, based on the extracted data and the time difference, the applying including accessing the memory to retrieve at least the state data for the entity and the category identifier associated with the transaction, A step of determining a binary output based on the scalar value output by the machine learning system, wherein the binary output indicates whether the proposed transaction is approved or rejected; A step of returning the binary output to the client transaction processing system; A method comprising the above.
28. A category history module for use in a machine learning system for processing data corresponding to incoming transactions associated with an entity, the category history module being configured to process at least some of the data, the category history module comprising: i) A memory configured to store state data for a plurality of categories indexed by respective category identifiers, wherein for each category identifier, the state data stored in the memory corresponds to the entity and previous transactions associated with each category; ii) A decay logic stage configured to modify the state data stored in the memory based on the time difference between the time of the incoming transaction and the time of the previous transaction, the decay logic stage comprising: a) When the memory contains state data for the entity and category identifier pair associated with the incoming transaction, the decay logic stage is configured to retrieve the state data, apply a decay function to the state data to generate a decayed version of the state data, and output the decayed version of the state data, the decay function depending on the time difference; iii) An update logic stage configured to use the input tensor generated from the data corresponding to the incoming transaction to update the state data output from the decay logic stage to generate updated state data, store the updated state data in the memory for the entity and category identifier pair associated with the incoming transaction, and output an output tensor including the updated state data. A category history module comprising the above.
29. A method for processing data associated with a proposed transaction associated with an entity, comprising: i) Storing in memory state data for a plurality of categories indexed by respective category identifiers, wherein for each category identifier, the state data stored in the memory corresponds to the entity and the previous transaction associated with each category; ii) Modifying the state data stored in the memory based on a time difference between the time of the arriving transaction and the time of the previous transaction, wherein the step of modifying the state data: a) When the memory contains state data for the entity and category identifier pair associated with the arriving transaction, retrieving the state data, applying an attenuation function to the state data to generate an attenuated version of the state data, and outputting the attenuated version of the state data, wherein the attenuation function depends on the time difference; iii) Updating the state data output using the input tensor generated from the data corresponding to the arriving transaction to generate updated state data, storing the updated state data in the memory for the entity and category identifier pair associated with the arriving transaction, and outputting an output tensor containing the updated state data; A method comprising the above steps. **Claim 30** A non-transitory computer-readable medium comprising instructions which, when executed by a processor, cause the processor to perform the method according to any one of claims 21 to 27 or claim 29.
Citation Information
Patent Citations
Interleaved sequence recurrent neural networks for fraud detection
US20210248448A1
Neural network architecture for transaction data processing
WO2022008130A1