Method and device for detecting fraudulent user transactions

A machine learning method that collectively analyzes transaction parameters through a coding and generative model to generate reference transactions, enhancing the accuracy and completeness of fraud detection in financial systems.

WO2025193116A1PCT designated stage Publication Date: 2025-09-18PUBLICHNOE AKTSIONERNOE OBSHCHESTVO SBERBANK ROSSII (PAO SBERBANK)

Patent Information

Application Number
PCT/RU2024/000098
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-15
Filing Date
2024-03-27
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing fraud detection systems in financial institutions lack the ability to analyze transaction parameters as a whole and often predict user transactions individually, leading to incomplete and inaccurate identification of fraudulent activities.

Method used

A machine learning-based approach that considers all transaction parameters collectively, using a coding machine learning model to generate a hidden state matrix and a generative model to create a reference transaction, which is then compared with actual transactions to determine anomalies and generate signals for fraudulent activity.

Benefits of technology

Enhances the completeness and accuracy of fraudulent transaction detection by analyzing transaction parameters as a unified set, improving the precision of identifying fraudulent transactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure RU2024000098_18092025_PF_FP_ABST
    Figure RU2024000098_18092025_PF_FP_ABST
Patent Text Reader

Abstract

The invention relates to the field of information security. A method for detecting fraudulent user transactions comprises the steps of: obtaining a sequence of transactions made by a user within a given time frame, each transaction being characterized by a set of attributes; converting each obtained transaction into an attribute vector on the basis of the corresponding set of transaction attributes; processing the obtained attribute vectors with the aid of a coding machine learning model, with a representation of a sequence of attribute vectors being produced in the form of a hidden state matrix; processing the hidden state matrix with the aid of a generative machine learning model (MLM) trained on hidden state matrices of sequences of transaction attribute vectors, with the generative MLM generating from the hidden state matrix interrelated attribute vectors of one or several transactions; generating the last transaction of the user on the basis of the generated attribute vector; comparing the last transaction made by the user and the generated last transaction of the user by calculating an anomaly assessment of the transaction made; signalling a fraudulent transaction if the anomaly assessment value is above a threshold value. The invention provides for more complete and accurate detection of fraudulent transactions.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND DEVICE FOR DETERMINING FRAUDULENT USER TRANSACTIONS AREA OF TECHNOLOGY

[0001] The declared solution relates to the field of information security, in particular to anti-fraud systems. LEVEL OF TECHNOLOGY

[0002] With the development of information technology, IT solutions have begun to have a significant impact on all areas and sectors of life. Currently, various companies and organizations are actively implementing and utilizing IT solutions within their structures.

[0003] For large financial institutions (such as banks), identifying fraudulent user transactions is always a labor-intensive and painstaking undertaking. The quality of this solution determines both the amount of direct losses the bank incurs from fraudulent activity and the loyalty of customers, who are faced with the need to repeatedly verify the legitimacy of their transactions.

[0004] Known in the art is patent US 10482464 BI, "Identification of Anomalous Transaction Attributes in Real-Time with Adaptive Threshold Tuning," patent holder: Wells Fargo Bank, NA, published: November 19, 2019. This patent describes a process for analyzing transaction parameters for comparison with a certain threshold. If the threshold is exceeded, the transaction is considered fraudulent.

[0005] A disadvantage of the known solution in this area of ​​technology is the lack of the ability to consider a specific set of transaction parameters as a single whole and analyze this combination for anomalies and fraud, as well as the fact that the prediction of user transactions is carried out only within the framework of individual transaction parameters. ESSENCE OF THE INVENTION

[0006] The proposed technical solution proposes a new approach to identifying fraudulent transactions. This solution utilizes a machine learning algorithm that, based on previous transaction history, considers all transaction parameters as a whole, generates the next transaction with the most probable parameters, and, by comparing the generated transaction with the actual transaction, issues a verdict on the anomalous nature of the actual transaction. transactions. Then, based on the characteristics of the anomaly of the actual transaction, a signal about a fraudulent transaction or some other signal is generated.

[0007] This solves the technical problem of determining transaction anomalies and identifying fraudulent transactions.

[0008] The technical result achieved by solving this problem is an increase in the completeness and accuracy of identifying fraudulent transactions.

[0009] The claimed technical result is achieved by implementing a computer-implemented method for determining fraudulent user transactions, performed by at least one processor, and containing the following steps: • receive a sequence of completed user transactions over a given time period, with each transaction characterized by a set of attributes; • transform each received transaction into an attribute vector based on the corresponding set of transaction attributes; • process the obtained attribute vectors using a coding machine learning model (CMLM), during which a representation of the sequence of attribute vectors is obtained in the form of a matrix of hidden states; • process the matrix of hidden states using a generative machine learning model (GMLM) trained on matrices of hidden states of sequences of transaction attribute vectors, during which the GMLM generates interconnected attribute vectors of one or more transactions from the matrix of hidden states; generate the last transaction of the user based on the generated attribute vector; compare the last completed transaction of the user and the generated last transaction of the user by calculating an anomaly score for the completed transaction; • generate a signal about a fraudulent transaction if the anomaly score value is greater than the threshold value.

[0010] In one particular example of implementation, KMMO is trained on sequences of attribute vectors. [OOP] In another particular example of implementation, the calculation of the anomaly score of a completed transaction is carried out based on one or more attributes by comparing the distance between the vectors of transaction attributes.

[0012] In another particular example of implementation, the method further comprises the step of restricting access of a user performing at least one fraudulent transaction.

[0013] The claimed technical result is also achieved by implementing a device for determining fraudulent user transactions, containing at least one processor, at least one memory associated with the processor and containing machine-readable instructions, which, when executed by at least one processor, ensure the implementation of the method for determining fraudulent user transactions. BRIEF DESCRIPTION OF DRAWINGS

[0014] NNAa FFiiigg.. 1 shows a block diagram of a computer-implemented method for identifying fraudulent user transactions.

[0015] Fig. 2 illustrates an example of a machine learning coding model.

[0016] Fig. 3 illustrates an example of a generative machine learning model.

[0017] Fig. 4 shows the general diagram of the computing device. IMPLEMENTATION OF THE INVENTION

[0018] Below, concepts and terms necessary for understanding the present invention will be described.

[0019] A machine learning (ML) model is a set of artificial intelligence methods whose characteristic feature is not a direct solution to a problem, but learning through the process of applying solutions to many similar problems.

[0020] A transaction (operation) is an action by a client with a bank or other financial account, which is characterized by a set of attributes (otherwise: characteristics, "features", properties, parameters, variables). Examples of such attributes may be: transaction time; amount; transaction type (withdrawal / replenishment from an ATM, transfer to a client of a certain bank, online payment, purchase in a store with payment through a terminal, opening a debit account, opening a credit account, opening a current account, loan application, etc.); client age; client name; name of the recipient client (in the case of a transfer to a client of a certain bank); IP address of the Internet connection of the device for online payment; number of the bank branch that opened the bank account for the client; type The client's bank card information, the card balance, whether the transfer recipient is in the client's mobile app contact list, whether the client's mobile device running the online app has superuser (root) rights, and other data that characterizes the transaction, the client, the payment recipient, the transaction purpose, the location and time of the transaction, the device used to perform the transaction, and more. These attributes can be quantitative (transaction amount and many others) or categorical (qualitative) (username, IP address, and many others).

[0021] A hidden state matrix (LSM) is the output of a machine learning model or part of it, representing a matrix of numbers, each element of which is calculated based on the set of all attributes of all input data objects. Within the framework of this invention, it is specified that the LSM is calculated based on the set of all attributes of all client transactions. However, it is impossible to identify a single matrix element that would correspond to a single transaction attribute or a single transaction in the history. It is impossible to reconstruct the original transaction attribute or transaction history from this matrix, except by training a decoding or generative model.

[0022] A machine learning coding model (MLCM) is a machine learning model that allows one to obtain a hidden state matrix from an ordered set (sequence) of objects characterized by a set of attributes.

[0023] A generative machine learning model (GMLM) is a machine learning model that can extract at least one attribute of a single transaction from a hidden state matrix of a set of transactions. It is also known as a "generative model." "decoding model".

[0024] A completed (actual) transaction is a customer's transaction that has occurred (real, true, valid).

[0025] A generated (reference) transaction is a simulated (imaginary, estimated, predicted) client transaction, in the form of at least one attribute, obtained as a result of the operation of the GMMO from the MCC, given that the MCC was obtained on the basis of at least one previous actual transaction, including, possibly, on the basis of previous generated (reference) transactions.

[0026] Recall is a measure of what proportion of objects of the target class out of all objects of the target class the algorithm found.

[0027] A fraudulent transaction is at least one event that has no apparent economic meaning to the user, and is either not initiated or not confirmed by the user (for example, carried out using the user's device or an account belonging to the user, but without his knowledge), or committed by a user under psychological and / or physical influence of the fraudster, and acting in the interests of the fraudster.

[0028] Preparing data for training.

[0029] The model was trained using historical data on client transactions, where each transaction is divided into two classes: • fraudulent (class 1), • legitimate (class 0).

[0030] The ML model is trained on data in the form of pairs of a historical (time-ordered) set of customer transactions and the class of each operation.

[0031] To assess the model's quality, the dataset was split into two parts: training and validation sets. The split was random, with 70% of the data allocated to the training set and 30% to the validation set.

[0032] The fraud detection rate is around 0.9.

[0033] Fig. 1 shows a computer-implemented method (100) for determining fraudulent user transactions. The method (100) is executed using at least one processor. In a particular example of implementation, the method (100) is executed using a device for determining fraudulent user transactions, which can be implemented on the basis of a computing device modified in its hardware and software so as to perform the functions of a device for determining fraudulent user transactions. A more detailed description of the computing device is disclosed below with reference to Fig. 4.

[0034] At the first stage (101), data is obtained containing at least one completed (actual) transaction kklliiieennttaa. In the preferred embodiment, a sequence of completed user transactions is obtained over a given time period, wherein each transaction is characterized by a set of attributes.

[0035] At step (102), each received transaction is transformed into a vector of attributes based on the corresponding set of transaction attributes. At this stage, categorical variables are processed and represented as a vector.

[0036] At step (102), attributes are represented as sets of categorical variables. At this stage, each attribute in a transaction is represented as an independent categorical variable. In this case, each transaction position In history, a variable is considered a categorical variable. Each such variable at each position can be realized as one and only one value of that variable. Moreover, this variable is not a number or a numerical vector, but rather a category at that position. This category can be specified by a string value or any other means.

[0037] In one embodiment of the invention, step (102) utilizes a "Label Encoder" or "Embedding" method for generating a numerical vector. Embedding types can be generated using technologies such as fully connected, recurrent, convolutional, and transformer neural networks, or pretrained vectors such as Word2Vec, Glove, FastText, and others.

[0038] In one particular embodiment of the invention, a vector representation (embedding) of each categorical variable is obtained, including a contextualized representation (e.g., based on recurrent neural networks, transformers, networks with a self-attention layer, or weighting with neighboring embeddings). The embeddings of the categorical variables are then averaged, and the maximum value is selected among each embedding parameter.

[0039] In one particular embodiment of the invention, a sequence of embeddings of categorical variables is fed to a neural network with convolutional layers such as Convld, Conv2d, Conv3d, ConvTransposeld, ConvTranspose2d, ConvTranspose3d, LazyConvld, LazyConv2d, LazyConv3d, LazyConvTransposeld, LazyConvTranspose2d, LazyConvTranspose3d, Unfold, Fold.

[0040] Next, in one of the particular variants of the invention, there are one or more pooling layers: MaxPool Id, MaxPool2d, MaxPool3d, MaxUnpoolld, MaxUnpool2d, MaxUnpool3d, AvgPoolld, AvgPool2d, AvgPool3d, FractionalMaxPool2d, FractionalMaxPool3d, LPPoolld, LPPool2d, AdaptiveMaxPoolld, AdaptiveMaxPool2d, AdaptiveMaxPool3d, AdaptiveAvgPoolld, AdaptiveAvgPool2d, AdaptiveAvgPool3d [1].

[0041] In one of the particular examples of the implementation of the claimed solution, pre-trained neural networks are used to obtain a vector representation of a command in a sequence of categorical variables, such as DistilBERT: smaller, faster, cheaper, lighter; ALBERT (Lite BERT Google); TinyBERT; T-NLG (Turing Natural Language Generation); USE (Universal Sentence Encoder); ELMo (Embeddings from Language Models), T5, GPT (Generative Pre-trained Transformer), GPT2, GPT3, GPT4, ChatGPT, GigaChat or networks inherited from them.

[0042] In one embodiment of the invention, the TF-IDF (term-frequency times inverse document-frequency) vectorization method is used at step (102) to obtain a numerical vector. The vectorizer was trained on a training set with / without the use of stop values ​​(stop_words=None), with / without the use of IDF, and n-skip-m-grams of categorical variables were used.

[0043] In an alternative embodiment of the claimed solution, step (102) utilizes the One-Hot Encoder method for generating a numerical vector. This encoding method is based on the creation of binary features that indicate membership in a unique value.

[0044] In a particular embodiment of the claimed invention, at step (102), the Binary Encoder method of vector representation of categorical variables is used as a method for obtaining a numerical vector. In this case, N unique values ​​of a categorical feature are encoded using log(N) binary features, representing the ordinal number of the unique value of the categorical variable as a string of the binary representation.

[0045] In one of the particular examples of implementation at step (102), the method for obtaining a numerical vector is a method of vector representation of categorical variables in at least one of the following ways: Contrast Encoding, Helmert Encoder, Backward-Difference Encoder, Target Encoding, Leave-One-Out Encoder, James-Stein Encoder.

[0046] In one alternative embodiment of the invention, at step (102), the vector representations are concatenated, resulting in a single vector of transaction attributes.

[0047] At step (103), the obtained attribute vectors are processed using a coding machine learning model (CMLM), during which a representation of the sequence of attribute vectors is obtained in the form of a matrix of hidden states.

[0048] In a particular embodiment of the claimed invention, a coding machine learning model (CMLM) is trained on sequences of attribute vectors.

[0049] To obtain the hidden state matrix, technologies such as fully connected, recurrent, convolutional, and transformer neural networks are used. In one particular embodiment of the invention, the encoder block is at least one layer of a neural network of the following types: Recurrent neural networks (RNN), Long short-term memory (LSTM), Supervised Gated recurrent units (GRUs), Neural Turing machines (NMTs), Bidirectional RNNs, LSTMs, and GRUs (BiRNNs, BiLSTMs, and BiGRUs), Deep residual networks (DRNs), Echo state networks (ESNs), Liquid state machines (LSMs), and Kohonen self-organizing maps (KNs, or organizing (feature) maps, SOMs, or SOFMs).

[0050] In one particular embodiment of the invention, neural networks such as GNN (Graph Neural Network), DeepWalk, Line, Node2vec, Hope, a neural network with a graph attention mechanism (GAT - Graph ATtention), an inductive learning model (GraphSAGE), RGCN (Relational Graph Convolutional Networks), GCN (Graph Convolutional Networks) or networks inherited from them are used to obtain a matrix of hidden states.

[0051] At step (104), the matrix of hidden states is processed using a generative machine learning model (GMLM) trained on matrices of hidden states of sequences of transaction attribute vectors, during which the GMLM generates interconnected attribute vectors of one or more transactions from the matrix of hidden states.

[0052] The above processing is performed by the generator (decoder) block of the generative machine learning model (GMLM). This block represents at least one of the following: a Variational Autoencoder (VAE) decoder, a deep learning method, a Restricted Boltzmann Machine (RBM), a Deep Belief Network (DBN), or neural networks: Fully Connected, Recurrent, Convolutional, or Transformer.

[0053] In a particular embodiment of the claimed invention, the generator block is at least one layer of a neural network of the following types: Convolutional neural networks (CNN), Recurrent neural networks (RNN), Long short term memory (LSTM), Gated recurrent units (GRU), Neural Turing machines (NMT), Bidirectional RNN, LSTM and GRU (BiRNN, BiLSTM and BiGRU), Deep residual networks (DRN), Echo state networks (ESN), Liquid state machines (LSM), Kohonen self-organizing map (KN, or organizing (feature) map, SOM, SOFM).

[0054] At step (105), the last user transaction is generated based on the generated attribute vector, resulting in a generated (reference) last user transaction.

[0055] In a particular embodiment of the claimed invention, a reference transaction is generated on the basis of all previously completed (actual) transactions.

[0056] In one of the alternative embodiments of the claimed solution, the last (reference) transaction is generated on the basis of some previously completed transactions and reference transactions generated on their basis.

[0057] At step (106), the last completed user transaction and the generated last user transaction are compared by calculating an anomaly score for the completed transaction.

[0058] In a particular embodiment of the claimed invention, the calculation of the anomaly assessment of a completed transaction is carried out based on one or more attributes by comparing the distance between the vectors of transaction attributes.

[0059] In this case, the following metrics are used to calculate the distances: Bray-Curtis, Canberra, Chebyshev, City Block (Manhattan metric, metric of city blocks, cityblock), Correlation, Cosine, Euclidean metric and those inherited from it (euclidean, seuclidean, sqeuclidean), Jensen-Shannon divergence, Kullback-Leibler divergence, Mahalanobis, Minkowski of any norm p, Dice metric, Hamming, Jaccard, Kulczynskil coefficient, Rogers-Tanimoto metric, Russell-Rao metric (russellrao, Russell-Rao), Sokal-Michener metric (sokalmichener, Sokal-Michener), Sokal-Sniff metric (sokalsneath, Sokal-Sneath), Yule-Walker metric (yule, Yule-Walker) and its inherited ones (YuleBonett, YuleCor, Yule2phi, Yule2tetra, Yule2tetra) and others.

[0060] In an alternative embodiment of the claimed solution, at step (106) one of the following anomaly detection methods is used to assess the transaction anomaly: One-class SVM (support vector machine), Isolation Forest (isolation forest), Elliptic envelope, k-nearest neighbors, k-nearest neighbor, ABOD (angle-based outlier detection), or LOF (local outlier factor). In this case, the output value of the method is used as the transaction anomaly measure.

[0061] At step (107), a signal is generated regarding a fraudulent transaction if the anomaly score is greater than the threshold. In a particular embodiment of the claimed invention, a signal is generated regarding a fraudulent transaction if the anomaly score is greater than the threshold for at least one selected distance metric.

[0062] In an example implementation, the solution restricts access for a user who performs at least one fraudulent transaction. Examples of user access restrictions may include, but are not limited to, the following: restricting / prohibiting user access to the target system; restricting specific access rights; suspending and / or canceling user transactions; blacklisting the user; restricting / prohibiting client access to online banking; or a combination thereof.

[0063] Training a machine learning model

[0064] To obtain the hidden state matrix of the transaction history, the encoder (coding) block of the machine learning model is used. A generator (decoder) block is used to generate a reference transaction. Each of these blocks can contain an embedding layer, randomly trained or initialized. Further neural networks are used in the blocks: fully connected, recurrent, convolutional, and transformer.In one particular embodiment of the invention, the encoder block is at least one layer of neural networks of the following types: Recurrent neural networks (RNN), Long short-term memory (LSTM), Gated recurrent units (GRU), Neural Turing machines (NMT), Bidirectional RNN, LSTM and GRU (BiRNN, BiLSTM and BiGRU), Deep residual networks (DRN), Echo state networks (ESN), Liquid state machines (LSM), Kohonen self-organizing map (KN, or organizing (feature) map, SOM, SOFM).

[0065] Training occurs as follows: a transaction history for a given client is obtained from some, not necessarily all, N transactions that have occurred. The first N-1 transactions are fed into a coding machine learning model. A hidden state matrix (HSM) is derived from these transactions. This HSM is fed as input to a generative machine learning model (GMLM). The GMLM output is a generated (reference) transaction for the Nth completed (actual) transaction. A loss function is calculated between the generated (reference) transaction for the Nth completed (actual) transaction and the Nth completed (actual) transaction itself. The GMLM and then the CMLM are trained based on the loss function.

[0066] In one particular embodiment of the invention, at least one of the loss functions is used: KLD (Calculates the Kulbeck-Leibler divergence loss between the true value and the predicted value), MAE (Calculates the mean absolute error between labels and predictions), MAPE (Computes the mean absolute percentage error between the true value and the predicted value), MSE (Computes the root mean squared error between labels and predictions), MSLE (Computes the logarithmic mean squared error between the true value and the predicted value), binary crossentropy (Computes the binary crossentropy loss), binary_focal_crossentropy (Computes the binary focal crossentropy loss), categorical_crossentropy (Computes the categorical crossentropy loss), categorical_hinge (Computes the categorical lasso loss between the true value and the predicted value), cosine_similarity (Computes the cosine similarity between labels and predictions), hinge (Computes the hinge loss between the true value and the predicted value), huber (Calculates the Huber loss), kl_divergence (Calculates the Kulbeck-Leibler divergence loss between the given value and the predicted value),kullback-leibler divergence (Computes the Kullback-Leibler divergence loss between the true value and the predicted value), log_cosh (Log of the hyperbolic cosine of the prediction error), logcosh (Log of the hyperbolic cosine of the prediction error), mean absolute error (Computes the mean absolute error between labels and predictions), mean_absolute_percentage_error (Computes the mean absolute percentage error between the true value and the predicted value), mean_squared_error (Computes the mean squared error between labels and predictions), mean_squared_logarithmic_error (Computes the mean squared logarithmic error between the true value and the predicted value), poisson (Computes the Poisson loss between the true value and the predicted value), sparse_categorical_crossentropy (Computes sparse categorical crossentropy loss),squared_hinge (Calculates the squared loss on the lasso between the true value and the predicted value).

[0067] These loss functions are as follows: binary_cross_entropy, binary_cross_entropy_with_logits, poisson_nll_loss, cosine_embedding_loss, cross_entropy, ctc_loss, gaussian_nll_loss, hinge embedding loss, kl_div, ll_loss, mse_loss, margin_ranking_loss, multilabel_margin_loss, multilabel_soft_margin_loss, multi_margin_loss, nll loss, huber_loss, smooth_ll_loss, soft_margin_loss, triplet_margin_loss, triplet_margin_with_distance_loss.

[0068] In one particular embodiment of the invention, at least one of the following learning algorithms is used: Adadelta, Adagrad, Adam, AdamW, Sparse Adam, Adamax, ASGD, LBFGS, NAdam, RAdam, RMSprop, Rprop, FTRL, SGD, FastSGD, SGD-Nesterov, SAGA, SAGA+.

[0069] In one embodiment of the invention, the dataset used for training was augmented with fraudulent and / or legitimate examples. The augmentations were performed by randomly permuting the parameters and characteristics of the transactions.

[0070] In one embodiment of the invention, the training dataset was augmented with fraudulent and / or legitimate examples based on oversampling of the original sample of a certain class. Techniques such as RandomOverSampler, random sample duplication, and others were used.

[0071] In one embodiment of the invention, the dataset used for training was augmented with fraudulent and / or legitimate examples based on undersampling of the original sample of a certain class. Techniques used included: random seed deletion in RandomUnderSampler, NearMiss, EditedNearestNeighbors, RepeatedEditedNearestNeighbors, CondensedNearestNeighbor, OneSidedSelection, NeighborhoodCleaningRule, InstanceHardnessThreshold, and others.

[0072] In one embodiment of the invention, the dataset used for training was augmented with fraudulent and / or legitimate examples based on the generation of new samples from the original sample of a certain class. Techniques used included SMOTE (Synthetic Minority Oversampling Technique), ADASYN (Adaptive Synthetic Sampling), legacy methods such as BorderlineSMOTE, SVMSMOTE, KMeansSMOTE, SMOTEENN, SMOTETomek, and TomekLinks, and methods based on generative neural networks such as GAN, WCGAN (Wasserstein conditional generative adversarial network), and WCGAN-GP (Wasserstein conditional generative adversarial network with gradient penalty), among others.

[0073] Fig. 2 shows an example of the step-by-step operation of the machine learning coding model (200).

[0074] 1) A sequence of elements is fed to the input (201) of the coding machine learning model (CMLM) (200), and a sequence is created based on it embeddings, where each x tThis is a vector representation of the element

[0075] 2) Using a positional encoder, add (202) positional vectors. This is necessary in order to display the information. about the position of an element in the original sequence. The main property of positional coding is that the further two vectors are from each other in the sequence, the greater the distance between them.

[0076] 3) The resulting vector h. t are fed (203) to the input of the multi-headed self-attention unit t ^ ) where the training matrices are: Q for the query, K for the key, V for the value. Then concatenation is needed to return to the original dimension

[0077] 4) Add (204) skip connections - adding from the input vector to the output tAfterwards, layer normalization is performed (204): It has two trainable parameters, for each vector dimension the mean and variance are calculated.

[0078] 5) Next, add (205) a transformation that will be trainable - a fully connected two-layer neural network (a feed forward neural network):

[0079] 6) Repeat (206) step 4) again: add a through link and layer normalization:

[0080] In the machine learning encoding model (200), steps (203–206) are repeated several more times, transforming context one after another. This enriches the model and increases its number of parameters.

[0081] Fig. 3 shows an example of the step-by-step operation of the generating machine learning model (300).

[0082] The output of the ML model (200) is fed to the input of the ML model (300). The main difference in the ML model (300) architecture is the addition of an attention layer to the vector obtained from the last ML model (200) block. The ML model architecture is also multilayered, and each ML model block receives the vector from the last ML model block as its input. The following are the ML model (300) operation stages in order:

[0083] 1) To parallelize the GMMO (300) and avoid recurrence, but still generate elements sequentially, a technique called data masking from the future is used. The idea is to prevent oneself from peeking at elements that have not yet been generated, taking into account the order. When generating element number t, one is only allowed to look at the first t - 1 elements.

[0084] 2) Next comes the multi-dimensional self-attention stage: linear normalization and multi-dimensional self-attention. The peculiarity is that in the attention layer, keys and values ​​are applied not to all vectors, but only to those whose values ​​have already been synthesized. composition.

[0085] 3) In the next step, we do multidimensional attention on the Z encoding, the result of the KMMO(200): ,j

[0086] 4) The transformation is carried out (similar to KMMO (200)) using a linear fully connected network (feed forward neural network):

[0087] 5) Finally, a probabilistic generative model for the elements is obtained. The result (the index of the attribute value with the highest probability) is the learnable parameters of the linear transformation. For each position t in the output transaction sequence, each of which is represented by a set of attributes, a probabilistic model of the occurrence of values ​​for each attribute is constructed; that is, all elements from the output dictionary are assigned a probability value. These values ​​are obtained from the vectors y. t from the previous point, which are obtained from the last block of GMMO (300).

[0088] The final step is performed only after steps 1–4 have been repeated for all decoders. The output is the probability of each attribute value.

[0089] Let's consider an example of implementing the claimed invention using a user's transactions on a Friday evening. Let's assume the user is a young man heading out on a date with a woman.

[0090] In the course of implementing the claimed method, the following sequence of completed user transactions is obtained: "purchase at a flower shop" -> "withdraw money from an ATM" -> "pay the bill at a restaurant" -> "paying the bill at the bar"

[0091] Each transaction is characterized by a set of attributes, such as transaction time, amount, transaction type, Merchant Category Code, location, etc.

[0092] The goal of the proposed solution is to determine whether a user's most recent transaction is fraudulent. To do this, a reference transaction is generated using a GMMO. The generated (reference) transaction matches the client's history and behavior, consistent with their previous transactions. Thus, the user's most recent transaction, as a result of implementing the method, is "Taxi Payment" (since in similar circumstances—Friday evening—the user often called a taxi).

[0093] Next, we compare the generated last transaction "Taxi payment" with the last completed (actual) transaction of the user. In the case under consideration An example of the last completed (actual) transaction of a user is the transaction - "Receiving a loan for 1,000,000 rubles."

[0094] The comparison result (e.g. by comparing the distance between attribute vectors) shows that the anomaly score of the completed transaction exceeds the threshold value, which indicates that the transaction “Obtaining a loan for 1,000,000 rubles” is fraudulent.

[0095] It should be noted that in the prior art, a separate forecast is built for each attribute (with the forecast being built based on historical data). For example, a transaction amount is predicted based on historical transaction amounts of 1500, 2300, 1900, 2700, and 3100. The forecast will be 3200. If the actual attribute value is 50,000, then this is an anomaly. Similarly, the next attribute is predicted, for example, Merchant Category Code: store, supermarket, fast food, pharmacy, supermarket. Then, a purchase at a fast food restaurant is predicted. In other words, we are talking about separate forecast models for each attribute. Next, the anomaly of a specific attribute is determined, and based on the aggregate of anomaly estimates for all attributes, a prediction is made whether the entire transaction is anomalous. Thus, the prior art does not disclose any mechanism for predicting one attribute taking into account another, for example, predicting the transaction amount taking into account the Merchant Category Code.

[0096] In the claimed invention, in contrast to the prior art, the user's last (reference) transaction is generated based on the totality of all attributes and their interrelations, which contributes to increasing the accuracy and completeness of the detection of fraudulent transactions.

[0097] Furthermore, the claimed invention combines all transaction attributes using a CMMO by processing attribute vectors and obtaining a sequence of attribute vectors in the form of a hidden state matrix. The CMMO is then used to derive related attribute vectors for one or more transactions from the hidden state matrix, which are used to generate the user's final transaction. Thus, the user's final transaction will be generated not based on individual, disparate attributes (as is the case with the prior art), but on related attribute vectors (for example, where at least one transaction takes into account the relationship between transaction time, amount, transaction type, Merchant Category Code, and location). This approach improves the accuracy and completeness of fraudulent transaction detection.

[0098] In turn, the anomaly of a transaction is determined by comparing the distance between the attribute vectors of the user's last completed transaction and the last generated user transaction, i.e., the attribute vectors are compared, not the attributes themselves, as is typical for the prior art.

[0099] It should be noted that, unlike the prior art, which used only an encoder and obtained a matrix of hidden states, and then predicted a transaction based on this matrix, in the claimed invention, a matrix of hidden states is used to generate interconnected attribute vectors using a generative machine learning model.

[0100] A distinctive feature of the proposed solution is that the hidden state matrix is ​​constructed simultaneously across all transaction attribute vectors using the KMMO, while the GMMO generates all transaction attribute vectors simultaneously, with the generated attribute vectors being interconnected. This improves the accuracy and completeness of fraudulent transaction detection by enabling the processing of a set of transaction attributes as a single unit and analyzing this combination for anomalies.

[0101] In a specific example of implementing the claimed solution, unlike the prior art, historical data may not be explicitly used to generate a reference transaction. For example, in the prior art, 99 transactions are taken and the attributes of the 100th transaction are predicted. In the claimed solution, for example, a hidden state matrix is ​​obtained from these 99 transactions. Using a GMMO, related attribute vectors are generated from the hidden state matrix, and all transactions—from the 1st to the 100th—can be generated. This means that the nth transaction is not predicted based on the previous n-1 transactions (in the general case, n, where i is always less than n, i.e., prediction is always based on previous data). In the general case, all n transactions are generated from scratch. Consequently, a hidden state matrix is ​​obtained for n-1 transactions, all transactions up to n are generated starting from the very first transaction, and then this generated nth transaction is compared with the actual one.Thus, the proposed solution presents a completely different approach - the anomaly is not compared with all previous transactions, but one specific (reference) last transaction is generated and the last completed (actual) transaction is compared with it.

[0102] Although the presented implementation examples relate to the financial sector, the claimed invention is not limited to application in this area. The claimed invention can be applied in any field of technology. where there is a need to protect sensitive data that may be subject to fraudulent attacks, such as healthcare, social services, etc.

[0103] Fig. 4 shows a general view of a computing device (400), on the basis of which a device for determining fraudulent user transactions can be implemented, providing a realizable method for determining fraudulent user transactions.

[0104] In general, the computing device (400) comprises one or more processors (401) connected by a common information exchange bus, memory means such as RAM (402) and ROM (403), input / output interfaces (404), input / output devices (405), and means for network interaction (406).

[0105] The processor (401) (or several processors, a multi-core processor) can be selected from a range of devices that are widely used at the present time, for example, from Intel™, AMD™, Apple™, Samsung Exynos™, MediaTEK™, Qualcomm Snapdragon™, etc. A graphics processor can also be used as the processor (401), for example, from Nvidia, AMD, Graphcore, etc., the type of which is also suitable for the full or partial implementation of the method, and can also be used for training and applying machine learning models in various information systems.

[0106] RAM (402) is random access memory (RAM) and is designed to store machine-readable instructions executed by the processor (401) to perform the necessary logical data processing operations. RAM (402) typically contains executable instructions from the operating system and corresponding software components (applications, software modules, etc.).

[0107] ROM (403) is one or more permanent data storage devices, such as a hard disk drive (HDD), a solid-state drive (SSD), flash memory (EEPROM, NAND, etc.), optical storage media (CD-R / RW, DVD-R / RW, Blu-Ray Disc, MD), etc.

[0108] To organize the operation of the device components (400) and to organize the operation of external connected devices, various types of I / O interfaces (404) are used. The selection of the appropriate interfaces depends on the specific design of the computing device, which may include, but are not limited to: PCI, AGP, PS / 2, IrDA, FireWire, LPT, COM, SATA, IDE, Lightning, USB (2.0, 3.0, 3.1, micro, mini, type C), TRS / Audio jack (2.5, 3.5, 6.35), HDMI, DVI, VGA, Display Port, RJ45, RS232, etc.

[0109] To ensure user interaction with the computing device (400), various I / O information devices (405) are used, for example, a keyboard, a display (monitor), a touch display, a touchpad, a joystick, a mouse, a light pen, a stylus, a touch panel, a trackball, speakers, a microphone, augmented reality tools, optical sensors, a tablet, light indicators, a projector, a camera, biometric identification tools (a retinal scanner, a fingerprint scanner, a voice recognition module), etc. [IT] The network interaction means (406) ensures the transmission of data by the device (400) via an internal or external computer network, for example, an Intranet, the Internet, a LAN, etc. One or more means (406) may be, but are not limited to: an Ethernet card, a GSM modem, a GPRS modem, an LTE modem, a 5G modem, a satellite communication module, an NEC module, a Bluetooth and / or a BEE module, a Wi-Fi module, etc.

[0111] Additionally, the device (400) may also include satellite navigation tools such as GPS, GLONASS, BeiDou, Galileo.

[0112] The submitted application materials disclose preferred examples of the implementation of the technical solution and should not be interpreted as limiting other, particular examples of its implementation that do not go beyond the scope of the requested legal protection, which are obvious to specialists in the relevant field of technology.

[0113] Sources of information: [1] torch.nn - PyTorch 2.2 documentation. https: / / pytorch.Org / docs / stable / nn.html#pooling-layers

Claims

FORMULA 1. A computer-implemented method for determining fraudulent user transactions, performed by at least one processor, comprising the steps of: • receive a sequence of completed user transactions over a given time period, with each transaction characterized by a set of attributes; • transform each received transaction into an attribute vector based on the corresponding set of transaction attributes; • process the obtained attribute vectors using a coding machine learning model (CMLM), during which a representation of the sequence of attribute vectors is obtained in the form of a matrix of hidden states; • process the matrix of hidden states using a generative machine learning model (GMLM) trained on matrices of hidden states of sequences of transaction attribute vectors, during which the GMLM generates interconnected attribute vectors of one or more transactions from the matrix of hidden states; generate the next user transaction based on the generated attribute vector; compare the last completed user transaction and the generated last user transaction by calculating an anomaly score for the completed transaction; • generate a signal about a fraudulent transaction if the anomaly score value is greater than the threshold value.

2. The method of p.1 is characterized by the fact that the CMMO is trained on sequences of attribute vectors.

3. The method according to paragraph 1, characterized in that the calculation of the assessment of the anomaly of a completed transaction is carried out according to one or several attributes by comparing the distance between the vectors of transaction attributes.

4. The method according to paragraph 1, characterized in that it additionally comprises the step of restricting access of a user who is performing at least one fraudulent transaction.

5. A device for determining fraudulent user transactions, comprising at least one processor, at least one memory associated with the processor and containing machine-readable instructions that, when executed by at least one processor, ensure the implementation of the method according to any one of paragraphs 1-4.

Citation Information

Patent Citations

  • A negative sample adversarial generation method with noisy learning

    CN111428853B

  • Method for obtaining low-dimensional numerical representations of sequences of events

    EA040376B1

  • Identification of anomalous transaction attributes in real-time with adaptive threshold tuning

    US10482464B1

  • Methods and arrangements to detect fraudulent transactions

    US20200065812A1

  • Data augmentation in transaction classification using a neural network

    US20200210808A1

Cited By

  • Payment fraud real-time identification and disposal method and system based on machine learning

    CN120996814A