Sampling framework for unbalanced transactions

By encoding and sampling transaction data on the server computer and generating a smaller training data set, the problem of high collection and processing costs of large data sets is solved, efficient data processing and model training is achieved, and resource requirements and costs are reduced.

CN120239869APending Publication Date: 2025-07-01VISA INTERNATIONAL SERVICE ASSOCIATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380076583.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-01
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The collection, cleaning and labeling of large training datasets is expensive, storage needs are high, transmission is time-consuming and expensive, and the demand for powerful hardware leads to high costs, limiting rapid development and experimentation of models by organizations with limited resources or researchers.

Method used

Provides an efficient sampling technology/framework that encodes transaction data through server computers, performs difficult negative sampling and diversity sampling processes, generates smaller input training data sets, and evaluates the performance of the machine learning model through iteratively until predefined conditions are met.

Benefits of technology

Reduces the size of the training dataset, reduces the need for storage and computing resources, improves the efficiency of data collection and model training, reduces costs, and supports rapid development and experiments by organizations or researchers with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239869A_ABST
    Figure CN120239869A_ABST
Patent Text Reader

Abstract

The present invention relates to a server computer implementing efficient sampling techniques for unbalanced transactions. Transaction data including untagged transaction data and fraudulent transaction data is encoded to form (i) a first encoded data set associated with fraudulent transactions and untagged transactions similar to fraudulent transactions, and (ii) a second encoded data set associated with untagged transactions. A first sampling process is performed for the first encoded data set to obtain a first sampled encoded data set, and a second sampling process is performed for the second encoded data set to obtain a second sampled encoded data set. An optimal sampling size of the transaction data is determined based on whether performance of a machine learning model that classifies the first sampled encoded data set and the second sampled encoded data set satisfies a condition.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] With the recent explosion of data, organizations have been continuously making considerable efforts to process and manage data. End-user applications, increased network bandwidth, and technological advancements in communication devices (e.g., mobile devices) are some of the factors contributing to the data explosion. Training data is one such type of data and is a key component in building machine learning models. Training data corresponds to a dataset used to teach a model to make predictions or classifications based on input data. The quality and quantity of training data have a significant impact on the performance of machine learning models.

[0002] In various fields, especially in the fields of machine learning and artificial intelligence, the size of training datasets has been a subject of continuous discussion and change. The trend in the size of training datasets can vary depending on the specific application. However, the use of increasingly large datasets to train machine learning models has become a general trend. This trend is driven in part by the availability of more data, advancements in data collection methods, and the view that larger datasets can produce more accurate and robust models.

[0003] While large training datasets are generally beneficial for improving the performance and generalization of machine learning models, they have several drawbacks and challenges. For example, collecting, cleaning, and labeling large datasets can be expensive and time-consuming. Large datasets require a large amount of storage capacity, which can result in high storage costs. Transmitting such datasets to a cloud platform or across a network can also be time-consuming and expensive. In addition, training on large datasets requires powerful hardware, which can be costly. For smaller organizations or researchers with limited resources, the need for more powerful processors (e.g., graphics processing units) can be a limiting factor. Additionally, training on large datasets can be time-consuming, taking days or even weeks to complete. Such factors can hinder the rapid development and experimentation of models. Embodiments of the present invention address these and other problems individually and jointly. Summary of the Invention

[0004] Embodiments provide an efficient sampling technique / framework for unbalanced transactions.

[0005] One embodiment includes a method that includes: encoding, by a server computer, transaction data that includes unlabeled transaction data and fraudulent transaction data, the encoding forming encoded transaction data that includes: (i) a first encoded data set associated with fraudulent transactions and unlabeled transactions similar to fraudulent transactions, and (ii) a second encoded data set associated with unlabeled transactions; performing, by the server computer, a first sampling process on the first encoded data set to obtain a first sampled encoded data set; performing, by the server computer, a second sampling process on the second encoded data set to obtain a second sampled encoded data set; evaluating, by the server computer, the performance of a machine learning model that classifies the first sampled encoded data set and the second sampled encoded data set; and in response to the performance of the machine learning model not meeting a condition, repeating, by the server computer, the first sampling process and the second sampling process by increasing the sizes of the first sampled encoded data set and the second sampled encoded data set until the condition is met.

[0006] Another embodiment includes a server computer that includes: a processor; and a non-transitory computer-readable medium coupled to the processor and including code that can be executed by the processor to implement a method that includes: encoding transaction data that includes unlabeled transaction data and fraudulent transaction data, the encoding forming encoded transaction data that includes: (i) a first encoded data set associated with fraudulent transactions and unlabeled transactions similar to fraudulent transactions, and (ii) a second encoded data set associated with unlabeled transactions; performing a first sampling process on the first encoded data set to obtain a first sampled encoded data set; performing a second sampling process on the second encoded data set to obtain a second sampled encoded data set; evaluating the performance of a machine learning model that classifies the first sampled encoded data set and the second sampled encoded data set; and in response to the performance of the machine learning model not meeting a condition, repeating, by the server computer, the first sampling process and the second sampling process by increasing the sizes of the first sampled encoded data set and the second sampled encoded data set until the condition is met.

[0007] More detailed information about embodiments of the present invention can be found in the detailed description and the drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 Shows a high - level block diagram of an interaction system according to some embodiments.

[0009] Figure 2 Shows a block diagram of a server computer according to an embodiment.

[0010] Figure 3 Depicts a sampling framework implemented by a server computer according to an embodiment.

[0011] Figure 4A and 4B Depicts different types of transactions according to some embodiments.

[0012] Figure 5A Depicts a schematic diagram showing the process of training an encoder according to an embodiment.

[0013] Figure 5B Schematically shows the operations performed on certain types of transactions according to an embodiment.

[0014] Figure 6 Depicts an exemplary flowchart showing the steps of a process (e.g., a first sampling process) performed by a server computer according to an embodiment.

[0015] Figure 7A 、 7B and 7C schematically show certain steps of a first sampling process performed by a server computer according to an embodiment.

[0016] Figure 8 Depicts an exemplary flowchart showing the steps of another process (e.g., a second sampling process) performed by a server computer according to an embodiment.

[0017] Figure 9 Schematically shows the steps of a second sampling process performed by a server computer according to an embodiment.

[0018] Figure 10A Depicts an exemplary flowchart showing the steps of an evaluation process performed to determine the optimal sample size (of input data) of a machine - learning model.

[0019] Figure 10B Depicts an exemplary graph showing the performance of a machine - learning model versus different sample sizes. DETAILED DESCRIPTION

[0020] Before discussing embodiments of the present invention, some terms may be described in further detail.

[0021] "Transaction data" can refer to information or records generated as a result of transactions (e.g., financial or non-financial transactions) between parties or entities. Transaction data typically includes details about the exchange of goods, services, or financial assets. For example, transaction data can include any suitable data corresponding to the transaction, such as account information of a payment account (e.g., PAN, payment token, expiration date, card verification value (e.g., CVV, CVV2), dynamic card verification value (dCVV, dCVV2), identifier of the issuer of the holding account), etc.

[0022] "Credential" can be any suitable information that serves as reliable evidence of value, ownership, identity, or permission. A credential can be a string of numbers, letters, or any other suitable characters that can be presented or contained within any object or document that can serve as confirmation.

[0023] "Token" can be a substitute value for a credential. A token can be a string of numbers, letters, or any other suitable characters. A token can be generated as a hash value by inputting the credential into a cryptographic hash function. Examples of tokens include access tokens, such as payment tokens, data that can be used to access a secure system or location, etc.

[0024] "Payment credential" can include any suitable information associated with an account (e.g., a payment account and / or payment device associated with the account). Such information can be directly related to the account or can be derived from information related to the account. Examples of account information can include PAN (primary account number or "account number"), username, expiration date, CVV (card verification value), dCVV (dynamic card verification value), CVV2 (card verification value 2), CVC3 card verification value, etc. CVV2 is generally understood to be a static verification value associated with a payment device. The CVV2 value is typically visible to the user (e.g., the consumer), while the CVV and dCVV values are typically embedded in memory or authorization request messages and are not easily known to the user (although they are known to the issuer and payment processor). A payment credential can be any information that identifies or is associated with a payment account. A payment credential can be provided to make a payment from a payment account. A payment credential can also include a username, expiration date, gift card number or code, and any other suitable information.

[0025] "Server computer" is typically a powerful computer or a cluster of computers. For example, a server computer can be a mainframe, a small computer cluster, or a group of servers acting as a unit. In one example, a server computer can be a database server coupled to a web server.

[0026] "Encoding" corresponds to the process of converting data from one format or representation to another. Such conversions are typically performed for various purposes, such as data compression, data security, or ensuring compatibility with a particular system or application. An "encoder" is a device, algorithm, or component that performs the encoding task, which involves converting information or data from one format or representation to another. In the present disclosure, data corresponding to a transaction can be encoded and represented as a vector in an N-dimensional (e.g., N = 5-dimensional) vector space.

[0027] "Sampling" can refer to the process of selecting a subset of data points or observations from a larger population or dataset for the purpose of analysis, modeling, or making inferences about the entire population. Sampling is a technique used when it may be impractical or infeasible to collect and process data from the entire population. A "sampling method" or "sampling process" is a procedure for selecting individuals or data points from a population to form a sample, and the "sample size" is the number of data points or individuals selected for the sample.

[0028] The "hard negative sampling" process can include a first sampling process performed by a first sampling module. Hard negative sampling can be considered a variant of negative sampling and contrastive learning, and is useful in scenarios where there is an imbalance between the number of positive and negative samples, making it challenging to effectively train a model. The goal of hard negative sampling is to focus the training (e.g., sampling) of the model on the most challenging or informative negative examples, thereby improving the model's ability to distinguish positive instances from negative instances and ultimately leading to better classification performance.

[0029] The "diversity sampling" process can include a second sampling process performed by a second sampling module. Diversity sampling is a data sampling technique that aims to select different subsets of data points from a larger dataset. The goal of diversity sampling is to ensure that the selected subsets contain a variety of representative examples that cover all aspects of the data distribution. Specifically, diversity sampling aims to improve the quality of the selected subsets and prevent overrepresentation of specific patterns or classes.

[0030] A "fraudulent transaction", also known as a fraudulent activity or fraudulent operation, can be an unauthorized or deceptive transaction carried out with the intention of deceiving, obtaining an unfair advantage, or committing an illegal act. Fraudulent transactions can occur in various contexts, such as banking, e-commerce, insurance, credit card transactions, etc. The data corresponding to a fraudulent transaction is referred to herein as "fraudulent transaction data".

[0031] "Training data" can be used to build a machine learning model. Training data corresponds to a dataset used to teach a model to make predictions or classifications based on input data. It should be noted that the quality and quantity of training data can have a significant impact on the performance of a machine learning model.

[0032] "Model training" or "machine learning training" corresponds to the process by which a machine learning algorithm or model learns from data to make predictions or perform a specific task. It involves presenting the model with a dataset that includes input features and corresponding target outputs (also known as labels), and the model learns to make predictions or classifications by adjusting its internal parameters.

[0033] In the context of machine learning and data analysis, "unlabeled data" refers to a dataset in which individual data points or instances do not have an associated label or target value. In other words, the data does not have explicit annotations or categories assigned to each observation. Unlabeled data is different from labeled data, which includes data with predefined target values or classifications.

[0034] "Imbalanced transactions" can refer to a situation where there is a significant difference in the number or value of transactions between different classes or categories within a dataset. Such imbalances can occur in various contexts, including financial transactions, fraud detection, healthcare, etc. It should be noted that addressing imbalanced transactions in machine learning is important as it can lead to biased models.

[0035] "Memory" can be any suitable one or more devices that can store electronic data. Suitable memory can include non-transitory computer-readable media that stores instructions executable by a processor to implement the desired method. Examples of memory can include one or more memory chips, disk drives, etc. Such memory can operate using any suitable electrical, optical, and / or magnetic operating modes.

[0036] "Processor" can refer to any suitable one or more data computing devices. The processor can include one or more microprocessors that work together to achieve the desired function. The processor can include a CPU, which includes at least one high-speed data processor sufficient to execute program components for executing user- and / or system-generated requests. The CPU can be a microprocessor, such as AMD's Athlon, Duron, and / or Opteron; IBM and / or Motorola's PowerPC; IBM and Sony's Cell processor; Intel's Celeron, Itanium, Pentium, Xeon, and / or XScale; and / or similar processors.

[0037] "User" can include an individual. In some embodiments, the user may be associated with one or more personal accounts and / or user devices.

[0038] Figure 1A block diagram of an interaction system 100 according to an embodiment is shown. The interaction system 100 includes one or more user devices, such as client devices 101, 103, and 105, and a server computer 107. In some embodiments, the server computer 107 may host a machine learning model and be configured to train the machine learning model based on training data received from different data sources (e.g., client device 101). Training data is a key component in building a machine learning model. The training data corresponds to a dataset used to teach the model to make predictions or classifications based on input data. The quality and quantity of the training data have a significant impact on the performance of the machine learning model. The trend of the size of the training dataset can vary according to a specific application. However, using increasingly large datasets to train machine learning models has become a general trend.

[0039] Although large training datasets are generally beneficial for improving the performance and generalization of machine learning models, they have several drawbacks and challenges as previously described. According to some embodiments, the server computer 107 is configured to determine the optimal size of the training data set without including the performance of the machine learning model. As will be discussed below, the server computer implements different data sampling processes to obtain input training data that is significantly smaller in size. It should be understood that the server computer 107 ensures that different sample variants in the input data set (i.e., training data) are considered when evaluating the performance of the machine learning model while reducing the size of the input data set.

[0040] The functions of the server computer 107 described herein can be applied to different transactions, such as highly imbalanced credit card transactions. Specifically, in the sense that most transactions (e.g., normal transactions) have a similar pattern, i.e., similar characteristics, while fraudulent transactions have abnormal behavior (i.e., different characteristics), the data associated with credit card transactions (also referred to as transaction data herein) is highly imbalanced. Thus, through some embodiments, the server computer 107 generates different data sets based on the training data received from different client devices, i.e., generates a first normal transaction set and implements a first sampling process (e.g., diversity sampling) to sample the data covering various normal transactions. In doing so, the server computer ensures that sampling of normal transactions does not impair (i.e., negatively impact) the performance of the machine learning model.

[0041] In addition, the server computer 107 generates a second data set corresponding to normal transactions that have behavior (i.e., characteristics) similar to fraudulent transactions. A second sampling process is implemented on the generated second data set. This enables the server computer to distinguish fraudulent patterns from normal transactions in the transaction data. As Figure 1As depicted, the output of the server computer can correspond to an optimal size of training data (e.g., a reduced training data size) that can be used to train a machine learning model without significantly affecting its performance. Next, refer to Figure 2 A- Figure 10B Describe details related to the server computer 107.

[0042] Figure 2 Depicts the server computer 200. In some embodiments, Figure 2 The server computer 200 of Figure 1 Corresponds to the server computer 107 of

[0043] The computer-readable medium 204 can include an encoding module 204A, a communication module 204B, a first sampling module 204C, a second sampling module 204D, a merging module 204E, and an evaluation module 204F. The communication module 204B can include code that enables the processor 202 to generate messages, forward messages, reformat messages, and / or otherwise communicate with other entities. Specifically, the communication module 204B can include various communication means, such as short-range antennas, long-range antennas, etc., to communicate with other devices such as Figure 1 The user or client devices 101, 103, and 105 of

[0044] In some embodiments, the communication module 204B can include one or more RF transceivers and / or connectors that can be used by the server computer 200 to communicate with other devices and / or connect to an external network. The short-range antenna of the communication module 204B can be configured to communicate with external entities via a short-range communication medium (e.g., using Bluetooth, Wi-Fi, infrared, NFC, etc.). The long-range antenna of the communication module 204B can be configured to communicate over the air with a remote base station and a remote cellular or data network. An example of a communication channel formed by the server computer 200 can be a communication channel formed with a user device (e.g., Figure 1 The client devices 101, 103, and 105 depicted in Figures 3 to 10B In such a communication channel, the client device can be programmed to transmit training data for training a specific machine learning model to the server computer 200. In response, the server computer 200 can perform the processing described below with reference to Figure 1of the client device 101).

[0045] The encoding module 204A is a component or layer of a model responsible for encoding or extracting basic information from input data. Specifically, the encoding module 204A corresponds to an encoder that converts raw input data (e.g., transaction data) into a more compact and meaningful representation (e.g., an N-dimensional vector representation in a vector space) suitable for subsequent processing or analysis. Such an encoded representation may have dimensionality reduction and focus on capturing the significant features of the input data, which can be used for tasks such as classification, clustering, or generating reconstructions of the original data.

[0046] The first sampling module 204C and the second sampling module 204D correspond to respective components of the server computer 200 configured to extract or select a subset of items or data points from a larger set of data points. It should be understood that the first sampling module 204C and the second sampling module 204D may correspond to a device, component, or process (usually involving elements of randomness or selection criteria) that performs the task of selecting or extracting items or data points from a larger set of data points. The selection criteria correspond to the process implemented by the sampling module when generating a subset of data points from a larger set of data points. For example, in some embodiments, the first sampling module 204C may implement a first sampling process (e.g., hard negative sampling), while the second sampling module 204D may implement a second sampling process different from the first sampling process (e.g., diversity sampling). Refer to Figure 6 and Figures 7A - 7C details the details related to the first sampling process performed by the first sampling module 204C, while referring to Figure 8 and 9 details the details related to the second sampling process performed by the second sampling module 204D. It should be understood that although the server computer 200 is depicted as including two separate sampling modules (each performing a unique sampling process), this should not limit the scope of the present disclosure. For example, the server computer 200 may include a single sampling module configured to perform different sampling processes for different input data sets.

[0047] The merging module 204E of the server computer 200 is a component configured to generate a merged data set (e.g., a concatenated data set) from the outputs generated by the first sampling module 204C and the second sampling module 204D. Specifically, in some embodiments, the first sampling module 204C generates a first sampled data set (i.e., a first set including a first plurality of data points), and the second sampling module 204D generates a second sampled data set (i.e., a second set including a second plurality of data points). The merging module 204E is configured to generate a concatenated set including the first sampled data set and all the elements (data points) included in the second sampled data, i.e., the merged set includes the first plurality of data points and the second plurality of data points.

[0048] The evaluation module 204F of the server computer 200 is programmed to evaluate (i.e., assess) the performance of a machine learning model to determine the extent to which the model can make predictions or classifications on new unseen data. The evaluation process helps to understand the strengths, weaknesses, and suitability of the model for a particular task. In some embodiments, common metrics and techniques can be used to evaluate the performance of the model, and the choice of the evaluation method depends on the type of machine learning task, such as classification, regression, or clustering.

[0049] In some embodiments, the evaluation module 204E is configured to evaluate the performance of the machine learning model in an iterative manner. Specifically, to obtain the optimal sample size of the input data, the evaluation module 204F obtains the performance of the machine learning model (e.g., a first performance) for a first sample size of the input data. The evaluation module 204F further determines whether the first performance meets some predefined conditions. In response to determining that the first performance does not meet the predefined conditions, the evaluation module 204F is configured to repeat another evaluation process (i.e., another iteration of evaluating the machine learning model) with an increased sample size of the input data. Details regarding the iterative evaluation of the machine learning model are described later with reference to Figure 10A and 10B The details of the iterative evaluation of the machine learning model are described. It should be noted that the server computer 200 may be associated with the database 210. The server computer can utilize the database 210 to store data obtained from the encoding module (i.e., encoded data), from the first / second sampling module (i.e., sampled data sets), from the merging module (i.e., merged data sets), and / or data obtained from the evaluation module (i.e., the performance of the machine learning model).

[0050] Figure 3 Depicts a sampling framework implemented by a server computer (e.g., Figure 2 the server computer 200) according to one embodiment. As Figure 3As shown in, the sampling framework 300 includes an encoder 303, a first sampling module 309, a second sampling module 311, a merging module 317, and an evaluation module or evaluator 319. It should be noted that the components / modules listed above are included in Figure 2 the server computer 200.

[0051] The input to the encoder 303 is the transaction data 301. The transaction data 301 corresponds to fraudulent transactions and unlabeled transactions. It should be noted that such transaction data 301 can be obtained from different data sources, as previously referenced Figure 1 as described. Specifically, the transaction data 301 includes a plurality of data points (also referred to herein as transaction data points), the plurality of data points including – one or more fraudulent transaction data points corresponding to known fraudulent transactions (e.g., fraudulent transaction data point 301A) and one or more unlabeled transaction data points corresponding to some transactions (e.g., unlabeled transaction data point 301B). The encoder 303 is configured to convert information or data from one format or representation to another. For example, the encoder 303 can receive the transaction data 301 in one format (e.g., the plaintext vector / matrix format as shown in Figure 4A and 4B ), and encode each transaction data point included in the transaction data into a vector (e.g., an N-dimensional vector) represented in an N-dimensional vector space.

[0052] In some embodiments, the encoder 303 encodes the transaction data 301 based on a contrastive learning method. Specifically, the goal of contrastive learning is to learn a meaningful data representation by training the encoder to distinguish pairs of data points or separate positive examples (pairs that should be similar) from negative examples (pairs that should not be similar). It should be noted that the contrastive learning method does not require labeled data for supervised learning. Instead, it exploits the inherent structure and similarity in the data to learn useful representations. As shown in Figure 3 , the encoder 303 generates two encoded data sets, namely a first encoded data set 305 and a second encoded data set 307. The first encoded data set 305 is associated with fraudulent transactions and unlabeled transactions that are 'similar' to the fraudulent transactions. In other words, the first encoded data set includes encoded fraudulent transaction data points (represented as bold circles) and encoded unlabeled transaction data points (represented as non-bold circles). The second encoded data set 307 is associated with unlabeled transactions, i.e., the second encoded data set includes one or more encoded unlabeled transaction data points. Details regarding the operations performed by the encoder 303 and the concept of unlabeled transaction data points that are 'similar' to fraudulent transaction data points (corresponding to fraudulent transactions) will be referred to later with reference toFigure 5A and 5B are described.

[0053] The first encoded data set 305 is input to the first sampling module 309, while the second encoded data set is input to the second sampling module 311. According to some embodiments, the first sampling module 309 performs a first sampling process (e.g., hard negative sampling process) to generate a first sampled encoded data set 313. The second sampling module 311 performs a second sampling process (e.g., diversity sampling process) to generate a second sampled encoded data set 315. As Figure 3 shown, for illustrative purposes, the transaction data 301 is depicted as including a total of 17 data points, i.e., 3 fraudulent transaction data points (represented by solid circles) and 14 unlabeled transaction data points (represented by hollow circles). The encoder 303 generates two encoded data sets, namely, a first encoded data set 305 and a second encoded data set 307, where the first encoded data set includes 3 fraudulent transaction data points and 4 unlabeled transaction data points (which are 'similar' to fraudulent transactions), and the second encoded data set includes the other 10 unlabeled transaction data points.

[0054] The first sampling module 309 generates the first sampled encoded data set 313 by sampling the first encoded data set 305. Similarly, the second sampling module 311 generates the second sampled encoded data set 315 by sampling the second encoded data set 307. In some embodiments, the first sampling module 309 is configured to sample / select all fraudulent transaction data points (i.e., the points represented by solid circles in the encoded data set 305), and sample one or more unlabeled transaction data points (which are similar to fraudulent transactions) from the encoded set 305. The one or more sampled unlabeled transaction data points similar to fraudulent transactions are represented by shaded circles (e.g., shaded data point 301C). In Figure 3 the example depicted, the first sampling module 309 samples 3 out of 4 unlabeled transaction data points (in total) to include in the first sampled encoded data set 313. In a similar manner, the second sampling module 311 is depicted as selecting / sampling 5 out of 10 unlabeled transaction data points (in total) from the second encoded data set 307 to form the second sampled encoded data set 315.

[0055] In some embodiments, the generated first sampled and encoded data set 313 and the second sampled and encoded data set 315 are input to a merging module 317. The merging module 317 generates a concatenated (or combined) data set that includes the samples or data points included in the first sampled and encoded data set 313 and the second sampled and encoded data set 315.

[0056] The combined data set generated by the merging module 317 can be input to an evaluator 319. In some embodiments, the evaluator 319 is configured to evaluate the performance of a machine learning model that classifies the data points included in the concatenated data set (i.e., the first sampled and encoded data set and the second sampled and encoded data set). The evaluator 319 determines whether the performance of the machine learning model meets a condition. Based on the condition not being met, the evaluator 319 triggers the first sampling module 309 and the second sampling module 311 to increase the sample sizes of the first sampled and encoded data set 313 and the second sampled and encoded data set 315, respectively, in order to perform another round (i.e., iteration) of evaluating the performance of the model. In this way, the evaluator 319 repeatedly performs model evaluation (with the increased sample sizes of the first sampled and encoded data set and the second sampled and encoded data set) until the condition is met. After the condition is met, the evaluator 319 ends the evaluation of the machine learning model and outputs the sample sizes of the first sampled data set and the second sampled data set relative to the transaction data 301 as the optimal sample sizes (e.g., the overall sample size).

[0057] By way of some examples, the above-indicated condition corresponds to the point at which the performance of the machine learning model begins to level off with respect to the sample size (e.g., the overall sample size of the transaction data) (e.g., on a performance chart of the machine learning model). For example, referring to Figure 10B , an exemplary chart 1050 is depicted that shows the performance of a machine learning model with respect to different sample sizes. The chart shows the performance of the machine learning model (plotted on the Y-axis) with respect to different sample sizes (plotted on the X-axis).

[0058] The performance of the machine learning model is evaluated for different sample sizes of the transaction data. For example, as Figure 10BAs shown in Chart 1050, for the first sample size denoted as S1, the point marked as 1051 (marked as solid line 'X') corresponds to the performance of the machine learning model (denoted as P1). Similarly, for the sample sizes denoted as S2 - S7 respectively, the points marked as 1052, 1053, 1054, 1055, 1056, and 1057 correspond to the performance of the machine learning model (denoted as P2 to P7) respectively. It should be noted that by gradually increasing the sample size of the transaction data, the performance of the machine learning model is obtained iteratively at each sample size.

[0059] For the sample sizes of S1 to S4, the points 1051, 1052, 1053, and 1054 corresponding to the performance of the machine learning model (P1 to P4) are marked by the solid line 'X', while for the sample sizes of S5 to S7, the points 1055, 1056, and 1057 corresponding to the performance (P5 to P7) are marked by the dashed line 'X'. It can be seen from Chart 1050 that there is a significant improvement (e.g., above the predetermined performance threshold level) in the process from sample size S1 to S2, then from S2 to S3, further from S3 to S4, and from S4 to S5. However, in the process from sample size S5 to S6 and further from S6 to S7, the performance improvement of the machine learning model is negligible (e.g., below the predetermined performance threshold level).

[0060] Therefore, according to some embodiments, the point marked as 1055 corresponds to the point at which the performance of the machine learning model with respect to the sample size begins to level off (i.e., the performance of the machine learning model begins to reach a steady state). In other words, in the process of the sample size from S5 to S6 (corresponding to the points marked as 1055 and 1056 respectively), the improvement in the performance of the machine learning model is below the predetermined performance threshold level. Therefore, by increasing the sample size beyond the point marked as 1055, the improvement (or enhancement) in the performance of the machine learning model is negligible. This can be further confirmed by comparing the performance of the machine learning model at point 1055 with the performance of the model at point 1060 (denoted as a star and corresponding to the complete set of transaction data), where the performance improvement is negligible. Therefore, according to some embodiments (and referring to Figure 3 )), the evaluator 319 of the server computer can iteratively increase the sample size from S1 to S5 and determine that the sample size of S5 corresponds to the optimal size of the set of transaction data that produces an acceptable performance level of the machine learning model. It should be understood that in some embodiments, in order to determine that the sample size of S5 is optimal, for the purpose of verifying that the performance of the machine learning model has reached a steady state (e.g., for confirmation purposes), the evaluator can increase the sample size to S6 or S7.

[0061] Go to Figure 4Aand 4B , depicts different types of transactions according to some embodiments. Specifically, the different types of transactions that can be included in the transaction data (e.g., Figure 3 transaction data 301) can be transactions that occur at the user level or transactions that occur at the user's account level. Transactions that occur at the user level correspond to real-time transactions that occur between two entities 401 and 403. For example, Figure 4A depicts a real-time transaction that occurs between two entities (i.e., user 401 and user 403). Such a transaction can be characterized by a one-dimensional array 405 that includes entries corresponding to the sender, recipient, and transaction amount. In addition, the one-dimensional array can include one or more features associated with the transaction. The one or more features (e.g., feature 1, feature 2, and feature 3) can correspond to parameters associated with the transaction between the two entities (e.g., the transaction speed between the entities, the amount turnover speed between the entities, etc.).

[0062] According to some embodiments, the transactions included in the transaction data can correspond to account-level transactions as depicted in Figure 4B . Specifically, Figure 4B depicts transactions associated with the account of user 401 (e.g., incoming and outgoing transactions). Such transactions are depicted in Figure 4B as a two-dimensional matrix 415. It should be understood that each row of the two-dimensional matrix 415 is similar to the one-dimensional array 405 as depicted in Figure 4A . In other words, each row of the two-dimensional matrix 415 corresponds to a single transaction associated with the account of user 401. Therefore, each row can be characterized by the sender, recipient, transaction amount, and one or more features associated with the transaction.

[0063] Figure 5A depicts a schematic diagram showing the process of an encoder included in a training server computer according to an embodiment. In some embodiments, the encoder included in the training server computer is trained based on a contrastive learning method. The goal of contrastive learning is to learn a meaningful data representation by training the encoder to distinguish pairs of data points or separate positive examples (pairs that should be similar) from negative examples (pairs that should not be similar).

[0064] Figure 5ADepicts a framework for training the encoder 303. The framework includes a sample processing unit 512, fraudulent data samples 510, and unlabeled data 501. It should be noted that transaction data may often be extremely imbalanced, i.e., fraudulent transactions occur in a small number (and may have different behaviors, i.e., characteristics), while normal transactions (i.e., non-fraudulent transactions) occupy most of the data samples and share similar behaviors or characteristics. It should be understood that fraudulent transactions are important in the training of the encoder because ignoring them may cause the encoder to be biased. As described below, the encoder 303 of the present disclosure is trained by using sampling of normal (i.e., non-fraudulent) transactions and considering all fraudulent transactions. Therefore, it should be understood that the trained encoder should be able to separate fraudulent transactions from non-fraudulent transactions.

[0065] According to some embodiments, the encoder is trained by providing a first input of fraudulent data samples 510 to the encoder 303. For illustrative purposes, the fraudulent data samples are depicted as including an amount entry (i.e., transaction amount) corresponding to the characteristics of the fraudulent data samples and three feature entries (i.e., feature 1, feature 2, and feature 3). In one implementation, the fraudulent data samples 510 are also input to the sample processing unit 512. The sample processing unit 512 performs certain operations on the fraudulent data samples 510 to generate one or more additional samples of the fraudulent data samples 510, i.e., variants of the fraudulent data samples 510. For example, the sample processing unit 512 may perform a first operation corresponding to a random swap operation on the fraudulent data samples 510. As Figure 5A shown, a different sample of the fraudulent data samples 510 can be obtained by swapping the entry of feature 1 with feature 2 of the fraudulent data samples 510.

[0066] The sample processing unit 512 may perform a second operation corresponding to a masking operation, where specific entries of the fraudulent data samples 510 may be masked. As Figure 5A shown, another different sample of the fraudulent data samples 510 can be obtained by masking the entry associated with feature 2. Additionally, the sample processing unit 512 may perform a third operation corresponding to adding Gaussian noise to the fraudulent data samples 510. Such operations correspond to adding random numbers (i.e., noise) to one or more entries of the fraudulent data samples 510. As Figure 5A shown, a Gaussian noise variant of the fraudulent data samples 510 is obtained by adding noise (i.e., small variants) to the entries of features 1 to 3 of the fraudulent data samples 510. The variants of the fraudulent data samples 510 generated by the sample processing unit 512 are Figure 5A collectively denoted as 515 in. By some embodiments, each variant of the fraudulent data samples 510 generated by the sample processing unit 512 is input to the encoder 303.

[0067] In addition, through an embodiment, unlabeled data 501 is randomly sampled and input into the encoder 303. The encoder 303 is configured to map each input in the vector space. For example, the encoder encodes each input into an N-dimensional vector in the vector space (e.g., N = 5 dimensions). Through an embodiment, the output of the encoder 303 is a first encoded data set 520 and a second encoded data set 530. The first encoded data set 520 includes all fraudulent data samples (e.g., fraudulent data samples / points labeled as 520A) and one or more unlabeled data samples / points, e.g., unlabeled data point 520B. One or more unlabeled data points included in the first encoded data set 520 are referred to herein as unlabeled data points similar to fraudulent transactions (corresponding to unlabeled transactions). Based on the unlabeled data point (corresponding to an unlabeled transaction) being within a predetermined distance of the encoded plurality of fraudulent data points (corresponding to fraudulent transactions) in the vector space, it is determined that the unlabeled data point (e.g., sample 520B) is similar to the fraudulent data point. In other words, the unlabeled data point, e.g., sample 520B behaves similarly to the fraudulent data samples included in the first encoded data set 520 (i.e., has the same characteristics). The second encoded data set 530 includes other unlabeled data samples that are not similar to the fraudulent data sample portion. In this way, the encoder generates two encoded data sets 520 and 530 respectively, such that the distance between the two combinations (e.g., the distance measured from the center (average value) of the two sets) is maximized.

[0068] As described below, the first encoded data set 520 is input into the first sampling module of the server computer. The first sampling module performs or implements a first sampling process on the first encoded data set 520. On the other hand, the second encoded data set is input into the second sampling module of the server computer. The second sampling module performs or implements a second sampling process (different from the first sampling process) on the second encoded data set 530. In addition, it should be understood that the operations performed by the sample processing unit 512 described above (e.g., random swapping, masking, Gaussian noise, etc.) are intended to be illustrative and do not limit the scope of the present disclosure in any way. Instead, the sample processing unit can perform any number of different operations to generate variants of the fraudulent data sample 510. Similarly, although the fraudulent data sample 510 used for training the encoder is depicted as a one-dimensional array, it should be understood that a two-dimensional fraudulent data sample 550 (as Figure 5BAs shown in [reference], it can be equally applicable to training the encoder. In such scenarios, the sample processing unit 512 can perform similar operations (e.g., random masking, addition of Gaussian noise, etc.) on the fraudulent data sample 550 to generate one or more of its variants, such as the samples denoted as 560 and 570 in [reference]. Figure 5B The samples.

[0069] Figure 6 FIG. [figure number] depicts an exemplary flowchart showing the steps of a process (e.g., the first sampling process) performed by a server computer according to an embodiment. Specifically, Figure 6 the flowchart of [figure number] corresponds to the steps performed by the first sampling module 309 included in the server computer. The first sampling module performs a first sampling process (referred to herein as the 'hard negative' sampling process) on the first encoded data set output by the encoder. Figure 6 The processes depicted in [figure number] can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding system, hardware, or a combination thereof. The software can be stored on a non-transitory storage medium (e.g., on a memory device). In Figure 6 The methods presented in [figure number] and described hereinafter are intended to be illustrative and non-limiting. Although Figure 6 FIG. [figure number] depicts various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In certain alternative embodiments, the steps can be performed in some different order, or some steps can also be performed in parallel.

[0070] It should be noted that the first encoded data set includes data points corresponding to fraudulent transactions (e.g., fraudulent transaction data points) and data points determined to be similar to fraudulent transactions (e.g., unlabeled transaction data points). The goal of the first sampling process is to select (i.e., sample) one or more data points from the set of unlabeled transaction data points that are similar to fraudulent transaction data points (i.e., known hard negative data points). In some embodiments, when the first sampling module performs the first sampling process on the first encoded data set, it generates a first sampled encoded data set that includes one or more unlabeled transaction data points (regarded as similar to fraudulent transaction data points) and all fraudulent transaction data points. As Figure 6 The process depicted in [figure number] includes two phases: (a) a pre-filtering phase corresponding to steps 601 to 607 in Figure 6 [figure number], and (b) a selection phase corresponding to steps 609 to 619 in Figure 6 [figure number]. Additionally, for clarity, Figure 7A 7B7C schematically shows certain steps of a first sampling process performed by a first sampling module according to some embodiments.

[0071] The pre-filtering phase begins at step 601, where the sample size is defined. The sample size corresponds to Figure 6 the number of samples extracted from the first encoded data set for each iteration of the process. For example, the sample size K 样本大小 = 2 can be defined in this step. The process in step 603 calculates the centroid for the fraudulent transaction data points included in the first encoded data set (i.e., the input set). It should be noted that the centroid is the central point or arithmetic mean position of a set of data points in a multi-dimensional space (i.e., a vector space). As Figure 7A shown, for the fraudulent transaction data points 701A, 701B, and 701C, the centroid 703 is calculated.

[0072] The process then moves to step 605, where for each unlabeled transaction data point included in the input set, its distance to the centroid (calculated in step 603) is calculated. For example, as Figure 7B shown, for the unlabeled transaction data points 705A, 705B, 705C, 705D, and 705E, their respective distances to the centroid 703 (denoted as 'a', 'b', 'c', 'd', and 'e') are calculated. The process in step 607 filters the first predetermined number of unlabeled transaction data points based on their calculated distances in step 605. For example, the first predetermined number (e.g., N = 50 * K 样本大小 ) of the closest points to the centroid can be filtered out. As Figure 7B shown, the unlabeled transaction data points 705A, 705B, 705C are filtered out, and the other unlabeled transaction data points 705D and 705E are excluded / discarded because these transaction data points correspond to points that are far from the centroid (e.g., beyond a threshold distance). In this example, the unlabeled transaction data points 705A, 705B, 705C (circled by the dashed ellipse 706) correspond to the first predetermined number of unlabeled transaction data points.

[0073] Then, the process proceeds to the selection phase starting at step 609. In step 609, for the unlabeled transaction data points included in the first predetermined number of unlabeled transaction data points (from step 607), the distance of the unlabeled transaction data point to each fraudulent transaction data point is measured. For example, referring to Figure 7C it can be seen that the unlabeled transaction data points 705A, 705B, and 705C are included in the first predetermined number of unlabeled transaction data points (determined in step 607). Taking a particular unlabeled transaction data point as an example, for instance Figure 7CThe point 705A therein depicts a scenario where the calculated distances from the point 705A to each fraudulent transaction data point (i.e., points 701A, 701B, and 701C) are considered. In Figure 7C it, the measured distances from the point 705A to the fraudulent transaction data points are respectively labeled as 'x', 'y', and 'z'.

[0074] In step 611, for an unlabeled transaction data point (e.g., 705A), based on the distances calculated in step 609, the minimum distance is determined. For example, considering the unlabeled transaction data point 705A, the minimum distance associated with this point is calculated as: minimum(x, y, z). Additionally, in step 613, for the unlabeled transaction data point, its distance to the fraudulent transaction data point is set to the minimum distance associated with this point (calculated in step 611). Such a distance is referred to herein as the separation distance between the unlabeled transaction data point and the fraudulent transaction data point. In step 615, for other unlabeled transaction data points included in the first predetermined number of unlabeled transaction data points (e.g., points 705B and 705C), steps 609, 611, and 613 are repeated.

[0075] Then the process moves to step 617, where a ranking is assigned to each unlabeled transaction data point included in the first predetermined number of unlabeled transaction data points. It should be understood that the ranking can be assigned based on the ascending order (i.e., an increasing sequence) of the minimum distances of the corresponding unlabeled transaction data points to the fraudulent transaction data points. The rankings of the unlabeled transaction data points form a set of sorted unlabeled transaction data points. In step 619, the first N number of unlabeled transaction data points are selected from the set of sorted unlabeled transaction data points, where the parameter N = K 样本大小 , i.e., the sample size defined in step 601. In this way, according to one embodiment, the first sampling module of the server computer samples the N nearest unlabeled transaction data points.

[0076] Figure 8 depicts an exemplary flowchart showing the steps of another process (e.g., a second sampling process) performed by a server computer according to an embodiment. Specifically, Figure 8 the flowchart corresponds to the steps performed by the second sampling module 311 included in the server computer. The second sampling module performs a second sampling process (referred to herein as a 'diversity' sampling process) on the second encoded data set output by the encoder. Figure 8 The processing depicted in it can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of the corresponding system, hardware, or a combination thereof. The software can be stored on a non - transitory storage medium (e.g., on a memory device). InFigure 8 The methods presented and described hereinafter are illustrative and not restrictive. Although Figure 8 various processing steps are depicted as occurring in a particular sequence or order, this is not intended to be restrictive. In certain alternative embodiments, the steps may be performed in some different order, or some steps may be performed in parallel.

[0077] It should be noted that the second encoded data set includes unlabeled data points corresponding to non-fraudulent transactions. The goal of the second sampling process is to cover various samples from different clusters, where each cluster is characterized by a unique behavior (such as high transaction amount, high speed, etc.). In other words, the unlabeled data points belonging to a particular cluster share the same characteristics (i.e., behavior). In some embodiments, when the second sampling module performs the second sampling process on the second encoded data set, a second sampled and encoded data set is generated, and the second sampled and encoded data set includes one or more unlabeled data points. Additionally, for clarity, Figure 9 certain steps of the second sampling process performed by the second sampling module according to some embodiments are schematically shown.

[0078] The process begins at step 801, where the second sampling module generates clusters of encoded unlabeled data samples (i.e., unlabeled transaction data points). It should be noted that different clusters can be generated based on the characteristics associated with the unlabeled transaction data points. For example, as Figure 9 shown, three clusters 901, 903, and 905 can be generated, and each of the clusters is associated with a specific characteristic, that is, the unlabeled transaction data points included in a specific cluster share the same characteristics (such as high speed, high transaction amount, etc.). In step 803, for each cluster generated in step 801, the average value of the cluster is calculated. Additionally, in step 805, the calculated average value (from step 803) is designated as the seed of the cluster (i.e., the seed data point of the cluster). For example, referring to Figure 9 , with respect to cluster 901, the average value of cluster 901S is designated as the seed of the cluster. Similarly, for clusters 903 and 905, their respective average values (i.e., 903S and 905S) are designated as the respective seeds of the clusters.

[0079] According to some embodiments, the seed of each cluster is used as the starting point for performing sampling for the cluster. Specifically, for each unlabeled transaction data point in the cluster, a probability is calculated based on the distance of the unlabeled transaction data point from the seed of the cluster. The probability of the unlabeled transaction data point can be calculated as follows: where D(x i, s) is an unlabeled transaction data point x i The distance from the seed (s) of the cluster, and the variable j iterates over all unlabeled transaction data points included in the cluster. In other words, the probability of the unlabeled transaction data point P(x i ) of x i is calculated as the ratio of the distance of the unlabeled transaction data point to the seed of the cluster to the sum of the corresponding distances of the other unlabeled transaction data points in the cluster to the seed of the cluster. Thus, according to the above equation, for a particular cluster (e.g., Figure 9 cluster 901), compared to the second unlabeled transaction data point (e.g., 901B), the first unlabeled transaction data point (e.g., 901A) located farther from the seed (901S) of the cluster has a higher probability of being sampled from the cluster.

[0080] The process in step 809 continues to sample one or more unlabeled transaction data points from each cluster based on the probabilities calculated in step 807. In some embodiments, a predetermined sample size (B 样本大小 ) can be defined for each cluster, e.g., B 样本大小 = 1. Thus, in step 809, B 样本大小 number of unlabeled transaction data points can be selected. Additionally, the process moves to step 811, where for each cluster, the sampled unlabeled transaction data points of the cluster are designated to operate as the seeds of the cluster in the next iteration of the second sampling process. It should be understood that duplicate samples can be ignored during the above sampling process of unlabeled transaction data points from a particular cluster.

[0081] Figure 10A Depicts an exemplary flowchart showing the steps of an evaluation process according to an embodiment, which is performed to determine the optimal sample size of a machine learning model. Figure 10B The processing depicted in can be implemented in software (e.g., code, instructions, programs) executed by one or more processing units (e.g., processors, cores) of a corresponding system, hardware, or a combination thereof. The software can be stored on a non-transitory storage medium (e.g., on a memory device). In Figure 10B The methods presented and described below are intended to be illustrative and non-limiting. Although Figure 10B depicts various processing steps occurring in a particular sequence or order, this is not intended to be limiting. In certain alternative embodiments, the steps can be performed in some different order, or some steps can also be performed in parallel.

[0082] The process begins at step 1001, where an initial value of the sample size (N) is determined. For example, the value of the sample size can be set to N = 10 x C, where C corresponds to the number of clusters in the encoded data set (e.g., Figure 3 the first encoded data set 305 and the second encoded data set 307). Additionally, in step 1001, the value of the iteration counter (K) is initialized to 1. In step 1003, the process evaluates the performance of a machine learning model with a sample size of N.

[0083] In step 1005, a query is implemented to determine whether the sample size needs to be increased. Note that as previously referenced Figure 3 and Figure 10B such a determination can be made based on verifying whether the performance of the machine learning model has reached a steady state. If the response to the query is affirmative (i.e., the sample size needs to be increased), the process moves to step 1007. However, if the response to the query is negative (i.e., the sample size does not need to be increased), the process moves to step 1011, where the value of the sample size (N) is output as the optimal sample size.

[0084] The process in step 1007 continues to increase the sample size (N). In some embodiments, the sample size can be increased by a value of N x 2 K . In step 1009, the value of the iteration counter (K) is incremented by one, after which the process loops back to step 1003 to perform the next sampling iteration by the first sampling module and the second sampling module. It should be understood that through one example, samples with an increased number obtained from the clusters of the first encoded data set and the second encoded data set can be obtained in a unified manner (i.e., N x 2 K ). However, this in no way limits the scope of the present disclosure. Instead, other mechanisms (e.g., weighted methods) can also be applied to determine how many samples will be obtained from the first sampling module and the second sampling module.

[0085] Embodiments of the present disclosure provide an encoding framework for differentiating different transaction behaviors. Features of the present disclosure also provide two different sampling techniques, namely, hard negative sampling technique and diversity sampling technique, which each operate on different encoded data sets obtained from an encoder. The encoding and sampling frameworks provided here have many advantages. For example, it enables a reduction in the size of the training data used to train a machine learning model. Such a reduction in the size of the training data (without significantly degrading the performance of the model) provides advantages such as a reduction in the total storage space required to store the data and a reduction in computational resources. Additionally, as previously mentioned, collecting, cleaning, and labeling large data sets can be expensive and time-consuming. Therefore, the reduction in the size of the training data in various aspects of the present disclosure also provides a means to train a machine learning model in an economically viable manner.

[0086] Any software component or function described in this application can be implemented as software code executed by a processor using any suitable computer language such as Java, C, C++, C#, Objective-C, Swift or a scripting language such as Perl or Python using, for example, conventional or object-oriented techniques. The software code can be stored on a computer-readable medium as a series of instructions or commands for storage and / or transmission, suitable media including random access memory (RAM), read-only memory (ROM), magnetic media such as hard disk drives or floppy disks, or optical media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, and the like. The computer-readable medium can be any combination of such storage or transmission devices.

[0087] Such programs can also be encoded and transmitted using a carrier signal suitable for transmission via a wired network, optical network, and / or wireless network conforming to various protocols including the Internet. Thus, a computer-readable medium according to an embodiment of the present invention can be created using a data signal encoded with such a program. The computer-readable medium encoded with program code can be packaged together with a compatible device or provided separately from other devices (e.g., downloaded via the Internet). Any such computer-readable medium can reside on or within a single computer product (e.g., a hard disk drive, a CD, or an entire computer system), and can exist on or within different computer products in a system or network. The computer system can include a monitor, a printer, or other suitable display for providing any of the results mentioned herein to a user.

[0088] The above description is illustrative and not restrictive. Many variations of the present invention will become apparent to those skilled in the art after reading this disclosure. Therefore, the scope of the present invention should not be determined with reference to the above description, but should be determined with reference to the pending claims together with their full scope or equivalents.

[0089] Without departing from the scope of the present invention, one or more features of any embodiment can be combined with one or more features of any other embodiment.

[0090] As used herein, unless expressly indicated to the contrary, the use of "a", "an", or "the" is intended to mean "at least one / kind".

Claims

1. A method, comprising: encoding, by a server computer, transaction data including unlabeled transaction data and fraudulent transaction data, the encoding forming encoded transaction data including: (i) a first set of encoded data associated with fraudulent transactions and unlabeled transactions similar to fraudulent transactions, and (ii) a second set of encoded data associated with unlabeled transactions; performing, by the server computer, a first sampling process on the first set of encoded data to obtain a first sampled set of encoded data; performing, by the server computer, a second sampling process on the second set of encoded data to obtain a second sampled set of encoded data; evaluating, by the server computer, the performance of a machine learning model that classifies the first sampled set of encoded data and the second sampled set of encoded data; and responsive to the performance of the machine learning model not meeting a condition, repeating, by the server computer, the first sampling process and the second sampling process by increasing the sizes of the first sampled set of encoded data and the second sampled set of encoded data until the condition is met.

2. The method of claim 1, wherein the unlabeled transaction data includes a first plurality of unlabeled transaction data points and a second plurality of unlabeled transaction data points, and the fraudulent transaction data includes a third plurality of fraudulent transaction data points, and wherein each of the first plurality of unlabeled transaction data points, the second plurality of unlabeled transaction data points, and the third plurality of fraudulent transaction data points is encoded as an N-dimensional vector in a vector space.

3. The method of claim 2, wherein the first plurality of unlabeled transaction data points and the third plurality of fraudulent transaction data points are encoded to be included in the first set of encoded data, and wherein the second plurality of unlabeled transaction data points are encoded to be included in the second set of encoded data.

4. The method of claim 2, wherein the first unlabeled transaction is determined to be similar to a fraudulent transaction based on the first encoded unlabeled transaction data point corresponding to the first unlabeled transaction being within a predetermined distance of a second plurality of encoded fraudulent transaction data points in the vector space.

5. The method of claim 2, wherein the first sampling process includes: calculating, by the server computer, a centroid for the third plurality of fraudulent transaction data points; for each unlabeled transaction data point included in the first plurality of unlabeled transaction data points, calculating a first distance of the unlabeled transaction data point to the centroid; and and generating a set of unlabeled transaction data points by filtering the first plurality of unlabeled transaction data points based on their respective first distances to the centroid.

6. The method of claim 5, further comprising: For each unlabeled transaction data point included in the set of unlabeled transaction data points compute a second distance of the unlabeled transaction data point to each fraudulent data point included in the third plurality of fraudulent transaction data points; determine a minimum distance based on the respective second distances associated with the unlabeled transaction data point; and for the unlabeled transaction data point, set the minimum distance as the separation distance of the unlabeled transaction data point from the third plurality of fraudulent transaction data points.

7. The method according to claim 6, further comprising: for each unlabeled transaction data point included in the set of unlabeled transaction data points, assign a ranking based on the separation distance of the unlabeled transaction data point; and based on the assignment, sample a first sample size of unlabeled transaction data points from the first sampled and encoded data set.

8. The method according to claim 3, further comprising: generating, by the server computer, one or more clusters of unlabeled transaction data points from the second plurality of unlabeled transaction data points; for each cluster, assign a seed data point to the cluster, the seed data point corresponding to an average of the unlabeled transaction data points in the cluster; and sample the cluster based on the seed data point.

9. The method according to claim 8, further comprising: for each unlabeled transaction data point included in the cluster, compute a probability corresponding to sampling the unlabeled transaction data point from the cluster; and sample a second sample size of unlabeled transaction data points based on the computation.

10. The method according to claim 9, wherein the probability associated with the unlabeled transaction data point is computed as a ratio of a third distance of the unlabeled transaction data point to the seed data point of the cluster to a sum of the respective distances of the other unlabeled transaction data points in the cluster to the seed data point of the cluster.

11. The method according to claim 9, wherein a first unlabeled transaction data point positioned farther from the seed data point as compared to a second unlabeled transaction data point is associated with a higher probability of being selected from the cluster as compared to the second unlabeled transaction data point.

12. The method according to claim 1, wherein the condition corresponds to a point at which the performance of the machine learning model starts to level off with respect to the sample size.

13. The method according to claim 1, wherein the first sampling process is different from the second sampling process.

14. A server computer, comprising: a processor; and a non-transitory computer-readable medium coupled to the processor and including code executable by the processor to implement a method including the following: Encode transaction data, the transaction data including unlabeled transaction data and fraudulent transaction data, the encoding resulting in encoded transaction data, the encoded transaction data including: (i) a first set of encoded data associated with fraudulent transactions and unlabeled transactions similar to fraudulent transactions, and (ii) a second set of encoded data associated with unlabeled transactions; Perform a first sampling process on the first set of encoded data to obtain a first sampled set of encoded data; Perform a second sampling process on the second set of encoded data to obtain a second sampled set of encoded data; Evaluate the performance of a machine learning model that classifies the first sampled set of encoded data and the second sampled set of encoded data; and In response to the performance of the machine learning model not meeting a condition, repeat the first sampling process and the second sampling process by the server computer by increasing the sizes of the first sampled set of encoded data and the second sampled set of encoded data until the condition is met.

15. The server computer according to claim 14, wherein the unlabeled transaction data includes a first plurality of unlabeled transaction data points and a second plurality of unlabeled transaction data points, and the fraudulent transaction data includes a third plurality of fraudulent transaction data points, and wherein Each of the first plurality of unlabeled transaction data points, the second plurality of unlabeled transaction data points, and the third plurality of fraudulent transaction data points is encoded as an N-dimensional vector in a vector space.

16. The server computer according to claim 15, wherein the first plurality of unlabeled transaction data points and the third plurality of fraudulent transaction data points are encoded to be included in the first set of encoded data, and wherein the second plurality of unlabeled transaction data points are encoded to be included in the second set of encoded data.

17. The server computer according to claim 15, wherein the first unlabeled transaction is determined to be similar to a fraudulent transaction based on the first encoded unlabeled transaction data point corresponding to the first unlabeled transaction being within a predetermined distance of the second plurality of encoded fraudulent transaction data points in the vector space.

18. The server computer according to claim 15, wherein the first sampling process includes: Calculating, by the server computer, a centroid for the third plurality of fraudulent transaction data points; For each unlabeled transaction data point included in the first plurality of unlabeled transaction data points, calculating a first distance of the unlabeled transaction data point to the centroid; And Generating a set of unlabeled transaction data points by filtering the first plurality of unlabeled transaction data points based on their respective first distances to the centroid.

19. The server computer according to claim 18, further comprising: For each unlabeled transaction data point included in the set of unlabeled transaction data points calculate a second distance of the unlabeled transaction data point to each fraudulent data point included in the third plurality of fraudulent transaction data points; determine a minimum distance based on the respective second distances associated with the unlabeled transaction data point; and for the unlabeled transaction data point, set the minimum distance as the separation distance between the unlabeled transaction data point and the third plurality of fraudulent transaction data points.

20. The server computer according to claim 14, wherein the condition corresponds to a point at which the performance of the machine learning model begins to level off with respect to the sample size.